39 Commits
Author SHA1 Message Date
Seth Jennings 4c1b16a4a1 feat(server): add operator-only provider credential retrieval (#4357)
* feat(server): add operator-only provider credential retrieval

Authenticate operators exclusively through verified direct gateway mTLS. Export selected runtime credentials with freshness validation and all-or-nothing delivery.

Coordinate refresh across callers and replicas, preserve rotated grants across RPC cancellation, and fence credential delivery against reconfiguration. Add generated bindings, security and regression tests, and operator documentation.

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* feat(sdk): add curated operator credential retrieval

Expose provider credential clients in Python and TypeScript with explicit workspace and key selection, freshness options, optional expiration, and preserved gateway errors. Reuse existing mTLS transports without introducing retries or secret caching.

Add generated-wire and error-mapping tests, SDK package validation, and operator usage documentation.

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(server): honor recovery state for queued automatic refreshes

Signed-off-by: Seth Jennings <sjenning@redhat.com>

---------

Signed-off-by: Seth Jennings <sjenning@redhat.com>
2026-10-09 16:16:47 +00:00
Emilien Macchianddivesh 8b3cc3fdc0 feat(server): add gateway capacity metrics and optional HPA (#3978)
* feat(server): add gateway capacity metrics and optional HPA

Expose per-replica supervisor sessions, pending relay capacity, relay
rejections and claim latency, and outbound peer request outcomes and
latency. Use bounded labels and Prometheus histograms for the new latency
metrics while preserving existing summary metrics.

Add an optional Helm HPA with external-database and resource validation,
conservative scale-down defaults, and support for custom metrics. Keep
certificate hook pods outside gateway workload selectors.

Document per-pod scraping, scaling limits, upgrade behavior, and
PostgreSQL connection sizing.

Part of #3528

Signed-off-by: Emilien Macchi <emacchi@redhat.com>

* feat(server): unify routing metrics and clarify replica capacity

Combine local relay setup and outbound peer requests in one counter, labeled by operation, route, target, outcome, and gRPC status. Preserve peer latency metrics.

Rename the rejection reason from global_capacity to replica_capacity to reflect the per-replica relay budget. Keep capacity limits unchanged.

Update tests and documentation.

Signed-off-by: divesh <dgude@nvidia.com>

* fix(server): align routed request metric labels and outcomes

Count a local relay as successful only when its supervisor claims it, the
same event that answers a peer relay on the owner, and rename the outcomes
to local_error and remote_error so they say where an attempt failed. Label
the peer latency histogram by operation, like the routed request counter,
and rename the target label to relay_kind so it does not read as the
Prometheus scrape target.

Rename RelayCapacity.global to per_replica, drop the per-sandbox relay
capacity gauge, which no per-sandbox series can pair with, describe the
latency histogram as peer-only, and restore the note that unavailable
spikes are expected during rollouts.

Part of #3528

Signed-off-by: Emilien Macchi <emacchi@redhat.com>

---------

Signed-off-by: Emilien Macchi <emacchi@redhat.com>
Signed-off-by: divesh <dgude@nvidia.com>
Co-authored-by: divesh <dgude@nvidia.com>
2026-10-05 15:38:40 +00:00
0a770d9173 feat(kubernetes): support HA gateway rebalancing (#1868)
* feat(kubernetes): support HA gateway rebalancing

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* perf(server): cache peer connections, tokens, and owner lookups

Every forwarded relay rebuilt its setup from scratch: an owner lookup, a
blocking read of the peer token, a TLS connect to the owning replica, and
a TokenReview plus Pod GET on the receiving side. Sandbox service routing
does this per HTTP request, so the apiserver calls scaled with traffic.

Cache all of it on ServerState:

- peer channels pooled per endpoint, so relays multiplex over one
  connection instead of redialing
- peer tokens keyed by SHA-256, expiring at min(ttl, token exp) so a hit
  cannot accept an expired token
- owner records for 3s against a 45s ownership TTL, still freshness
  checked before use

Entries are evicted when a relay fails. Also raise HTTP/2
max_concurrent_streams to 1024, since pooling funnels every relay between
two replicas onto one connection and hyper's default of 200 sits below
the 256 pending-relay budget.

Signed-off-by: divesh <dgude@nvidia.com>

* perf(server): pool upstream connections for sandbox services

Each HTTP request to a sandbox service opened its own supervisor relay,
paying a new TCP connection and HTTP/1 handshake every time. Worse, it
counted against the 32 in-flight relay cap, so a service handling more
than 32 concurrent requests failed outright.

Pool idle upstreams per endpoint and port, up to 8 each for 15s. Reuse is
safe because the pool only returns a connection hyper reports as ready,
and HTTP/1 cannot start a request until the previous body has drained.
Upgrades are never pooled since they take the connection over, and a
failed send evicts that endpoint. Pruning is bounded per key, with the
full sweep limited to once per 30s.

Signed-off-by: divesh <dgude@nvidia.com>

* fix(server): address HA gateway review findings (#3449)

- Let a gateway own supervisor sessions without a peer endpoint. Requiring
  one whenever the store is PostgreSQL broke every single-instance
  PostgreSQL deployment, because no sandbox supervisor could connect.
  A cross-replica request to an owner that advertises no endpoint now fails
  immediately naming the cause, instead of retrying until the wait timeout.
- Close a supervisor session on heartbeat only when another replica owns it,
  or after renewals fail for the ownership TTL. A database error no longer
  drops every session heartbeating during an outage.
- Clamp owner record ages at zero so a skewed or corrupt stored timestamp
  cannot produce a negative age.
- Bound the cross-object advisory lock with a lock timeout, so a stuck holder
  fails instead of blocking every mutation in the fleet.
- Refuse to start when a peer endpoint is configured on a multi-replica
  backend but peer authentication is unavailable, and warn when a
  multi-replica backend has no peer endpoint at all.
- Reject a plaintext peer endpoint when the gateway serves TLS.
- Skip the sandbox watch poller on single-replica backends, where the local
  update bus already sees every write.
- Rate-limit the peer owner cache sweep so an insert no longer scans the
  whole map under the lock.
- Retry GET and HEAD on a pooled upstream the sandbox closed, instead of
  returning 502, and drop an emptied endpoint from the pool right away.
- Document the gateway peer environment variables and the post-rollout
  ownership skew operators should expect.

Signed-off-by: divesh <dgude@nvidia.com>

* fix(server): harden HA supervisor ownership

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Signed-off-by: divesh <dgude@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: divesh <dgude@nvidia.com>
Co-authored-by: Divesh Chowdary <47188680+FrostGod@users.noreply.github.com>
2026-09-21 21:22:12 +00:00
John T. Myers 1d010f4187 feat(sandbox): validate configuration before workload activation (#3259)
* feat(sandbox): validate configuration before workload activation

Closes #3145

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(sandbox): bound startup failures and preserve activation history

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(sandbox): enforce deadlines on startup RPC attempts

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(sandbox): isolate provider auto-create policy fixture

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(sandbox): remove unnecessary fixture string delimiters

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(sandbox): capture startup logs before ephemeral cleanup

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(sandbox): preserve credential revocation after rebase

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(server): retain provider revision helper for endpoint reports

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(server): import provider object trait in production

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(supervisor): preserve admission across boundary extraction

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(sandbox): use existing boundary discovery imports

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(sandbox): distinguish workload and supervisor containers

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(server): reconcile admission schema with timestamp migration

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(server): reconcile admission with typed deletion schema

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(server): reconcile admission with mutation request IDs

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(supervisor): adapt local startup fixture to admission state

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(supervisor): deduplicate startup quarantine diagnostics

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(tui): show configuration blockers in sandbox notes

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* feat(sandbox): expire provisioning repair attempts after five minutes

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(supervisor): reconcile admission with provider readiness

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(tui): separate configuration summaries from full diagnostics

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(tui): shorten invalid configuration note

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(server): refresh schema inventory after rebase

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-09-18 00:16:18 +00:00
Shiju 9708ba9999 feat(providers): report applied sandbox provider changes (#3391)
* feat(providers): report applied sandbox provider changes

Record exact provider mutation targets in shared configuration operations.
Require authenticated evidence that credentials, effective policy, and the
workload launch environment have been installed before reporting readiness.

Add bounded CLI and Rust SDK status and wait support, preserving ordinary
revision-scoped references for existing processes. Verify new-client rotation
and acknowledged detach revocation without external-stable resolver changes.

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(providers): align readiness times with protobuf contracts

Represent readiness receipts, status, and operation times with Timestamp
and report intervals with Duration. Reserve the scalar field tags, update
all consumers and generated bindings, and preserve timestamp presence
and nanosecond identity through storage and client validation.

Qualify both empty-map constructors in the Linux boundary test so its
module compiles while retaining the explicit default required by Clippy.

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(cli): preserve provider mutation storage uncertainty

Recognize the gateway's exact structured storage-uncertainty reason for
provider attach, detach, and update. Explain that the change may already
be saved and must be reconciled before retrying, without exposing server
messages or metadata. Preserve uncertainty ahead of generic retry hints.

Exercise saved mutations through the CLI and verify single submission,
redaction, missing receipt handling, and untrusted error-detail rejection.
Document the recovery guidance for users and the public CLI skill.

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(cli): explain denied provider profile lookups

Report exact and alias profile lookup denials with fixed permission and workspace guidance. Keep backend details redacted and stop before provider mutations.

Cover denied create and update calls through the CLI. Verify the complete provider list independently in the cross-workspace OIDC regression, extracting its JSON object from surrounding startup diagnostics.

Signed-off-by: Shiju <shiju@nvidia.com>

---------

Signed-off-by: Shiju <shiju@nvidia.com>
2026-09-17 19:10:43 +00:00
Mrunal Patel c502be9fd7 feat(api): return typed deletion outcomes with explicit missing-target semantics (#3317)
* feat(api): return typed deletion outcomes with explicit missing-target semantics

Implement phase 2 of #3051 across the public gateway API and first-party SDKs. Preserve asynchronous sandbox deletion and observed resource identities, reserve legacy wire fields, and document the coordinated migration.

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(sdk): bind deletion waits to sandbox identity

Track accepted deletions by original sandbox identity in Rust and TypeScript. Preserve name-only waits and cover replacement races and lookup failures.

Merge the phase-one cleanup fix and adapt its interceptor regressions to typed deletion outcomes.

Refs #3051.

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* test(e2e): accept asynchronous stopped sandbox deletion

Allow either completed or accepted deletion output, then continue polling for actual sandbox absence. This preserves the lifecycle assertion under the phase-two deletion contract.

Refs #3051.

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

---------

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
2026-09-16 22:11:19 +00:00
Shiju fd3fd9cf74 feat(sandbox): explain failed calls to external tool servers (#3207)
Show configured tool server addresses and their last observed connection
results together in sandbox status. Keep sandbox lifecycle readiness
separate so an external connection failure does not mark the sandbox unready.

Expose direct endpoint records through the CLI and SDKs, with plain-language
failure explanations and gateway acceptance times. Keep observation tracking,
runtime reporting, and gateway validation in dedicated endpoint status modules.

Preserve bounded reporting, request attribution, retry ordering, and
configuration and supervisor authority checks. Clear obsolete observations
while retaining the configured addresses, and document the distinction
between an observed HTTP response, current availability, and tool success.

Signed-off-by: Shiju <shiju@nvidia.com>
2026-09-15 13:46:38 +00:00
Simon Scatton 48c449d8c8 chore(deps): replace ring with AWS-LC (#3243)
* chore(deps): replace ring with AWS-LC

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* fix(lint): address warnings after dependency upgrades

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* fix(tls): limit provider initialization to reqwest clients

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

---------

Signed-off-by: Simon Scatton <sscatton@nvidia.com>
2026-09-09 17:57:42 +00:00
Philippe Martin 118b250f01 feat(sandbox): support rootfs tar as --from source for VM driver (#2863)
* feat(sandbox): support rootfs tar as --from source for VM driver

Accept flat rootfs tar archives (.tar, .tar.gz, .tgz) via the --from
flag for VM-backed gateways. The CLI detects the archive extension,
validates that the gateway uses the VM compute driver, and passes the
tar path through driver_config. The VM driver copies the tar into its
staging area and feeds it into the existing rootfs extraction and ext4
disk creation pipeline, skipping the container image pull/export steps.

Closes #2175

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(sandbox): validate rootfs tar path at the VM driver boundary

The rootfs_tar_path field in driver_config was passed from the API
caller directly to tokio::fs::copy without validation. An authenticated
user bypassing the CLI could supply arbitrary host paths (e.g.
/dev/zero for disk exhaustion, or readable host files for data
exfiltration).

Introduce a trusted staging directory that the VM driver creates on
startup and advertises via GetCapabilities. The CLI now copies the tar
into the staging directory before creating the sandbox, and the driver
validates that the received path is a regular file inside the staging
root and within a configurable size limit (default 10 GiB) before any
I/O.

New VmDriverConfig options:
- rootfs_tar_staging_dir: override the staging directory
  (default: <state_dir>/rootfs-tar-staging)
- rootfs_tar_max_bytes: override the size limit (default: 10 GiB)

Addresses GATOR-28b5152e-01.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(sandbox): request-scoped staging, size pre-check, and cleanup for rootfs tar

Tighten the rootfs tar staging flow to address the remaining GATOR-01
obligations:

- Request-scoped staging: the CLI creates a unique per-request
  subdirectory (req-<pid>) under the staging root instead of placing
  files directly in the shared directory. The driver enforces that the
  tar path is at depth 2 (staging_root/<subdir>/<file>), preventing
  cross-request path selection.

- Size pre-check: the driver advertises rootfs_tar_max_bytes via
  GetCapabilities. The CLI reads this limit and rejects oversized files
  before copying, avoiding disk exhaustion in the staging directory.

- Cleanup: the driver removes the request staging subdirectory after
  consuming the tar (on cache hit, copy success, or copy failure),
  ensuring staged data does not persist beyond the request.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(vm): restore rootfs-tar sandboxes from persisted image identity

On restore or restart, the one-shot staged tar archive has already been
cleaned up. Reading the persisted image identity from the sandbox state
directory and resolving the cached disk path directly avoids re-accessing
the deleted staging path.

Addresses GATOR-168b9210-01.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(cli): use random staging dirs and enforce byte limit during rootfs tar copy

Replace PID-based request staging directories with tempfile-generated
random names to prevent collisions and make paths unpredictable.
Replace bare tokio::fs::copy with a streaming copy loop that enforces
the advertised max_bytes limit during transfer, closing the TOCTOU gap
between the pre-copy size check and the actual copy.

Signed-off-by: Philippe Martin <phmartin@nvidia.com>
Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(sandbox): issue rootfs tar staging slots from the gateway

A caller could name any host path in `driver_config.vm.rootfs_tar_path`,
which the privileged VM driver then read. The CLI-side locality check did
not apply to direct API requests.

The gateway now owns staging. `BeginRootfsTarStaging` allocates a
request-scoped directory under the driver-advertised staging root and
returns an opaque single-use token; `CreateSandbox` carries the token, and
the gateway substitutes the path it allocated before dispatching to the
driver. `template.driver_config.<driver>.rootfs_tar_path` is rejected
outright in request validation, so a caller-supplied path never reaches
privileged I/O.

Tokens are bound to the issuing workspace and subject, consumed once, and
expire after 30 minutes. Outstanding slots are capped per caller and
overall, so one caller can neither exhaust the staging filesystem nor
starve others. An RAII guard reclaims the directory on every failure path
after consumption, and an age-gated sweep runs at startup and on each
reconcile pass for directories whose driver died before its own cleanup.

The token is stripped from the public sandbox before persistence: the
stored copy is returned verbatim by GetSandbox, ListSandboxes and
WatchSandbox to every member of the workspace.

Also fixes two defects this exposed:

- The CLI wrote `rootfs_tar_path` at the top level of `driver_config`, but
  the gateway forwards only `driver_config.<driver_name>`, silently
  dropping unmatched keys. The archive never reached the VM driver, so the
  documented `--from ./rootfs.tar` flow did not work at all. Config is now
  nested under `vm` and deep-merged, so a caller's existing VM settings
  survive instead of being clobbered by a shallow extend.

- Staging previously required `GetGatewayInfo`, which is restricted to
  `platform_admin`, making the feature unusable for ordinary users on any
  RBAC-enabled gateway. The new RPC matches CreateSandbox at
  `sandbox:write` / `workspace_role: user`.

`compute_driver.proto` is unchanged; the gateway reads the staging root
from the capabilities it already stores.

Refs #2175

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(vm): derive rootfs tar cache identity from archive contents

The prepared-disk cache key combined the archive's full path with an mtime
truncated to seconds, then mapped punctuation to `-`. Distinct paths such as
`/tmp/a/b.tar` and `/tmp/a-b.tar` collapsed onto the same key and reused each
other's disk, a rewrite within the same second kept stale contents, and a long
path could exceed filesystem component limits.

Identity is now a SHA-256 of the archive contents. This is also what makes the
cache work at all now that the gateway allocates a fresh staging directory per
request: a path-derived key would miss on every create.

The archive is hashed, the cache checked, and only on a miss copied — so a hit
skips writing a multi-gigabyte file. The copy is hashed as it is written and
rejected if the digest differs from the first pass, which closes the window
where the source changes during staging rather than approximating it with a
re-stat.

Refs #2175

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(vm): decompress gzip rootfs tar archives during staging

`--from` accepts `.tar.gz` and `.tgz`, but the driver staged whatever bytes
it was given and the guest image-prep VM extracts the staged file with a
plain `tar -xpf`. Compressed sources therefore depended on the guest tar
auto-detecting gzip, and the prepared disk was sized from the compressed
length, which is far too small for the expanded rootfs.

Staging now detects gzip from the archive's magic bytes -- the driver only
ever sees a gateway-issued staging path, never the caller's file name -- and
writes an uncompressed tar. The digest still covers the source bytes, so the
"archive changed while staging" check is unaffected, and expansion is bounded
by `rootfs_tar_max_bytes` so a compression bomb cannot fill the host disk.

`extract_rootfs_archive_to` sniffs gzip as well, so the host-side extraction
path matches.

Adds unit coverage for gzip staging, bounded expansion, and gzip extraction,
plus an e2e sandbox created from a gzip-compressed export.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

---------

Signed-off-by: Philippe Martin <phmartin@redhat.com>
Signed-off-by: Philippe Martin <phmartin@nvidia.com>
2026-09-08 22:36:40 +00:00
grs 74960ebfae feat(server): add sandbox templates (#2833)
* feat(server): add sandbox workload templates

Signed-off-by: Gordon Sim <gsim@redhat.com>

* feat(go-sdk): add sandbox workload template support

Signed-off-by: Gordon Sim <gsim@redhat.com>

* feat(rust-sdk): add sandbox workload template support

Signed-off-by: Gordon Sim <gsim@redhat.com>

* feat(python-sdk): add sandbox workload template support

Signed-off-by: Gordon Sim <gsim@redhat.com>

* feat(typescript-sdk): add sandbox workload template support

Signed-off-by: Gordon Sim <gsim@redhat.com>

* docs(agents): document sandbox workload templates

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(cli): support default GPU requests in sandbox templates

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(server): cap sandbox templates per workspace

Signed-off-by: Gordon Sim <gsim@redhat.com>

* docs(architecture): document sandbox workload template boundaries

Signed-off-by: Gordon Sim <gsim@redhat.com>

* feat(cli+sdk): expose sandbox workload template provenance

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(cli): include sandbox template annotations in output

Signed-off-by: Gordon Sim <gsim@redhat.com>

* feat(sandbox): add label selectors to template listing

Signed-off-by: Gordon Sim <gsim@redhat.com>

* test(e2e): cover sandbox template failure paths

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(server): preserve command and ttl when creating sandbox from template

Signed-off-by: Gordon Sim <gsim@redhat.com>

* test(server): add field coverage test for template merge

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(go-sdk): add pagination support to fake client

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(cli): warn on env vars that looks like secrets

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(sdk-ts): propagate sandbox workspace through lifecycle calls

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(docs): update workspace management docs

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(ts-sdk): support command and tty when creating from template

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(go-sdk): support command and tty when creating from template

Signed-off-by: Gordon Sim <gsim@redhat.com>

* test(python-sdk): verify command and tty handling when creating from template

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(go-sdk): update docs and ClientInterface

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(server): validate sandbox create specs before I/O

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(go-sdk): guard empty DNS-1123 label validation

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(python-sdk): allow empty template builder mappings

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(cli): align template GPU JSON default output

Signed-off-by: Gordon Sim <gsim@redhat.com>

---------

Signed-off-by: Gordon Sim <gsim@redhat.com>
2026-09-02 00:55:29 +00:00
Drew Newberry 69a05ebb3b fix(sandbox): complete successful main processes (#2884) 2026-08-28 18:21:20 -07:00
grs 40d1b48666 feat(provider): support for SPIFFE backed token exchange (#1970)
* feat(provider): add ability to request token exchange instead of client credentials as OAuth grant_type

Signed-off-by: Gordon Sim <gsim@redhat.com>

* test(proxy): add further tests for token exchange

Signed-off-by: Gordon Sim <gsim@redhat.com>

* test(provider): add runnable example for token exchange

Signed-off-by: Gordon Sim <gsim@redhat.com>

* test(e2e): cover Podman token exchange grants

Signed-off-by: Gordon Sim <gsim@redhat.com>

* refactor(oauth): extract duplicated functionality from server and supervisor

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(provider): evict nearest-to-expiry entry from intermediate token cache

Signed-off-by: Gordon Sim <gsim@redhat.com>

* doc(supervisor): add podman example for token exchange

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(provider): withhold token-exchange subject credentials

Signed-off-by: Gordon Sim <gsim@redhat.com>

---------

Signed-off-by: Gordon Sim <gsim@redhat.com>
2026-08-24 05:42:30 +00:00
Drew Newberry ef296806f5 feat(sandbox): add canonical main process (#2726)
* feat(sandbox): add canonical main process

Closes #2710

Persist and supervise one canonical workload per sandbox, attach sandbox connect to its retained session, and make every unexpected main-process exit terminal.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(sandbox): simplify canonical main process contract

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): preserve legacy VM main compatibility

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): preserve main status across driver updates

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): satisfy macOS process lint

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): gate Linux exit acknowledgement publisher

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(sandbox): simplify main process plumbing

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(supervisor): make controlling tty ioctl portable

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(supervisor): initialize canonical process environment

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* perf(supervisor): optimize retained main session

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(sandbox): detach main session on ctrl-c

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): use explicit main detach keys

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(dev): atomically stage Docker supervisor

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(test): align Docker main environment assertion

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(sdk): expose canonical main process fields

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-08-20 22:18:23 +00:00
Jesse Jaggars c4b500a7de feat(helm): cert-manager external issuer + OpenShift passthrough Route (#2468)
* fix(core): trust public root CAs alongside the sandbox mTLS CA

The supervisor gRPC client only trusted the CA configured via
OPENSHELL_TLS_CA, since tonic ClientTlsConfig starts with an empty root
store unless with_native_roots()/with_webpki_roots() is also enabled.
Deployments where the gateway server certificate is issued by a public CA
(e.g. cert-manager against an ACME issuer) caused every supervisor
connection to fail the TLS handshake with "UnknownCA", since the sandbox
mTLS CA and the server cert issuer were no longer the same.

Enable both native and webpki roots in addition to the configured CA.
tonic root store is a union of all configured sources, so this does not
weaken verification for existing self-signed deployments. webpki-roots
(compiled in) is enabled alongside native-roots since the supervisor
binary may run in minimal sandbox images without a populated system CA
bundle.

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* feat(helm): support external cert-manager issuers and OpenShift Route passthrough

Add certManager.serverIssuerRef/clientIssuerRef so the gateway and mTLS
client certificates can be issued by a real Issuer/ClusterIssuer (e.g.
ACME) instead of only the chart built-in self-signed CA.

Add openshiftRoute template for exposing the gateway via a TLS
passthrough Route so the gateway keeps terminating its own TLS/mTLS.

The server Certificate excludes internal-only SANs (cluster-local,
localhost, loopback) when an external issuer is configured, since ACME
issuers reject those per CA/Browser Forum baseline requirements. A
template-time fail guard catches the misconfiguration at helm install
time rather than asynchronously at cert-manager issuance time.

Includes Helm unittest coverage for both issuerRef overrides and Route
rendering, plus a CI values overlay for lint coverage.

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* docs: document cert-manager external issuer and OpenShift Route

Update managing-certificates.mdx with the serverIssuerRef workflow and
install-time validation behavior. Add a production section to the
OpenShift guide covering passthrough Route with a real certificate.
Regenerate Helm README for new certManager and openshiftRoute values.
Sync debug-openshell-cluster skill with new troubleshooting steps for
ACME issuance failures and supervisor UnknownCA from mismatched CAs.

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(helm,core): address PR review feedback on cert-manager external issuer

Addresses all five blocking review items from #2468:

1. Remove .with_native_roots() from supervisor gRPC client -- the
   supervisor runs inside the user-selected sandbox image, so the
   image CA bundle is not operator-controlled. Keep .with_webpki_roots()
   (compiled-in, not user-controlled) alongside the configured CA.

2. Fail at render time when serverIssuerRef.name is set but
   clientCaFromServerTlsSecret is still true. Add negative Helm test.

3. Remove clientIssuerRef -- changing only clientIssuerRef breaks both
   directions because trust bundles are not modeled separately. Change
   serverIssuerRef.kind default from ClusterIssuer to Issuer.

4. Add server.oidc.issuer and server.oidc.audience to the documented
   OpenShift production Helm command. Add Access Control prerequisite.

5. Fail at render time when openshiftRoute.enabled and disableTls are
   both true. Add negative Helm test.

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(drivers): strip GATEWAY_TLS_SERVER_NAME from Docker and Podman env

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(helm,drivers): guard default clientCaSecretName and add env-strip tests

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* feat(tls): SNI-based dual certificate for internal and external server TLS

Split the gateway server certificate into two: an internal cert issued by
the chart's own CA (for supervisor connections via cluster-local SANs) and
an external cert issued by an operator-configured Issuer such as ACME/Let's
Encrypt (for CLI and Route access via public SANs).

The gateway uses SNI-based certificate selection: connections whose SNI
hostname matches external_server_names receive the external cert; all
others (including those with no SNI) receive the internal cert.

Security improvement: remove .with_webpki_roots() from the supervisor
gRPC client so supervisors trust only the chart CA, closing a MITM vector
via publicly-trusted certificates in user-supplied container images.

Key changes:
- Add DualCertResolver with SNI-based cert selection and full test coverage
- Add external_cert_path, external_key_path, external_server_names to TlsConfig
- Validate partial external cert config (error on cert-without-key or vice versa)
- Validate empty external_server_names when external cert is configured
- Split cert-manager templates into internal + external Certificate resources
- Add Helm guards for misconfigured external issuer (empty serverDnsNames,
  internal-only SANs with external issuer, conflicting clientCaFromServerTlsSecret)
- Update gateway-config.mdx, managing-certificates.mdx, openshift.mdx docs
- Update debug-openshell-cluster skill for dual-cert troubleshooting

Signed-off-by: Pi Agent <agent@openshell.local>

* fix(drivers): strip GATEWAY_TLS_SERVER_NAME in VM driver and correct comments

Add the same GATEWAY_TLS_SERVER_NAME environment stripping to the VM
compute driver that Docker, Podman, and Kubernetes drivers already
perform. Without this, a sandbox user on the VM driver could override
the TLS server name the supervisor verifies.

Fix stale comments in Docker and Podman drivers that referenced
'with WebPKI roots trusted' — WebPKI roots are explicitly not trusted
after the tls-webpki-roots removal.

Use tls-ring instead of bare channel for tonic in openshell-core so the
TLS API (ClientTlsConfig, Endpoint::tls_config) is available without
pulling in any root certificate store.

Signed-off-by: Pi Agent <agent@openshell.local>

* fix(tls,helm): wildcard SNI matching and Route host validation

Add RFC 6125 single-level wildcard matching to DualCertResolver so
external_server_names entries like *.example.com correctly match SNI
hostnames like gw.example.com. Previously only exact matches worked,
silently falling back to the internal cert for wildcard configurations.

Add a Helm fail guard in route.yaml that rejects openshiftRoute.host
values not listed in certManager.serverDnsNames when an external issuer
is configured — catches cert/route hostname mismatches at install time
instead of at TLS connect time.

Quote the host field in route.yaml for robustness.

Signed-off-by: Pi Agent <agent@openshell.local>

* fix(helm): address blocking review items — client-CA guard, wildcard Route, serverIssuerRef gate

1. Remove the obsolete guard rejecting serverIssuerRef + clientCaFromServerTlsSecret=true.
   The internal server certificate is always signed by the chart CA (the same
   CA that signs the client cert), so clientCaFromServerTlsSecret=true is
   correct — its filtered ca.crt is exactly the right trust anchor.  The old
   workaround (mounting openshell-ca-tls directly) unnecessarily exposed the
   CA private key to the gateway container.  Remove the client-CA overrides
   from docs, CI overlay, and production examples.

2. Route host validation now supports wildcard certificates per RFC 6125:
   single-level wildcards like *.example.com match gateway.example.com but
   not deep.sub.example.com.  Require an explicit openshiftRoute.host when
   an external issuer is configured — without one, OpenShift generates a
   hostname absent from serverDnsNames.

3. Reject serverIssuerRef.name when certManager.enabled is false — the
   external certificate, its Secret mount, and the gateway TLS config all
   require cert-manager to be enabled.

Validated on ROSA (dev.dyee.p3) with branch-built images:
- Fresh install with letsencrypt-prod ClusterIssuer
- SNI dual-cert: external hostname served Let's Encrypt cert
- Supervisor mTLS via internal cert path: ConnectSupervisor accepted
- Client CA volume: filtered ca.crt from internal server secret (no key)
- CLI connected via Route + OIDC

Helm tests: 81 pass across 7 suites.

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

---------

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>
Signed-off-by: Pi Agent <agent@openshell.local>
2026-08-14 00:12:11 +00:00
Seth Jennings 0f8fad23c4 feat(sandbox): add stop and start operations (#2653)
* feat(sandbox): add suspend and resume operations

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(server): preserve lifecycle work after cancellation

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(server): reconcile ambiguous lifecycle outcomes

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(server): complete suspended session cleanup

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(vm): preserve suspension state on resume failure

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(server): retry retained lifecycle transitions

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(server): clean sessions after suspend reconciliation

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* test(sandbox): cover deleting suspended sandbox

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(kubernetes): preserve progressing sandbox suspension

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(kubernetes): bound suspend status polling

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(kubernetes): detect legacy sandbox suspension

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(tui): render suspended sandbox phases

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* refactor(sandbox): rename suspend and resume lifecycle

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* perf(server): clean stopped sessions on transition

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(kubernetes): fail fast on rejected stop

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(compute): fence stale restart lifecycle events

Signed-off-by: Seth Jennings <sjenning@redhat.com>

---------

Signed-off-by: Seth Jennings <sjenning@redhat.com>
2026-08-13 06:13:54 +00:00
Derek Carr 9c019a93f5 Wire authorization into workspace model (#2445)
* feat(auth): implement RFC 0011 Phase 2 workspace authorization

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(auth): address PR review feedback on workspace authorization

- Docker e2e: add --health-port and switch readiness probe from
  `openshell status` to `curl /healthz`, fixing a false-positive
  readiness check in OIDC mode where the CLI exited 0 without
  actually contacting the gateway

- ListWorkspaces: move membership filtering from post-query N+1
  lookups into a SQL EXISTS subquery so pagination applies to the
  visible set, not the global ordering. Add generic
  list_with_membership to the persistence layer.

- Descriptor validator: reject role/scope fields on unauthenticated
  and sandbox auth modes, and allow-list workspace_role as
  user/admin and global_role as platform_admin to catch typos at
  startup

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(server): use authed request in delete telemetry test

The workspace authorization added by the Phase 2 auth changes requires
a Principal on every delete request. The delete-telemetry test was still
using a bare Request::new, so extract_principal failed before the handler
could acquire the delete gate, causing a 5-second timeout flake.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(auth): address gator review findings for workspace authorization

- Inject unauthenticated-local-dev principal in no-auth gateway mode so
  handlers that call extract_principal() always find one.
- Cap label-selector membership query at MAX_PAGE_SIZE instead of
  u32::MAX to bound the in-memory read.
- Authorize workspace membership before resolving workspace existence in
  all sandbox RPCs to prevent workspace-name enumeration by non-members.
- Remove dead_code allow on AuthorizedWorkspace.workspace now that
  callers use the normalized name from the authz result.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(auth): close workspace-name oracle and label-selector truncation

Swap authorize-before-resolve ordering in 27 handlers across
provider.rs, service.rs, policy.rs, and workspace.rs to prevent
CWE-203 workspace-name enumeration by non-members.

Add combined membership+label SQL query (list_with_membership_and_selector)
to both persistence backends so ListWorkspaces with label selectors no
longer silently drops results beyond the first page of membership matches.

Signed-off-by: Derek Carr <decarr@redhat.com>

* test(auth): add non-member rejection and membership+label persistence tests

Add comprehensive test coverage for workspace authorization changes:
- Non-member rejection tests across all 44 workspace-scoped handlers
  (sandbox, provider, service, policy, workspace, inference) verifying
  PERMISSION_DENIED is returned instead of NOT_FOUND to prevent
  CWE-203 workspace-name oracle
- Persistence test for list_with_membership_and_selector verifying
  SQL-level membership EXISTS + label filtering, multiple predicates,
  no-match cases, and pagination

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(auth): format merged import line in sandbox tests

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(auth): address gator re-review findings on workspace authorization

- Fix TUI unconditionally setting providers_v2_enabled after provider
  refresh; read the actual gateway setting via GetGatewayConfig at
  startup instead
- Fix SQLite json_extract with dotted label keys (e.g. example.com/env)
  by quoting the key in the JSON path
- Add authed_request wrappers to upstream OCI identity tests that were
  missing a principal after rebase
- Add test proving GetGatewayConfig is accessible without Platform Admin
- Add test for dotted/prefixed Kubernetes-style label key filtering

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(auth): address second gator re-review findings

- Loosen GetGatewayConfig from platform_admin to scope-only so workspace
  users can discover providers_v2_enabled during sandbox creation with
  inferred-provider commands; update proto descriptor, descriptor
  validation, and RFC 0011 access table
- Add validate_label_selector to handle_list_workspaces and escape
  single quotes in SQLite json_extract interpolation (CWE-89
  defense-in-depth)
- Re-fetch providers_v2_enabled after TUI gateway switch so the new
  gateway's capability is reflected
- Add e2e test for workspace user with inferred-provider command
- Add persistence test for adversarial label keys with SQL injection
  attempts
- Add handler test for invalid label selector rejection in
  ListWorkspaces

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(auth): address third gator review findings

- Cap label selector pairs at 64 (CWE-400) to bound SQLite dynamic SQL
- Add SCOPE_ONLY_METHODS allowlist for scope-without-role RPCs (CWE-863)
- Normalize ID-based data-plane handlers to return NOT_FOUND for
  unauthorized sandboxes, closing the cross-workspace oracle (CWE-203)
- Fix TUI provider profile cache lookup key mismatch for legacy
  providers with empty profile_workspace
- Add whoami to CLI skill reference command tree
- Update TUI skill doc with workspace, provider, and settings coverage
- Document scope/workspace orthogonality on GetGatewayConfig proto

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(auth): extend CWE-203 normalization to policy.rs sandbox handlers

GetSandboxConfig and GetSandboxLogs in policy.rs had the same
fetch-before-authorize pattern that leaked cross-workspace sandbox
existence. Promote fetch_and_authorize_sandbox to pub(super) and
use it from both sandbox.rs and policy.rs handlers.

Signed-off-by: Derek Carr <decarr@redhat.com>

* test(auth): update assertions for CWE-203 sandbox ID normalization

Cross-workspace sandbox access via ID-based handlers now returns
NOT_FOUND instead of PERMISSION_DENIED to prevent existence inference.
Update the unit test and OIDC e2e assertion to match.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(auth): narrow CWE-203 error mapping and correct whoami output formats

Only remap PERMISSION_DENIED to NOT_FOUND in fetch_and_authorize_sandbox
and RevokeSshSession, letting INTERNAL and UNAUTHENTICATED propagate
as-is. Fix whoami --output format values in cli-reference.md to match
the actual CLI (table/json/yaml, not text/json).

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(ci): share network namespace with Keycloak in containerized CI

In GitHub Actions job containers, Docker port publishing lands on the
host, not inside the job container. Detect this environment and attach
Keycloak to the job container's network namespace instead, with
hardened defaults (cap-drop ALL, no-new-privileges, loopback-only
listener).

Signed-off-by: Derek Carr <decarr@redhat.com>

---------

Signed-off-by: Derek Carr <decarr@redhat.com>
2026-07-30 00:31:32 +00:00
Derek Carr 5952a5a23f feat(workspace): add workspace resource model with scoping, membershi… (#2243)
* feat(workspace): implement workspace model (Phase 1 of RFC 0011)

Implements workspace and membership model providing hard isolation
boundaries for multi-player OpenShell deployments.

Workspace CRUD with Kubernetes-style Terminating phase for graceful
deletion. All resources scoped by workspace via ObjectMeta. Membership
RPCs for workspace access control. Persistence migration shifts name
uniqueness to (object_type, workspace, name). Provider profiles support
platform and workspace scoping. Service routing uses workspace-prefixed
DNS labels. Inference routes renamed and workspace-scoped with
DeleteInferenceRoute RPC. Python SDK with WorkspaceClient, two-method
list pattern (workspace-scoped and for_all_workspaces), and workspace
parameter on all methods. CLI workspace flags, TUI workspace cycling.
K8s driver filters unmanaged CRs and uses delete preconditions. Podman
driver uses immutable container IDs. Label serialization fixed across
all put_if call sites.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(cli): delegate sandbox upload command to existing upload function

The standalone `sandbox upload` command reimplemented upload logic
inline with two bugs: it used `Path::exists()` which follows symlinks
(rejecting dangling symlinks), and it ran git-aware filtering on
symlink sources. The `run::sandbox_upload()` function already handles
both cases correctly via `sandbox_upload_plan()`. Replace the inline
logic with a call to the existing function.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(e2e): shorten sandbox names and fix test compatibility

Shorten the sandbox name in initial_sparse_policy_is_acknowledged_as_loaded
from 'e2e-2159-sparse-enrich' (22 chars) to 'e2e-sparse-enrich' (17 chars)
to comply with MAX_ROUTABLE_NAME_LEN (19 chars).

Also capture stderr in create_keep_with_args so future sandbox creation
failures include the actual CLI error instead of reporting empty output.

Signed-off-by: Derek Carr <decarr@redhat.com>

* test(workspace): add test coverage for workspace CRUD and persistence isolation

Add unit tests for workspace create happy path, get round-trip, get
not-found, get empty-name rejection, already-exists error, and
resolve_workspace not-found. Add persistence test proving cross-workspace
name uniqueness (same name in different workspaces produces separate
records). Add workspace name max-length boundary tests. Fix e2e harness
to include stderr in name-parse-failure error path. Align Python e2e
test_workspace_crud with try/finally pattern. Document provider profile
catalog workspace scoping gap in RFC 0011.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(examples): update examples for workspace model compatibility

Shorten sandbox names in demo scripts to fit the 19-character
MAX_ROUTABLE_NAME_LEN limit: policy-demo prefix to pd-, multi-agent
notepad derives a short SANDBOX_TAG from the run ID, governance
interceptor uses gs-PID-RANDOM. Update vscode-remote-sandbox.md SSH
host aliases from openshell-{name} to openshell-{name}.{workspace}
format.

Signed-off-by: Derek Carr <decarr@redhat.com>

* feat(sdk): add workspace-scoped client and workspace CRUD

Add WorkspaceScopedClient modeled after kube::Api::namespaced — captures
workspace once and injects it into every sandbox request. Add workspace
CRUD methods (create, get, list, delete) and list_sandboxes_all_workspaces
on OpenShellClient. Extend SandboxRef with workspace field and add
WorkspaceRef type. Include mock tests for all new operations.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(lint): resolve clippy warnings in workspace test assertions

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(docs): convert indented code blocks to fenced in RFC 0011

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(lint): resolve clippy warnings and apply cargo fmt across workspace

Auto-format with cargo fmt and fix clippy warnings exposed by the
reformat: unnecessary qualifications, map_unwrap_or, identical match
arms, unused variable prefix, dead code annotations, and let-unit-value
in e2e harness.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(workspace): address workspace scoping issues from review

- Add workspace field to settings JSON output (CLI)
- Skip Podman containers missing workspace label instead of defaulting
  to empty string, matching K8s driver behavior
- Add resource_version to list_by_scope SELECT in both SQLite and
  Postgres backends, with regression test
- Gate PolicyLocalContext proposal/lookup routes on workspace readiness,
  returning 503 when workspace is not yet discovered
- Block sandbox and provider creation in TUI all-workspaces mode
- Clear workspace vectors in TUI reset_sandbox_state

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(workspace): make provider profile catalog workspace-aware

Thread workspace through snapshot_catalog so the
EffectiveProviderProfileCatalog enforces workspace boundaries on both
read and write paths. UserProviderProfileSource now loads platform-scoped
profiles (workspace "") plus the target workspace's profiles, preventing
cross-workspace duplicate profile ID collisions that previously caused
global catalog failures.

Update RFC 0011 to reflect catalog scoping is implemented in Phase 1
rather than deferred to future work.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(persistence): include workspace column in atomic policy revision INSERT

put_policy_revision_atomic omitted the workspace column from the INSERT
into the objects table in both SQLite and Postgres backends, causing
atomically-written policy revisions to lose their workspace association.
Add workspace field to AtomicPolicyRevisionWrite and thread it through
both backend INSERT statements, matching the non-atomic put_policy_revision
path which already included it.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(proxy): skip ancestor walk when socket owner is the entrypoint

collect_ancestor_identities walked the entire process tree above the
entrypoint when the connecting process was the entrypoint itself,
SHA256-hashing every ancestor binary (IDE, shell, container runtime).
On dev machines with large binaries in the ancestor chain this exceeded
the 30-second test timeout. When start_pid == stop_pid there are no
intermediate ancestors to verify, so return an empty list immediately.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(workspace): make provider profile catalog scope-aware

Allow the same profile ID at platform and workspace scopes by
introducing layered catalog entries where workspace profiles shadow
platform profiles. Add source and scope fields to the ProviderProfile
proto and CLI output. Migrate List/Get handlers to the catalog,
fixing divergence with runtime profile resolution.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(e2e): align podman e2e labels with centralized driver constants

The podman driver moved its container labels to the centralized
openshell.ai/ prefix, but the e2e test harness and cleanup script
still referenced the old openshell.sandbox-* keys, causing the
local_driver_token_restart test to fail on container lookup.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(e2e): align python profile isolation test with scope-aware catalog

Platform profiles are now visible in workspace listings as fallbacks
per the layered catalog design. Update the assertion to match.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(workspace): honor profile_workspace in runtime profile resolution

Runtime profile lookups now consult provider.profile_workspace via
get_type_profile_for_scope. Providers created with --global-profile
(profile_workspace="") resolve to the platform profile even when a
workspace profile shadows the same ID. All 6 runtime call sites
updated; type-only call sites remain scope-agnostic.

Signed-off-by: Derek Carr <decarr@redhat.com>

---------

Signed-off-by: Derek Carr <decarr@redhat.com>
2026-07-21 00:59:18 +00:00
Evan Lezar 21aaa89522 feat(gateway): add elevated gateway info (#2202)
* feat(cli)!: fold gateway metadata into list

BREAKING CHANGE: openshell gateway info no longer shows local gateway registration metadata. Use openshell gateway list or openshell gateway list -o json for local registration details.

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* feat(gateway): add elevated gateway info

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-07-15 22:04:57 +00:00
John T. Myers 48a7d09e82 feat(providers): support profile updates (#1914)
* feat(providers): support profile updates

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(providers): require CAS for profile updates

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(providers): serialize profile update validation

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(providers): serialize profile import invariants

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(providers): require profile update target id

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-06-23 11:34:46 -07:00
Yuedong Wu ff028ce0df feat(server): support TLS certificate hot-reload (#1870)
* feat(server): support TLS certificate hot-reload

Signed-off-by: Yuedong Wu <dwcn22@outlook.com>

* refactor(server): extract shared TLS test utilities

Signed-off-by: Yuedong Wu <dwcn22@outlook.com>

---------

Signed-off-by: Yuedong Wu <dwcn22@outlook.com>
2026-06-16 16:16:13 -07:00
Eric Curtin d5b79e5ba1 refactor(server): deduplicate test helpers and grpc utilities (#1708)
Remove three groups of copy-pasted code in openshell-server:

1. grpc/mod.rs had a private current_time_ms() wrapper identical to the
   one already exported from persistence/mod.rs. Remove the duplicate
   and update the three grpc sub-modules (policy, sandbox, service) to
   import directly from crate::persistence.

2. test_store() was repeated verbatim in seven #[cfg(test)] blocks.
   Promote a single canonical version to persistence/mod.rs (cfg-gated)
   and replace all copies with crate::persistence::test_store() calls or
   a thin Arc wrapper in supervisor_session.

3. grpc_client_mtls() and build_tls_root() were copy-pasted across
   edge_tunnel_auth.rs and multiplex_tls_integration.rs. Move both into
   the existing tests/common/mod.rs shared module and import from there.
2026-06-03 10:14:55 -07:00
alangou 2bdc968ed9 fix(gateway): make readiness health checks dependency-aware (#1328)
* feat(gateway): add readiness probe metrics and test-only store close

Emit Prometheus readiness metrics for database probes (healthy gauge and
outcome-labeled latency histogram) with coverage in health HTTP tests.
Restrict Store::close behind test support cfg to prevent accidental runtime
pool shutdown under live traffic.

Signed-off-by: Adrien Langou <alangou@nvidia.com>

* test(e2e): add simple e2e test with kubernetes to test /readyz

Signed-off-by: Adrien Langou <alangou@nvidia.com>

---------

Signed-off-by: Adrien Langou <alangou@nvidia.com>
2026-05-27 11:00:06 -07:00
Taylor Mutch a3b16c18ab feat(auth): per-sandbox authentication to gateway (#1404) 2026-05-21 17:58:55 -07:00
Eric Curtin 5620c8b229 refactor: deduplicate repeated patterns across crates (#1499)
Remove ~280 lines of duplicated code across 30 files in 5 areas:

- centered_rect: consolidate 5 identical TUI layout helpers into a
  single pub fn in openshell-tui/src/ui/mod.rs
- server test helpers: replace ~100 inline Store::connect() calls
  with local test_store() helpers; deduplicate test_server_state()
  in grpc/service.rs to use the shared test_support version
- rogue PKI: extract 20-line rogue CA+client cert generation block
  (duplicated in two integration tests) into generate_rogue_pki()
  in tests/common/mod.rs
- provider tests: replace 8 identical 28-line test modules with a
  single macro_rules! test_discovers_env_credential! invocation
- label constants: centralize openshell.ai/ container label keys
  in openshell-core::driver_utils; update Docker and Kubernetes
  drivers to import from there instead of redefining them locally
2026-05-21 10:50:49 -07:00
Eric Curtin 10af3e609f refactor: deduplicate shared test helpers (#1399)
Extract four categories of copy-pasted test code into single
canonical locations:

- POLICY_OBJECT_TYPE / DRAFT_CHUNK_OBJECT_TYPE constants that were
  defined identically in both persistence/sqlite.rs and
  persistence/postgres.rs are now owned by persistence/mod.rs.

- test_server_state() was duplicated across grpc/provider.rs,
  grpc/policy.rs, and grpc/sandbox.rs. A shared grpc::test_support
  module in grpc/mod.rs now owns the single implementation; each
  submodule imports it.

- The TestOpenShell stub (full OpenShell trait impl), install_rustls_provider(),
  PkiBundle, generate_pki(), and start_test_server() were copied across
  up to five openshell-server integration test files. A new
  tests/common/mod.rs module owns them; each test file uses mod common.

- StubResponse, unique_socket_path(), and spawn_podman_stub() were
  duplicated between openshell-driver-podman/src/driver.rs and
  openshell-driver-podman/src/grpc.rs. A new src/test_utils.rs
  (cfg(test)-gated) owns the shared helpers.
2026-05-19 12:26:12 -07:00
John T. Myers d255cdd9c9 feat(providers): add credential refresh foundation (#1349)
* feat(providers): add credential refresh foundation

* feat(providers): mint oauth refresh token credentials

* fix(providers): allow delegated refresh bootstrap

* fix(providers): guard refresh credential modes

* fix(providers): refine refresh lifecycle UX

* fix(cli): restore provider test gateway response import

* fix(server): clean auth endpoint test qualifications

* fix(providers): tighten refresh authorization and collisions

* chore(providers): trim bundled v2 profiles

* fix(providers): resolve refresh rebase fallout

* fix(cli): accept rfc3339 credential expiry

* test(providers): update inferred claude provider type

* test(providers): avoid removed outlook default profile

* test(providers): isolate attach limit fixtures
2026-05-19 07:42:10 -07:00
Florent BENOIT 442b0b6b92 feat(exec): add bidirectional streaming for interactive TTY sessions (#1331)
* feat(exec): add bidirectional streaming for interactive TTY sessions

The existing ExecSandbox RPC sends stdin upfront in the request body,
making interactive programs (bash, top, vim) unusable with --tty. This
adds an ExecSandboxInteractive bidirectional streaming RPC that forwards
live keystrokes, terminal resize events, and EOF to the sandbox.

- Add ExecSandboxInteractive RPC, ExecSandboxInput and
  ExecSandboxWindowResize messages to the proto definition
- Implement server handler with relay bridging, PTY allocation via
  russh channel.split(), and timeout support
- Implement CLI client with raw mode, spawn_blocking stdin reader,
  SIGWINCH resize forwarding, and process::exit for clean shutdown
- Route to interactive path only when --tty is explicitly passed

Signed-off-by: Florent Benoit <fbenoit@redhat.com>

* fix(exec): address interactive TTY streaming review feedback

Server-side (sandbox.rs):
- Close SSH channel after EOF so PTY-backed programs (bash, vim, top)
  terminate on client disconnect instead of leaking the session
- Break output loop when tx.send() fails (client gone) to tear down
  the streaming pipeline promptly
- Default zero cols/rows to 80x24 before request_pty, matching the
  proto "0 = use default" contract and preventing 1x1 terminals for
  non-CLI clients

Client-side (run.rs):
- Replace spawn_blocking stdin reader with a detached std::thread so
  the tokio runtime shutdown does not hang on a thread blocked in
  stdin.read(). This removes the need for std::process::exit(),
  allowing save_last_sandbox() in the caller to execute normally
- Switch from futures::channel::mpsc with try_send() to
  tokio::sync::mpsc with blocking_send() for proper backpressure —
  try_send silently stopped forwarding all input when the channel
  filled during paste or slow network
- Add TaskGuard (abort-on-drop) for the resize task to ensure cleanup
  on all return paths including early errors

---------

Signed-off-by: Florent Benoit <fbenoit@redhat.com>
2026-05-15 12:12:44 -07:00
Seth Jennings c94cddbfb8 feat(server): separate HTTPS from mTLS authentication (#1351)
Make --tls-client-ca optional and make client certificates always
optional when a CA is configured. This decouples HTTPS encryption
from mTLS authentication, allowing mTLS and OIDC bearer tokens to
coexist as parallel authentication mechanisms.

When --tls-client-ca is provided, client certificates are validated
against the CA when presented but never required. Clients may connect
with or without a certificate — authentication is handled at the
application layer (e.g. OIDC).

Two TLS modes are now supported:
- HTTPS with optional mTLS (--tls-client-ca provided)
- HTTPS-only (--tls-client-ca omitted)

The --disable-gateway-auth flag is preserved for backward
compatibility but is now a no-op. The allow_unauthenticated field
has been removed from TlsConfig. The Helm chart conditionally
includes the client-ca volume and env var based on whether
clientCaSecretName is configured.
2026-05-15 09:43:30 -07:00
Piotr Mlocek 0797fefa44 feat(gateway): add local-domain service routing (#1101) 2026-05-12 17:44:55 -07:00
Piotr Mlocek 5abc36c461 feat(relay): route forwarding through ForwardTcp (#1029) 2026-05-11 21:23:37 -07:00
John T. Myers 1d3b741ee3 feat(providers): support sandbox provider attach lifecycle (#1242)
* feat(providers): support sandbox provider attach lifecycle

Closes #1171

Adds sandbox provider list, attach, and detach API/CLI support while keeping provider policy and credential resolution derived from current sandbox attachments.

* fix(providers): refresh sandbox provider credentials

Adds provider environment revisions and generation-scoped sandbox credential snapshots so future SSH and exec launches pick up provider attach, detach, and credential updates without mutating already-running processes.

Also blocks provider deletion while attached to prevent stale sandbox provider references.

* fix(providers): serialize sandbox object mutations

* test(providers): cover sandbox provider attach lifecycle

* test(providers): accept versioned credential placeholders
2026-05-08 11:14:44 -07:00
John T. Myers cdb1de59ba feat(providers): add custom profile registry (#1170)
Add custom profile registry. Allow attaching custom profiles at sandbox start.
2026-05-07 09:41:25 -07:00
John T. Myers 043bde279a feat(providers): add profile-backed policy composition (#1037)
Foundation for providers v2. Add provider profiles and provider profile composition with user policies.
2026-05-04 18:34:33 -07:00
Saurabh Agarwal 04e48d585c feat(server): add request-ID middleware for request correlation (#1082)
Add a UUID-based request-ID middleware using tower-http's request-id
feature. Each inbound request receives a unique x-request-id header
(or preserves a client-supplied one), which is recorded in the tracing
span and propagated to the response.

This enables operators to correlate log lines across the middleware
stack for a single request under concurrent load, and lets clients
reference specific requests in bug reports.

Signed-off-by: sauagarwa <sauagarw@redhat.com>
2026-05-04 16:14:58 -07:00
Drew Newberry 24724742a8 ci(rust): enforce -D warnings on clippy (#1008) 2026-04-29 12:12:53 -07:00
Piotr Mlocek a6d45528c1 feat(server,sandbox): supervisor-initiated SSH connect and exec over gRPC-multiplexed relay (#867) 2026-04-21 08:38:18 -07:00
John T. Myers bbcaed2ea7 refactor(proto): rename UpdateSettings to UpdateConfig for consistency with read path (#515) 2026-03-20 16:45:59 -07:00
John T. Myers a831a8921b feat(settings): gateway-to-sandbox runtime settings channel (#474)
* feat(gateway/sandbox): add global and sandbox runtime settings flow
2026-03-20 14:08:57 -07:00
Drew Newberry fbd93a4632 refactor: rename navigator- crate prefix to openshell- (#277) 2026-03-13 02:02:18 -07:00