Commit Graph
174 Commits
Author SHA1 Message Date
Drew Newberry 4c5fce6e59 refactor(compute): support external driver parity (#2744)
* refactor(compute): negotiate external driver behavior

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(compute): revert external driver documentation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(compute): negotiate gateway-managed lifecycle

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(compute): remove driver feature negotiation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(compute): let drivers declare gateway lifecycle

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(compute): clarify lifecycle ownership

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-08-20 15:56:29 +00:00
Drew Newberry 7909fb5d0f refactor(compute): unify gateway restart reconciliation (#2743)
* refactor(compute): unify gateway restart reconciliation

Remove the Docker-specific gateway shutdown cleanup and reconcile persisted running intent through ComputeDriver::StartSandbox for Docker, Podman, and VM drivers. Explicitly stopped sandboxes remain stopped.

Refs #2417

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(compute): stop local sandboxes on shutdown

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(e2e): match managed Podman containers

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(compute): synchronize lifecycle sweeps

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-08-20 10:12:41 +00:00
Mrunal Patel 6e90f3d5a0 feat(providers): store refresh credentials in credential drivers (#2801)
* feat(providers): store refresh credentials in credential drivers

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(providers): harden refresh credential lifecycle

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(providers): migrate legacy refresh secrets before skip

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* refactor(providers): defer credential migration

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(providers): make refresh configuration atomic

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

---------

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
2026-08-19 21:37:24 +00:00
Drew Newberry 998db04780 feat(policy): allow non-root sandbox identities (#2785)
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-08-19 21:20:14 +00:00
Jesse Jaggars 2eb0880a00 feat(cli): support OIDC device authorization grant for headless login (#2795)
* feat(cli): support OIDC device authorization grant for headless login

Closes #2793

Add OAuth 2.0 Device Authorization Grant (RFC 8628) support to the OpenShell CLI's OIDC login flow. When running in a headless environment (OPENSHELL_NO_BROWSER=1) without a client secret configured, the CLI now uses the device code flow instead of the browser-based PKCE flow.

The device code flow:
- Requests a device code and user code from the IdP's device authorization endpoint
- Displays a verification URL and user code to the user
- Polls the token endpoint until the user completes authorization or the code expires
- Supports slow_down responses per RFC 8628 by increasing the polling interval

This implementation:
- Extends OidcDiscovery to optionally capture device_authorization_endpoint
- Adds oidc_device_code_flow function with proper error handling for all RFC 8628 error codes
- Updates gateway add and gateway login to dispatch to device flow when browser is suppressed
- Adds comprehensive unit tests for device flow structs and response parsing
- Updates gateway authentication documentation to describe the device code fallback

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(cli): validate OIDC device token responses

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(cli): add PKCE to OIDC device flow

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* docs(cli): document PKCE device flow

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

---------

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>
2026-08-19 17:44:46 +00:00
grs 3a16012dbe fix(cli): prompt for fresh OIDC login after logout (#2773)
Signed-off-by: Gordon Sim <gsim@redhat.com>
2026-08-19 15:46:27 +00:00
alangou 0d708d6d51 fix(policy): gate uninspected credentialed endpoints (#2493)
* fix(policy): gate uninspected credentialed endpoints

Signed-off-by: Adrien Langou <alangou@nvidia.com>

* refactor(cli): extract allowed-ip option parsing

Signed-off-by: Adrien Langou <alangou@nvidia.com>

* fix(policy): gate endpointless credential bindings

Signed-off-by: Adrien Langou <alangou@nvidia.com>

---------

Signed-off-by: Adrien Langou <alangou@nvidia.com>
2026-08-19 15:18:02 +00:00
Giuseppe Scrivano d51a653f9c feat(driver-podman): add userns config (#2562)
* refactor(driver): extract shared supervisor binary helpers

Move supervisor binary extraction, caching, and validation helpers from
the Docker driver into openshell-core::driver_utils so both Docker and
Podman drivers can reuse them.

Moved helpers: extract_first_tar_entry, write_cache_binary_atomic,
supervisor_cache_path, temp_extract_container_name, and
validate_linux_elf_binary.

The shared extract_first_tar_entry gains entry-type and empty-payload
checks that the Docker-local version lacked.  supervisor_cache_path
takes a driver_subdir parameter so each driver caches under its own
namespace (docker-supervisor vs podman-supervisor).

Signed-off-by: Giuseppe Scrivano <gscrivan@redhat.com>

* feat(driver-podman): add userns config

Add a `userns` option to the Podman compute driver that maps to
Podman's user namespace modes.  The mode string is split on the first
colon into the API's `nsmode` and `value` fields so parameterized
values like `auto:size=65536` and `keep-id:uid=1000,gid=1000` are
forwarded correctly.  When the mode is `auto`, the container spec
also sets `idmappings.AutoUserNs = true` as required by the API.

An allowlist validates the mode at startup: `auto` and `keep-id`
accept optional parameters; `host`, `private`, and `nomap` reject
them; everything else is an error.

Podman image volumes use overlay mounts internally and the kernel
does not support idmapped mounts on overlay (`mount_setattr` returns
EINVAL).  When userns is configured (any mode except `host`), the
driver extracts the supervisor binary from the image to a host-side
cache and bind-mounts it instead of using an image volume.

Configurable via TOML `userns = "auto"`, CLI `--userns`, or
environment variable `OPENSHELL_PODMAN_USERNS`.

Signed-off-by: Giuseppe Scrivano <gscrivan@redhat.com>

---------

Signed-off-by: Giuseppe Scrivano <gscrivan@redhat.com>
2026-08-15 17:10:47 +00:00
Piotr Mlocek 44bf0df485 feat(middleware): inspect WebSocket text messages (#2477)
* feat(middleware): inspect websocket text messages

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): address websocket review feedback

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(network): bound websocket message assembly

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(network): harden websocket upgrade lifecycle

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* refactor(middleware): unify in-process and remote transports

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* feat(middleware): support regex websocket redaction

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): bound persistent streaming sessions

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): accept websocket sequence gaps

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* refactor(middleware): refine websocket introspection contract

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): clarify websocket preflight lifecycle

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): clarify websocket coverage semantics

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(network): type websocket frame failures

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(network): return 503 when middleware admission is exhausted

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): align streaming API contract

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): clarify WebSocket event result scope

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(rfc): simplify middleware revision history

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* feat(examples): add WebSocket content guard support

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): unify binding payload limits

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* refactor(middleware): align payload limit terminology

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): address websocket review feedback

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(network): address websocket review findings

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* test(network): allow Linux handler setup in preflight regression

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(network): harden websocket relay finalization

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(network): inspect compressed websocket messages

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* test(network): stabilize compressed websocket regressions

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(go-sdk): regenerate middleware protobuf binding

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): clarify websocket skip lifecycle

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): address WebSocket review feedback

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

---------

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-08-14 21:51:42 +00:00
Derek Carr 59479f492a feat(k8s): add namespace-per-workspace support (RFC 0011 Phase 3) (#2656)
* feat(k8s): add namespace-per-workspace support (RFC 0011 Phase 3)

Implement three workspace namespace modes for the Kubernetes compute
driver: shared (default, preserves current single-namespace behavior),
managed (auto-creates/deletes namespaces per workspace), and operator
(pre-provisioned namespaces with dynamic discovery via label selector
or drop-in allowlist file).

Key changes:
- WorkspaceMode enum and namespace resolution in driver config
- Managed namespace lifecycle with ServiceAccount and OpenShift SCC
  annotation propagation
- Cluster-wide sandbox CR watchers for managed/operator modes
- NamespaceValidator (Exact/Prefix/Allowlist) for SA token auth
- Workspace-aware credential secret storage
- Helm ClusterRole for multi-namespace RBAC
- Gateway config, architecture, and reference docs

Signed-off-by: Derek Carr <decarr@redhat.com>

* test(k8s): add e2e tests for workspace namespace modes

Add end-to-end tests for managed and operator workspace modes
introduced in RFC 0011 Phase 3. The managed mode tests verify
namespace creation with correct labels, ServiceAccount provisioning,
sandbox CR placement, and namespace survival with remaining sandboxes.
The operator mode tests verify rejection of unlabeled and nonexistent
namespaces. The positive operator path (sandbox in labeled namespace)
is known to fail due to an RBAC gap and will be addressed separately.

Also fixes Helm 4 compatibility: move SPDX license headers inside
conditional guards in 8 chart templates to prevent empty comment-only
documents, and fix a trailing whitespace trimmer in clusterrole.yaml
that concatenated the license header with apiVersion.

Adds cleanup sweep in with-kube-gateway.sh to remove managed and
operator namespaces before Helm uninstall, and mise tasks for running
each mode independently.

Signed-off-by: Derek Carr <decarr@redhat.com>

* feat(k8s): add operator namespace label watcher

Spawn a background kube::runtime::watcher in the K8s driver that
watches namespaces matching the configured label selector and populates
the OperatorNamespaceAllowlist at runtime. The driver owns the
allowlist and exposes its Arc so the server can share the same set with
the SA token authenticator.

create_sandbox now gates pod creation on the allowlist in operator
mode — workspaces whose namespace is not yet labeled are rejected at
resource render time rather than silently proceeding. Workspace
lifecycle itself is unaffected; only sandbox (resource) creation is
gated.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(k8s): harden operator mode and address review findings

Close the fail-open gap in operator mode when only
operator_namespace_file is configured: the allowlist is now created
unconditionally in operator mode (fail-closed from startup).

Implement the namespace file watcher using the notify crate, following
the TLS hot-reload pattern (parent-directory watch, 1s debounce,
ConfigMap symlink-swap safe). The file format is a JSON array of
namespace name strings.

Additional fixes from the 10-reviewer audit:
- Change allowlist rejection from InvalidArgument to FailedPrecondition
  so callers know the request may succeed later once the namespace is
  provisioned.
- NamespaceValidator::Allowlist now holds the OperatorNamespaceAllowlist
  newtype instead of a raw Arc<RwLock<BTreeSet>>, eliminating silent
  denial on RwLock poison.
- Verify LABEL_MANAGED_BY and LABEL_GATEWAY_ID ownership before
  deleting a managed namespace.
- Replace fixed 5s sleep in operator e2e test with a 30s poll loop.
- Add Helm validation for workspaceMode values.
- Fix Helm README type column and description for operator fields.
- Add insert/remove methods to OperatorNamespaceAllowlist; label
  watcher now uses them instead of reaching through shared().
- Reject configs with both operator_namespace_label and
  operator_namespace_file set.

Signed-off-by: Derek Carr <decarr@redhat.com>

* feat(k8s): add workspace-level compute driver RPCs and harden RBAC

Decouple namespace lifecycle from sandbox lifecycle by adding
EnsureWorkspace/DeleteWorkspace RPCs to the ComputeDriver service.
Namespace creation now happens before credential storage and namespace
deletion happens on workspace delete, fixing credential storage in
managed workspace mode.

- Add EnsureWorkspace and DeleteWorkspace proto RPCs with
  implementations across all compute drivers (K8s managed delegates to
  ensure_namespace/delete_namespace_if_empty; others no-op)
- Wire ensure_workspace into provider create/update/refresh paths so
  the namespace exists before the credential driver writes secrets
- Wire delete_workspace into workspace deletion for cleanup
- Remove delete_namespace_if_empty from sandbox deletion path
- Scope ClusterRole secrets access to non-shared workspace modes
- Add TODO for TLS cert hot-reload in sandbox gRPC client
- Harden e2e tests with control-plane sandbox resolution assertions
- Fix docker image save --platform flag for OCI index manifests

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(k8s): address re-review findings and add test coverage

- Use server-side apply for TLS secret sync (fixes second sandbox
  creation failure when TLS is enabled)
- Scope gateway-ID label selector unconditionally across all workspace
  modes (fixes operator reads/watches/deletes seeing foreign sandboxes)
- Validate operator allowlist in EnsureWorkspace and DeleteWorkspace
  RPCs (prevents credential writes to namespaces outside the allowlist)
- Extend ClusterRole secrets patch+delete to all non-shared modes with
  credential driver enabled (fixes operator credential storage RBAC)
- Validate namespace ownership on 409 conflict in ensure_namespace
  (prevents adopting unowned namespaces in managed mode)
- Replace delete_namespace_if_empty with unconditional delete_namespace
  letting Kubernetes cascade cleanup (fixes stuck terminating CRs)
- Strengthen NetworkPolicy TODO to cover both managed and operator modes
- Extract selector and ownership logic into testable free functions
- Add unit tests for gateway-ID selectors and namespace ownership
- Add Helm ClusterRole RBAC tests for operator credential driver

Signed-off-by: Derek Carr <decarr@redhat.com>

* ci(k8s): add workspace managed and operator mode e2e to CI

Wire the existing e2e:kubernetes:workspace-managed and
e2e:kubernetes:workspace-operator mise tasks into the branch-e2e
workflow so they run alongside the other core Kubernetes e2e suites.
Both are gated by run_core_e2e and included in the Core E2E result
gate.

Signed-off-by: Derek Carr <decarr@redhat.com>

* test(k8s): add e2e tests for workspace namespace modes

Add 7 new e2e tests covering workspace namespace lifecycle, TLS secret
copying, ownership conflict detection, DNS-1123 validation, operator
namespace preservation, and dynamic label watcher behavior. Fix async
sandbox deletion race condition in existing tests by polling sandbox
list instead of asserting immediately after delete.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(k8s): grant secrets/patch unconditionally and backfill gateway-id labels

Address two review findings:

1. RBAC: server-side apply (PATCH) is used for TLS secret sync in
   multi-namespace modes, but the ClusterRole only granted patch when
   the kubernetes-secrets credential driver was enabled. Grant patch
   unconditionally for non-shared modes since TLS sync always needs it;
   keep delete gated on the credential driver.

2. Upgrade safety: the new gateway-id label selector would orphan
   legacy Sandbox CRs that predate its introduction. Add a startup
   backfill in shared mode that patches any managed Sandbox CR missing
   the gateway-id label before the driver begins serving requests.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(k8s): address workspace namespace review findings

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(k8s): address follow-up review findings

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(k8s): preserve workspace lookup after rebase

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(helm): allow managed secret creation

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(k8s): stop pods in workspace namespace

Signed-off-by: Derek Carr <decarr@redhat.com>

* test(k8s): scope pod deletion check to v1alpha1

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(k8s): address workspace namespace review findings

Signed-off-by: Derek Carr <decarr@redhat.com>

---------

Signed-off-by: Derek Carr <decarr@redhat.com>
2026-08-14 21:22:41 +00:00
Piotr Mlocek bdabb54cb3 fix(security): authenticate extension services (#2638)
* fix(security): authenticate extension services

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* feat(extension-core): verify gateway JWTs

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* refactor(extension-core): keep inbound verification external

This should become an extension SDK package.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(security): harden the extension authentication contract

Follow-up hardening on the alpha extension authentication mechanism.

Claim contract:
- Extension tokens carry an explicit `typ` of `openshell-ext+jwt`. They
  share a signing key with sandbox-to-gateway admission tokens and were
  otherwise separated by audience alone, so a verifier that neglects to
  check `aud` could accept a gateway credential. The header is a second,
  independent discriminator.
- Publish OIDC-shaped discovery at `/.well-known/openid-configuration`
  so a service configured with only the gateway URL can learn the exact
  expected issuer and the JWKS location. It is shaped, not compliant:
  `issuer` is the gateway identity, not the serving URL.

Audience agreement:
- `MiddlewareManifest` and `InterceptorManifest` gain `expected_audience`.
  The audience is otherwise configured independently on each side of the
  boundary, where a mismatch surfaces only as an opaque authentication
  failure on every call. OpenShell now compares the two and fails at
  startup. An empty field keeps the check off for existing services.

Compatibility:
- Add `allow_insecure_transport` per registration. Enabling gateway JWT
  signing previously made any plaintext endpoint a hard startup failure,
  including the endpoint form used in our own documentation. The opt-out
  attaches no credential, is refused by the gateway if a supervisor asks
  for one, and warns at every startup.
- Make the transport requirement kind-aware. A middleware endpoint must
  be reachable from every sandbox supervisor, so only interceptors may
  use a gateway-local Unix socket.

Credential lifecycle:
- Replace the process-global slot map with a supervisor-owned
  `ExtensionCredentialStore` shared explicitly across the gateway
  connections the supervisor opens, removing test-order coupling.
- Rotate only when a credential is missing or has passed four fifths of
  its lifetime. Configuration polling ran every ten seconds against
  fifteen-minute credentials, so each poll re-ran gateway effective-policy
  resolution and re-minted the gateway token.
- Bound credential minting per sandbox, since each request resolves the
  caller's effective policy.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs: record alpha extension authentication in RFC appendices

Restore the RFC 0009 and 0010 bodies to their accepted text and move
every extension-authentication update into appendices instead. An RFC
records a decision at a point in time; superseding detail belongs
alongside it rather than rewritten into it.

RFC 0009's appendix carries the shared contract: claims, authorization,
key distribution, the `allow_insecure_transport` replacement for the
body's `allow_insecure`, and residual risks. RFC 0010's records only
what differs for interceptors and links to it. The existing
protocol-extensions appendix, which parked the phase 2 transport
question, now points forward to what was built.

Also document the audience handshake, the discovery endpoint, the
`typ` requirement, and `jti` replay guidance in the extensibility and
gateway configuration pages, and correct the middleware transport
guidance: middleware endpoints must be reachable from sandbox
supervisors, so Unix sockets are not an option there.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* refactor(extension-core): abstract extension server trust

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(core): update middleware manifest example

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(extension-auth): preserve unsigned gateway compatibility

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(extension-auth): reject cross-domain token replay

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

---------

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-08-14 20:34:10 +00:00
Drew Newberry f12f3ef8d5 fix(macos): restore Homebrew sandbox callbacks (#2739)
* fix(macos): restore Docker gateway callbacks

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(gateway): reuse reachable primary callback listener

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-08-14 18:38:20 +00:00
LR90 d0c6dc3fd8 feat(kubernetes): support corporate upstream proxy (#2633)
* feat(kubernetes): support corporate upstream proxy

Signed-off-by: loveRhythm1990 <qiuweimin@126.com>

* fix(kubernetes): reject proxy_auth_secret_key values Kubernetes cannot create

Gateway validation accepted proxy_auth_secret_key values that Kubernetes
rejects when creating the Secret (keys longer than 253 bytes, or the
reserved "."/".." names), turning an invalid deployment setting into
repeated sandbox Pod-provisioning failures instead of a startup error.
Reject them in validate_upstream_proxy_config so they fail closed at
gateway startup.

Signed-off-by: loveRhythm1990 <qiuweimin@126.com>

* docs(skill): add corporate upstream proxy checks to debug-openshell-cluster

Add a Kubernetes corporate upstream proxy troubleshooting section covering
rendered [openshell.drivers.kubernetes] configuration, credential Secret
volume events, supervisor arguments and mounts confined to the network-
supervising container, and proxy reachability.

Signed-off-by: loveRhythm1990 <qiuweimin@126.com>

---------

Signed-off-by: loveRhythm1990 <qiuweimin@126.com>
2026-08-14 16:16:43 +00:00
Jesse Jaggars c4b500a7de feat(helm): cert-manager external issuer + OpenShift passthrough Route (#2468)
* fix(core): trust public root CAs alongside the sandbox mTLS CA

The supervisor gRPC client only trusted the CA configured via
OPENSHELL_TLS_CA, since tonic ClientTlsConfig starts with an empty root
store unless with_native_roots()/with_webpki_roots() is also enabled.
Deployments where the gateway server certificate is issued by a public CA
(e.g. cert-manager against an ACME issuer) caused every supervisor
connection to fail the TLS handshake with "UnknownCA", since the sandbox
mTLS CA and the server cert issuer were no longer the same.

Enable both native and webpki roots in addition to the configured CA.
tonic root store is a union of all configured sources, so this does not
weaken verification for existing self-signed deployments. webpki-roots
(compiled in) is enabled alongside native-roots since the supervisor
binary may run in minimal sandbox images without a populated system CA
bundle.

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* feat(helm): support external cert-manager issuers and OpenShift Route passthrough

Add certManager.serverIssuerRef/clientIssuerRef so the gateway and mTLS
client certificates can be issued by a real Issuer/ClusterIssuer (e.g.
ACME) instead of only the chart built-in self-signed CA.

Add openshiftRoute template for exposing the gateway via a TLS
passthrough Route so the gateway keeps terminating its own TLS/mTLS.

The server Certificate excludes internal-only SANs (cluster-local,
localhost, loopback) when an external issuer is configured, since ACME
issuers reject those per CA/Browser Forum baseline requirements. A
template-time fail guard catches the misconfiguration at helm install
time rather than asynchronously at cert-manager issuance time.

Includes Helm unittest coverage for both issuerRef overrides and Route
rendering, plus a CI values overlay for lint coverage.

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* docs: document cert-manager external issuer and OpenShift Route

Update managing-certificates.mdx with the serverIssuerRef workflow and
install-time validation behavior. Add a production section to the
OpenShift guide covering passthrough Route with a real certificate.
Regenerate Helm README for new certManager and openshiftRoute values.
Sync debug-openshell-cluster skill with new troubleshooting steps for
ACME issuance failures and supervisor UnknownCA from mismatched CAs.

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(helm,core): address PR review feedback on cert-manager external issuer

Addresses all five blocking review items from #2468:

1. Remove .with_native_roots() from supervisor gRPC client -- the
   supervisor runs inside the user-selected sandbox image, so the
   image CA bundle is not operator-controlled. Keep .with_webpki_roots()
   (compiled-in, not user-controlled) alongside the configured CA.

2. Fail at render time when serverIssuerRef.name is set but
   clientCaFromServerTlsSecret is still true. Add negative Helm test.

3. Remove clientIssuerRef -- changing only clientIssuerRef breaks both
   directions because trust bundles are not modeled separately. Change
   serverIssuerRef.kind default from ClusterIssuer to Issuer.

4. Add server.oidc.issuer and server.oidc.audience to the documented
   OpenShift production Helm command. Add Access Control prerequisite.

5. Fail at render time when openshiftRoute.enabled and disableTls are
   both true. Add negative Helm test.

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(drivers): strip GATEWAY_TLS_SERVER_NAME from Docker and Podman env

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(helm,drivers): guard default clientCaSecretName and add env-strip tests

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* feat(tls): SNI-based dual certificate for internal and external server TLS

Split the gateway server certificate into two: an internal cert issued by
the chart's own CA (for supervisor connections via cluster-local SANs) and
an external cert issued by an operator-configured Issuer such as ACME/Let's
Encrypt (for CLI and Route access via public SANs).

The gateway uses SNI-based certificate selection: connections whose SNI
hostname matches external_server_names receive the external cert; all
others (including those with no SNI) receive the internal cert.

Security improvement: remove .with_webpki_roots() from the supervisor
gRPC client so supervisors trust only the chart CA, closing a MITM vector
via publicly-trusted certificates in user-supplied container images.

Key changes:
- Add DualCertResolver with SNI-based cert selection and full test coverage
- Add external_cert_path, external_key_path, external_server_names to TlsConfig
- Validate partial external cert config (error on cert-without-key or vice versa)
- Validate empty external_server_names when external cert is configured
- Split cert-manager templates into internal + external Certificate resources
- Add Helm guards for misconfigured external issuer (empty serverDnsNames,
  internal-only SANs with external issuer, conflicting clientCaFromServerTlsSecret)
- Update gateway-config.mdx, managing-certificates.mdx, openshift.mdx docs
- Update debug-openshell-cluster skill for dual-cert troubleshooting

Signed-off-by: Pi Agent <agent@openshell.local>

* fix(drivers): strip GATEWAY_TLS_SERVER_NAME in VM driver and correct comments

Add the same GATEWAY_TLS_SERVER_NAME environment stripping to the VM
compute driver that Docker, Podman, and Kubernetes drivers already
perform. Without this, a sandbox user on the VM driver could override
the TLS server name the supervisor verifies.

Fix stale comments in Docker and Podman drivers that referenced
'with WebPKI roots trusted' — WebPKI roots are explicitly not trusted
after the tls-webpki-roots removal.

Use tls-ring instead of bare channel for tonic in openshell-core so the
TLS API (ClientTlsConfig, Endpoint::tls_config) is available without
pulling in any root certificate store.

Signed-off-by: Pi Agent <agent@openshell.local>

* fix(tls,helm): wildcard SNI matching and Route host validation

Add RFC 6125 single-level wildcard matching to DualCertResolver so
external_server_names entries like *.example.com correctly match SNI
hostnames like gw.example.com. Previously only exact matches worked,
silently falling back to the internal cert for wildcard configurations.

Add a Helm fail guard in route.yaml that rejects openshiftRoute.host
values not listed in certManager.serverDnsNames when an external issuer
is configured — catches cert/route hostname mismatches at install time
instead of at TLS connect time.

Quote the host field in route.yaml for robustness.

Signed-off-by: Pi Agent <agent@openshell.local>

* fix(helm): address blocking review items — client-CA guard, wildcard Route, serverIssuerRef gate

1. Remove the obsolete guard rejecting serverIssuerRef + clientCaFromServerTlsSecret=true.
   The internal server certificate is always signed by the chart CA (the same
   CA that signs the client cert), so clientCaFromServerTlsSecret=true is
   correct — its filtered ca.crt is exactly the right trust anchor.  The old
   workaround (mounting openshell-ca-tls directly) unnecessarily exposed the
   CA private key to the gateway container.  Remove the client-CA overrides
   from docs, CI overlay, and production examples.

2. Route host validation now supports wildcard certificates per RFC 6125:
   single-level wildcards like *.example.com match gateway.example.com but
   not deep.sub.example.com.  Require an explicit openshiftRoute.host when
   an external issuer is configured — without one, OpenShift generates a
   hostname absent from serverDnsNames.

3. Reject serverIssuerRef.name when certManager.enabled is false — the
   external certificate, its Secret mount, and the gateway TLS config all
   require cert-manager to be enabled.

Validated on ROSA (dev.dyee.p3) with branch-built images:
- Fresh install with letsencrypt-prod ClusterIssuer
- SNI dual-cert: external hostname served Let's Encrypt cert
- Supervisor mTLS via internal cert path: ConnectSupervisor accepted
- Client CA volume: filtered ca.crt from internal server secret (no key)
- CLI connected via Route + OIDC

Helm tests: 81 pass across 7 suites.

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

---------

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>
Signed-off-by: Pi Agent <agent@openshell.local>
2026-08-14 00:12:11 +00:00
Seth Jennings 0f8fad23c4 feat(sandbox): add stop and start operations (#2653)
* feat(sandbox): add suspend and resume operations

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(server): preserve lifecycle work after cancellation

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(server): reconcile ambiguous lifecycle outcomes

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(server): complete suspended session cleanup

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(vm): preserve suspension state on resume failure

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(server): retry retained lifecycle transitions

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(server): clean sessions after suspend reconciliation

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* test(sandbox): cover deleting suspended sandbox

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(kubernetes): preserve progressing sandbox suspension

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(kubernetes): bound suspend status polling

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(kubernetes): detect legacy sandbox suspension

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(tui): render suspended sandbox phases

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* refactor(sandbox): rename suspend and resume lifecycle

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* perf(server): clean stopped sessions on transition

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(kubernetes): fail fast on rejected stop

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(compute): fence stale restart lifecycle events

Signed-off-by: Seth Jennings <sjenning@redhat.com>

---------

Signed-off-by: Seth Jennings <sjenning@redhat.com>
2026-08-13 06:13:54 +00:00
John T. MyersandJohn Myers 0120535efc feat(proxy): bind static credentials to provider endpoints (#2510)
* feat(proxy): bind static credentials to provider endpoints

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* test(e2e): verify static credential endpoint isolation

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* docs(provider): explain static credential endpoint binding

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* fix(e2e): use valid endpoint isolation fixtures

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* docs(provider): explain static credential endpoint binding

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* fix(credentials): preserve binding identity across rotations

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(proxy): enforce bindings across request lifecycle

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(proxy): close credential relay gaps

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(credentials): clarify binding failure behavior

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(credentials): hash selected provider profile scope

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(proxy): resolve credentials after request admission

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(credentials): clarify binding failure diagnostics

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(proxy): align single-route credential denials

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(credentials): harden endpoint-bound rotation

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(credentials): enforce identity and authority binding

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(credentials): snapshot provider environment atomically

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(e2e): include authority port in query proxy requests

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(credentials): close credential revocation gaps

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(proxy): explain authority mismatch diagnostics

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(credentials): enforce binding lifecycle invariants

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(provider): reject credential config collisions

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): capture credential scope atomically

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): distinguish origin and absolute targets

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(provider): isolate endpointless profile credentials

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): normalize IPv6 request authorities

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(credentials): clarify endpointless profile isolation

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* feat(policy): bind endpointless provider credentials

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(credentials): use current GCP placeholder revision

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(providers): explain policy credential bindings

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(credentials): cover endpointless fail-closed invariant

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(policy): expect ambiguity rejection at creation

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(server): authenticate rebased policy requests

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* refactor(proxy): share credential mismatch finding builder

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(credentials): cover malformed binding metadata

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(credentials): verify multi-key endpoint isolation

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(e2e): cover same-host credential path denial

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(credentials): document serialized refresh contract

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* refactor(proxy): consolidate L7 log formatting

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* perf(credentials): precompile endpoint binding patterns

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* perf(credentials): share identity epoch revisions

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(proxy): require explicit request default ports

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(policy): validate SigV4 credential sources

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(credentials): preserve endpoint bindings for credential handles

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* feat(go-sdk): expose network credential bindings

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <johntmyers@users.noreply.github.com>
2026-08-10 19:19:45 +00:00
Matthew Grossman 4cb77a900e fix(e2e): separate Podman Machine loopback listeners (#2622)
* fix(e2e): separate Podman Machine loopback listeners

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* test(e2e): remove shallow harness checks

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* refactor(e2e): trim Podman listener workaround

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* fix(e2e): bypass proxies for Podman health probe

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

---------

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>
2026-08-07 17:39:44 +00:00
5548405fcb feat(credentials): add provider credential storage drivers (#2437)
* feat(credentials): add provider credential storage drivers

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(credentials): harden credential update handling

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(credentials): harden credential driver security, correctness, and performance

Address review findings from the credential storage drivers PR:

- Route additional_credentials through the driver on refresh to prevent
  silent data loss for multi-credential providers (e.g. AWS STS)
- Clean up stored credential handles on CAS failure during refresh to
  prevent orphaned secrets in external backends
- Enforce namespace validation in the Kubernetes Secrets driver to
  prevent cross-namespace credential access when allow_reference_namespace
  is not enabled
- Cache Vault Kubernetes auth tokens with 80% TTL to avoid re-authenticating
  on every credential operation
- Parallelize resolve_credentials in all three drivers using try_join_all
  for faster sandbox startup
- Add existingSecret support for the KEK Secret to fix helm template/GitOps
  workflows where lookup returns empty and regenerates the key
- Document RBAC blast radius for the Kubernetes Secrets credential driver
  and recommend a dedicated namespace

Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com>
Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>

* fix(credentials): add optimistic concurrency, fix thundering herd, parallelize operations

Use resourceVersion optimistic concurrency with retry loop for K8s
Secret ownership checks to prevent TOCTOU races. Switch Vault token
cache from RwLock to Mutex with double-check pattern to prevent
thundering herd on cache miss. Parallelize credential store and delete
operations across independent keys using try_join_all.

Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com>
Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>

* fix(credentials): handle partial failures, add delete retry, consolidate cleanup

Replace try_join_all with join_all in credential store/delete operations
to handle partial failures — successfully-stored handles are cleaned up
when another key fails. Add retry loop with conflict detection to
db-credstore delete_credential, matching the K8s driver pattern.
Consolidate 4 manual cleanup_pre_stored_provider_credentials call sites
into a single error handler using an async block. Remove inconsistent
.trim() from db-credstore validate_handle_owner.

Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com>
Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>

* fix(credentials): fix retry loop guard and remove unprotected validation

Remove attempt-count guard from 409/Aborted match arms in retry loops
so the post-loop Status::aborted error is reachable after exhausting
retries. Previously, last-attempt conflicts fell through to the
catch-all error arm, producing misleading Status::unavailable errors.

Remove duplicate validation calls that ran after
prepare_provider_credential_update but outside the cleanup-protected
async block, which would leak pre-stored handles on failure.

Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com>
Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>

* fix(credentials): add workspace/provider UUID to credential backend paths

Include workspace and provider ID in credential backend object paths to ensure
cross-workspace uniqueness and prevent credential collision (GATOR-1806c9be-01).

- Updated credential driver proto to include workspace and provider_id fields
- Modified Vault driver to include workspace/provider_id in managed_secret_path
- Modified Kubernetes Secrets driver to include workspace/provider_id in
  credential_owner_id and managed_secret_name
- Updated all credential runtime calls to pass workspace/provider_id
- Updated tests to use the new signatures

This prevents two workspaces sharing the same external credential store from
colliding on provider names, which was a critical security issue (CWE-639).

* fix(credentials): preserve provider-level expiration for handle-backed credentials

Compute effective expiration from both provider and driver values using the
earliest non-zero timestamp and skip expired values before insertion
(GATOR-1806c9be-02).

- Modified resolve_provider_handles to check provider credential_expires_at_ms
- Skip expired credentials during resolution instead of returning them
- Use effective expiration (min of provider and driver) in resolution results
- Fix inference.rs to preserve earliest expiration when merging

This ensures handle-backed credentials respect the same expiration semantics
as inline credentials.

* fix(credentials): stage refresh changes under new handles before validation

Stage credential replacements under new immutable handles instead of reusing
existing handles to prevent overwriting committed values before validation/CAS
(GATOR-1806c9be-03).

- Stage credentials with empty existing_handles map to force new handle creation
- Validate and CAS before the new values are committed to backend storage
- Delete old handles only after successful CAS
- On CAS failure, delete only the newly staged handles
- This prevents CWE-362/CWE-367 race conditions where failed refreshes could
  still modify or delete the active credential

The fix ensures that a rejected refresh cannot modify the backend object still
referenced by the committed provider record.

* fix(credentials): add timeouts to credential driver RPCs

Apply configured timeouts to both startup capability negotiation and runtime
RPCs to prevent indefinite hangs (GATOR-1806c9be-05).

- Add DEFAULT_CREDENTIAL_DRIVER_RPC_TIMEOUT_SECS constant (30s)
- Apply timeout to GetCapabilities during startup connection
- Apply timeout to all runtime RPCs (store, delete, resolve)
- Use tokio::time::timeout to bound the entire GetCapabilities operation
  during startup, not just the socket connection
- Return contextual deadline errors on timeout

This prevents a faulty or overloaded driver from hanging gateway operations
indefinitely.

* fix(credentials): fix test to use consistent workspace/provider identity

The Kubernetes auth Vault resolve test was constructing a managed path with
test-workspace/test-provider-id but sending default/prov-123 in the request,
causing validation to reject the request (GATOR-18e32351-01).

- Update test to use test-workspace and test-provider-id in the request to
  match the logical_path construction
- This ensures the test exercises the intended code path and validates
  Kubernetes auth resolution properly

The test now passes and correctly validates identity enforcement.

* fix(credentials): use unique staging ID for refresh to avoid overwrites

Stage refresh replacements under genuinely distinct immutable handles using
a unique staging ID to prevent overwriting committed values (GATOR-1806c9be-03).

- Generate a unique staging ID using UUID for each refresh operation
- Use this staging ID when storing credentials instead of the real provider ID
- Pass the same staging ID during cleanup on failure to delete only staged objects
- This ensures deterministic paths (Vault) and object names (K8s) don't collide
  with the committed provider's credentials

The fix prevents failed refreshes from silently replacing active credentials
or breaking providers by deleting still-referenced backend objects.

* fix(credentials): wrap credential driver RPCs in local timeouts

Add local tokio::time::timeout wrappers around credential driver RPCs to
bound non-compliant or stalled UDS peers (GATOR-1806c9be-05).

- Wrap StoreCredential, DeleteCredential, and ResolveCredentials in local timeouts
- Return contextual deadline_exceeded errors when timeouts occur
- Keep existing gRPC timeout metadata for compliant implementations
- GetCapabilities during startup was already wrapped in previous commit

This ensures a faulty local driver cannot hang gateway operations indefinitely,
even if it accepts the connection but never responds to the RPC.

* fix(credentials): preserve ownership for staged refreshes

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(credentials): bound startup capability probe

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* test(provider): authenticate credential handler requests

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(ci): grant actions read to credential driver e2e

Signed-off-by: Seth Jennings <sjenning@redhat.com>

---------

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com>
Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>
Signed-off-by: Seth Jennings <sjenning@redhat.com>
Co-authored-by: Taylor Mutch <taylormutch@gmail.com>
Co-authored-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>
2026-08-05 16:29:41 +00:00
Matthew Grossman 537805568d feat(sandbox): honor OCI image working directories (#2530)
* feat(sandbox): honor Docker OCI working directories

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* fix(sandbox): honor effective workspace access

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* test(sandbox): cover enforced workspace denial

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* docs(docker): explain effective workdir checks

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* fix(sandbox): validate effective workspace writes

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* fix(sandbox): reserve supervisor control roots

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* refactor(sandbox): centralize control paths

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* fix(sandbox): reserve OCI runtime mount roots

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

---------

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>
2026-08-04 17:38:51 +00:00
krishicks 8328412959 feat(vm): export driver traces over OTLP (#2564)
Continue distributed traces across the gateway-to-driver process boundary and
export VM driver spans to the same OTLP/gRPC collector. The driver reports as
the distinct openshell-driver-vm service.

Updated the gateway architecture and configuration reference with a generic
external-driver forwarding contract.

Instrumented:

- Every RemoteComputeDriver RPC injects the active W3C trace context into
  tonic metadata. Managed VM readiness and runtime initialization give startup
  capability probes stable parent operations rather than isolated root spans.
- A tonic service layer creates fixed, low-cardinality server spans for every
  ComputeDriver RPC. New handlers inherit tracing automatically; failures record
  OpenTelemetry error status and the gRPC status code.
- Background provisioning remains attached to CreateSandbox after the RPC
  returns without extending the RPC span lifetime.
- Provisioning records image preparation, bootstrap image resolution, overlay
  preparation, lifecycle configuration, pre-launch hooks, guest preparation,
  and launcher spawn as child spans.
- VM startup reconciliation roots one trace for the persisted-sandbox scan,
  with per-sandbox restore and provision operations beneath it. The root remains
  open until all spawned restore tasks finish.
- Delete cleanup records its own child operation.

Design notes:

- The gateway forwards its configured OTLP endpoint to managed external
  drivers. SDK `OTEL_*` variables continue to own sampling, batching, limits,
  headers, and transport tuning.
- The VM driver has its own tracer provider and service resource so trace
  backends preserve the service boundary.
- RPC operation names come from an explicit method mapping, keeping cardinality
  bounded without parsing the protobuf descriptor set at runtime.
- Propagation uses a remote SpanContext for spawned provisioning. This keeps
  one trace while allowing the CreateSandbox server span to finish when the
  RPC response is sent.
- Startup restoration is independent of gateway requests. It begins at the VM
  driver reconciliation span rather than attaching to an unrelated RPC.
- Existing tracing events remain on the logging path. The OpenTelemetry layer
  exports spans only and excludes the SDK exporter callsites to avoid recursive
  traces.
- Export configuration failures do not prevent the driver from serving, and
  buffered spans are drained during graceful shutdown.
- Trace fields identify drivers, sandboxes, images, lifecycle phases, and gRPC
  outcomes without recording credentials, sandbox tokens, or request query
  parameters.

Refs #2507

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-08-03 20:52:22 +00:00
John T. MyersandJohn Myers 905b554c7c refactor(network): consolidate proxy egress pipeline (#2373)
* refactor(network): introduce shared egress pipeline

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(network): cover shared proxy egress paths

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* refactor(network): make destination authorization explicit

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* refactor(network): pin proxy relay policy context

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(network): lock relay generation contracts

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(network): establish phase zero compatibility baseline

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* feat(policy): detect ambiguous network endpoints

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* feat(sandbox): fail closed on invalid policy updates

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* refactor(network): invalidate relays on policy changes

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(policy): document validation failure posture

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(network): cover validation and middleware egress

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(config): move policy failure mode to gateway toml

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(network): name proxy contracts by behavior

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): align overlap validation with endpoint selection

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(network): expect hard loopback denial

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(network): match declared endpoint denial

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(policy): preserve path-specific endpoint overrides

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(network): respect hard-blocked host gateways

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): reconcile proxy refactor with main

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): preserve CONNECT policy generation

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(policy): cover runtime endpoint glob semantics

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(server): reject ambiguous policies before persistence

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(policy): explain ambiguity preflight behavior

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(policy): compare body limits within protocol

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(proxy): avoid global tracing capture race

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* chore(server): format rebased provider tests

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(server): authenticate rebased policy requests

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(sandbox): retain runtime on middleware outage

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(server): preflight provider composition activation

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(sandbox): distinguish runtime failure transitions

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
Signed-off-by: John Myers <johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <johntmyers@users.noreply.github.com>
2026-07-31 20:59:00 +00:00
Evan LezarandDrew Newberry d220d89468 feat(compute): negotiate gateway callback listeners (#2492)
* feat(compute): query gateway listener requirements

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* feat(compute): add Podman listener requirements

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(docker): use default gateway bind address

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(gateway): avoid wildcard primary listener

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(podman): validate callback listener discovery

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(server): support split dual-stack listeners

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(podman): support legacy rootless listener discovery

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(e2e): accept loopback plaintext rejection

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* docs(agent): add callback listener diagnostics

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(server): restrict compute callback listeners

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(podman): validate local callback port

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(server): clarify callback listener contract

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(podman): require pasta for local callbacks

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* docs(gateway): document RPM listener default

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* refactor(server): keep listener provenance diagnostic-only

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(compute): preserve callback listener isolation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(e2e): remove Podman callback relay

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(packaging): preserve Podman callback loopback

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci(e2e): run VM smoke on nested-virt runner

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci(e2e): gate VM smoke on usable KVM

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci(e2e): probe KVM through VM driver

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci(e2e): tolerate hosted KVM denial

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(server): close traced futures before assertions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* revert: remove tracing test stabilization

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Evan Lezar <elezar@nvidia.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
2026-07-31 16:41:06 +00:00
krishicks fa24299092 feat(gateway): export traces over OTLP (#2534)
Add an opt-in OTLP/gRPC trace exporter to the gateway. Export is enabled
by the presence of an `[openshell.gateway.otlp]` table with an endpoint;
there is no separate toggle.

Instrumented:

- Inbound request server spans, named for the RPC (`$service/$method`) or
  `{method} {path}` for plain HTTP. They continue valid W3C `traceparent`
  context when present and start a new trace otherwise. gRPC spans also
  carry `rpc.system`, `rpc.service`, `rpc.method`, and trailer-derived
  `rpc.grpc.status_code`.
- Compute driver calls (create, delete, list, get, validate, watch) as
  client spans anchored on the `ComputeDriver` contract.
- Store reads and writes as children of the current request or loop span.
- Work with no inbound request: compute driver initialization, the sandbox
  reconcile sweep, provider credential refresh tick, and driver watch events.
  Each roots one operation trace so its child work does not arrive as anonymous
  single-span traces.

This is deliberately not exhaustive. Auth, policy evaluation, and
middleware remain uninstrumented, as do store lifecycle calls (`ping`,
`close`) that a readiness poll would turn into a span per tick. The aim is
a useful trace tree at a reviewable size; coverage can grow against real
traces.

Design notes:

- The TOML table owns whether and where to export. The SDK `OTEL_*`
  variables own how; sampling, batching, and limits are not mirrored into
  gateway config.
- The OpenTelemetry layer exports spans only. Existing `tracing` events
  remain on the stdout and sandbox-log paths and are not copied into trace
  payloads.
- Telemetry never blocks the gateway. A malformed endpoint logs an error
  and disables export rather than failing startup, and buffered spans are
  drained during graceful shutdown.
- Failed spans carry error status without a separate `error.type` attribute.
  Request spans use HTTP status and gRPC response trailers; driver spans use
  the returned gRPC status; autonomous loop spans record failed results
  explicitly. Store spans exempt `UniqueViolation` and `Conflict`, because
  those errors report expected contention such as a held lease or an
  optimistic-concurrency retry.
- The compute driver is reachable only through `TracedDriver::call`, so a
  call cannot skip its span. This is the client half of a client/server pair
  and the single place to inject context if drivers move out of process.
- Tests share one process-wide subscriber and in-memory exporter because
  `tracing` caches callsite interest globally.

Inbound W3C trace context is propagated into gateway request spans. Context
is not yet injected into outbound driver calls, so a future out-of-process
driver would still need propagation at the `TracedDriver` seam.

The Helm chart is intentionally unchanged, so OTLP export cannot yet be
enabled on a chart-deployed gateway.

Refs #2507

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-07-30 20:29:42 +00:00
Derek Carr 9c019a93f5 Wire authorization into workspace model (#2445)
* feat(auth): implement RFC 0011 Phase 2 workspace authorization

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(auth): address PR review feedback on workspace authorization

- Docker e2e: add --health-port and switch readiness probe from
  `openshell status` to `curl /healthz`, fixing a false-positive
  readiness check in OIDC mode where the CLI exited 0 without
  actually contacting the gateway

- ListWorkspaces: move membership filtering from post-query N+1
  lookups into a SQL EXISTS subquery so pagination applies to the
  visible set, not the global ordering. Add generic
  list_with_membership to the persistence layer.

- Descriptor validator: reject role/scope fields on unauthenticated
  and sandbox auth modes, and allow-list workspace_role as
  user/admin and global_role as platform_admin to catch typos at
  startup

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(server): use authed request in delete telemetry test

The workspace authorization added by the Phase 2 auth changes requires
a Principal on every delete request. The delete-telemetry test was still
using a bare Request::new, so extract_principal failed before the handler
could acquire the delete gate, causing a 5-second timeout flake.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(auth): address gator review findings for workspace authorization

- Inject unauthenticated-local-dev principal in no-auth gateway mode so
  handlers that call extract_principal() always find one.
- Cap label-selector membership query at MAX_PAGE_SIZE instead of
  u32::MAX to bound the in-memory read.
- Authorize workspace membership before resolving workspace existence in
  all sandbox RPCs to prevent workspace-name enumeration by non-members.
- Remove dead_code allow on AuthorizedWorkspace.workspace now that
  callers use the normalized name from the authz result.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(auth): close workspace-name oracle and label-selector truncation

Swap authorize-before-resolve ordering in 27 handlers across
provider.rs, service.rs, policy.rs, and workspace.rs to prevent
CWE-203 workspace-name enumeration by non-members.

Add combined membership+label SQL query (list_with_membership_and_selector)
to both persistence backends so ListWorkspaces with label selectors no
longer silently drops results beyond the first page of membership matches.

Signed-off-by: Derek Carr <decarr@redhat.com>

* test(auth): add non-member rejection and membership+label persistence tests

Add comprehensive test coverage for workspace authorization changes:
- Non-member rejection tests across all 44 workspace-scoped handlers
  (sandbox, provider, service, policy, workspace, inference) verifying
  PERMISSION_DENIED is returned instead of NOT_FOUND to prevent
  CWE-203 workspace-name oracle
- Persistence test for list_with_membership_and_selector verifying
  SQL-level membership EXISTS + label filtering, multiple predicates,
  no-match cases, and pagination

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(auth): format merged import line in sandbox tests

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(auth): address gator re-review findings on workspace authorization

- Fix TUI unconditionally setting providers_v2_enabled after provider
  refresh; read the actual gateway setting via GetGatewayConfig at
  startup instead
- Fix SQLite json_extract with dotted label keys (e.g. example.com/env)
  by quoting the key in the JSON path
- Add authed_request wrappers to upstream OCI identity tests that were
  missing a principal after rebase
- Add test proving GetGatewayConfig is accessible without Platform Admin
- Add test for dotted/prefixed Kubernetes-style label key filtering

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(auth): address second gator re-review findings

- Loosen GetGatewayConfig from platform_admin to scope-only so workspace
  users can discover providers_v2_enabled during sandbox creation with
  inferred-provider commands; update proto descriptor, descriptor
  validation, and RFC 0011 access table
- Add validate_label_selector to handle_list_workspaces and escape
  single quotes in SQLite json_extract interpolation (CWE-89
  defense-in-depth)
- Re-fetch providers_v2_enabled after TUI gateway switch so the new
  gateway's capability is reflected
- Add e2e test for workspace user with inferred-provider command
- Add persistence test for adversarial label keys with SQL injection
  attempts
- Add handler test for invalid label selector rejection in
  ListWorkspaces

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(auth): address third gator review findings

- Cap label selector pairs at 64 (CWE-400) to bound SQLite dynamic SQL
- Add SCOPE_ONLY_METHODS allowlist for scope-without-role RPCs (CWE-863)
- Normalize ID-based data-plane handlers to return NOT_FOUND for
  unauthorized sandboxes, closing the cross-workspace oracle (CWE-203)
- Fix TUI provider profile cache lookup key mismatch for legacy
  providers with empty profile_workspace
- Add whoami to CLI skill reference command tree
- Update TUI skill doc with workspace, provider, and settings coverage
- Document scope/workspace orthogonality on GetGatewayConfig proto

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(auth): extend CWE-203 normalization to policy.rs sandbox handlers

GetSandboxConfig and GetSandboxLogs in policy.rs had the same
fetch-before-authorize pattern that leaked cross-workspace sandbox
existence. Promote fetch_and_authorize_sandbox to pub(super) and
use it from both sandbox.rs and policy.rs handlers.

Signed-off-by: Derek Carr <decarr@redhat.com>

* test(auth): update assertions for CWE-203 sandbox ID normalization

Cross-workspace sandbox access via ID-based handlers now returns
NOT_FOUND instead of PERMISSION_DENIED to prevent existence inference.
Update the unit test and OIDC e2e assertion to match.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(auth): narrow CWE-203 error mapping and correct whoami output formats

Only remap PERMISSION_DENIED to NOT_FOUND in fetch_and_authorize_sandbox
and RevokeSshSession, letting INTERNAL and UNAUTHENTICATED propagate
as-is. Fix whoami --output format values in cli-reference.md to match
the actual CLI (table/json/yaml, not text/json).

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(ci): share network namespace with Keycloak in containerized CI

In GitHub Actions job containers, Docker port publishing lands on the
host, not inside the job container. Detect this environment and attach
Keycloak to the job container's network namespace instead, with
hardened defaults (cap-drop ALL, no-new-privileges, loopback-only
listener).

Signed-off-by: Derek Carr <decarr@redhat.com>

---------

Signed-off-by: Derek Carr <decarr@redhat.com>
2026-07-30 00:31:32 +00:00
Matthew Grossman bc14018cad feat(sandbox): use policy-first OCI image identity (#2509)
* feat(sandbox): use policy-first OCI image identity

Closes #2331

Preserve per-field policy omission, derive Docker and Podman fallbacks from the inspected immutable image, and resolve the final numeric identity before starting agent children.

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* fix(sandbox): preserve declared process identities

Keep explicit policy values and OCI-declared names intact, defer passwd lookup until a primary GID is required, and refresh stale policy examples.

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* fix(supervisor): reuse resolved OCI identity

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* fix(supervisor): allow Linux pre-exec arguments

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* fix(kubernetes): protect resolved sandbox identity

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* fix(sandbox): prepare workspace for OCI identity

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* refactor(sandbox): own only workspace root

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* fix(sandbox): harden partial identity drops

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* test(sandbox): scope OCI image e2e to Docker

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* fix(sandbox): narrow OCI identity fallback scope

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* test(podman): cover OCI identity launch

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* fix(podman): exercise OCI fallback in E2E

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

---------

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>
2026-07-29 05:27:21 +00:00
LR90 7955c8309b feat(k8s): support configuring workspace PVC storageClassName (#2463)
The Kubernetes driver's default workspace PVC never set storageClassName,
so on clusters with no default StorageClass the PVC stayed Pending and
sandbox creation failed.

Add a workspace_storage_class option to KubernetesComputeConfig, wired
through SandboxPodParams into the generated volumeClaimTemplates. When
non-empty it sets storageClassName; empty preserves the current behavior
of relying on the cluster default StorageClass.

Expose it via the OPENSHELL_K8S_WORKSPACE_STORAGE_CLASS env var on both
the standalone driver and the embedded gateway runtime defaults, and via
the server.workspaceStorageClass Helm value.

Closes #2442

Signed-off-by: lr90 <qiuweimin@matrixorigin.cn>
2026-07-29 03:35:00 +00:00
Evan Lezar f00ad23a26 fix(podman): tolerate shutdown transport closes (#2498)
Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-07-28 08:00:37 +00:00
Grace Smith deced8716b refactor(policy): extract shared L7 endpoint validation (#2389)
Move L7 endpoint semantic checks into the openshell-policy crate so both
profile lint and the runtime validator share one implementation. This
eliminates drift between the two validation paths.

The shared validator covers 9 checks: unknown protocol, rules/access
mutual exclusivity, JSON-RPC family access rejection, json-rpc requires
rules, non-JSON-RPC protocol requires rules or access, MCP requires
rules when allow_all is false, rules-would-deny-all detection,
deny_rules require protocol, and deny_rules require base allow set.

Changes rules/deny_rules fields to Option<Vec<...>> so absent vs empty
is distinguishable at lint time. Adds is_effectively_empty() to
L7AllowProfile for deny-all detection of allow: {} objects. Makes
rules_would_deny_all MCP-aware by checking tool/params.name selectors
before classifying a rule as deny-all. Adds params field to
L7AllowProfile so MCP tool selectors survive proto round-trip.

Signed-off-by: Grace Smith <gsmith@redhat.com>
Signed-off-by: Grace Smith <grasmith@redhat.com>
2026-07-24 17:48:24 +00:00
Philippe Martin 77e5c32217 feat(sandbox,gateway): route sandbox egress through corporate HTTP proxy (#2245)
* feat(sandbox,gateway): route sandbox egress through corporate HTTP proxy

- Chain sandbox egress through a corporate HTTP proxy so outbound
  traffic from within the sandbox respects the host proxy settings
- Forward sandbox proxy environment variables to the generated Podman
  config so the proxy is applied consistently to Podman-managed workloads

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(sandbox,podman): make corporate proxy routing operator-owned

The operator-configured corporate egress proxy was injected under the
conventional HTTPS_PROXY/HTTP_PROXY/NO_PROXY names as defaults beneath
sandbox spec/template environment, so a sandbox creator could redirect
egress at an arbitrary proxy or disable proxying with NO_PROXY=*.

Route the boundary through reserved, supervisor-only variables
(OPENSHELL_UPSTREAM_HTTPS_PROXY/HTTP_PROXY/NO_PROXY) written in the
Podman driver's required-variable tier. Any sandbox-supplied value under
a reserved name is stripped before the operator value is applied, so the
supervisor never observes a reserved proxy variable the operator did not
set. The supervisor now reads only the reserved names and ignores the
conventional proxy variables the sandbox controls.

Add the reserved proxy variables to the supervisor-only child-environment
denylist so the corporate proxy URL and any embedded credentials are not
inherited by the sandbox workload, which reaches egress through the local
policy proxy and never needs them.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* feat(sandbox,podman): deliver corporate proxy credentials via secret file

Proxy credentials were embedded inline in the proxy URL, so they were
stored in gateway.toml and exposed in container metadata via
'podman inspect'.

Reject inline 'user:pass@' credentials in https_proxy/http_proxy at
startup (parsed with the url crate rather than hand-splitting), and add a
proxy_auth_file option pointing at a 'user:pass' file. The driver stages
that file as a per-sandbox root-only Podman secret, mounts it at a fixed
path, and exports only the path in the reserved
OPENSHELL_UPSTREAM_PROXY_AUTH_FILE variable, so the credential never
appears in config, environment, or container metadata. Reading the file
fails closed on a missing, empty, or control-character-bearing value.

The supervisor reads the credential from the mounted file and builds the
Proxy-Authorization: Basic header, rejecting control characters, and no
longer derives credentials from URL userinfo. The auth-file path is added
to the child-environment strip list, and the generated gateway.toml is
written owner-only (mode 0600).

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(sandbox,podman): fail closed on invalid upstream proxy configuration

The reserved OPENSHELL_UPSTREAM_* variables are an operator-owned egress
boundary, but the supervisor treated present-but-invalid values as unset:
an unsupported or malformed proxy URL was ignored with a warning, an
unreadable auth file proceeded without credentials, and a malformed
credential silently became unauthenticated. Any of these could quietly
downgrade the corporate proxy boundary to direct dialing or
unauthenticated proxy access.

Make every configured-but-invalid proxy or auth setting fatal to
supervisor proxy startup, emit an OCSF ConfigStateChange failure event
before refusing, and share URL validation semantics between the Podman
driver and the supervisor through a single validator in
openshell-core (parse_upstream_proxy_url), so a value accepted at
sandbox-create time can never be rejected in-container or vice versa.

Inline user:pass@ URL credentials are now fatal in the supervisor too
(previously warn-and-strip), matching the driver. Unset or empty
variables still mean no proxy; only present-but-invalid values fail.
Error paths never include credential content.

Addresses the fail-closed review item on #2245.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(sandbox,podman): finish the fail-closed upstream proxy credential contract

The upstream proxy URL already had a single shared validator, but the
credential did not: the Podman driver rejected only CR/LF/NUL while the
supervisor rejected every control character, so a credential accepted at
sandbox-create time (e.g. one containing a tab) could still be rejected
in-container. Present-but-whitespace reserved OPENSHELL_UPSTREAM_*
values were also silently treated as unset, quietly downgrading the
operator's egress boundary to direct dialing.

Add parse_upstream_proxy_credential to openshell-core as the single
source of truth for the documented user:pass credential form (non-empty
user, no control characters, trimmed) and use it in both the Podman
driver's secret staging and the supervisor's Proxy-Authorization header
construction. Error variants carry no payload so credential content can
never leak into messages.

Make a present-but-empty reserved variable fatal to supervisor proxy
startup instead of meaning "unset"; only fully unset variables disable
the proxy. The driver correspondingly rejects an empty no_proxy at
config time so it can never inject a value the supervisor refuses.

Addresses the remaining fail-closed credential/config review item on

Signed-off-by: Philippe Martin <phmartin@redhat.com>
#2245.

* fix(sandbox,podman): close remaining fail-open upstream proxy config paths

Two configuration paths could still silently run without the proxy
boundary the operator believed was in effect.

A no_proxy bypass list configured without any https_proxy/http_proxy
was accepted by both the driver and the supervisor and simply meant
"dial everything directly". Reject it on both sides, exactly like the
existing proxy_auth_file-without-proxy rule: an operator who wrote a
bypass list assumed proxying was active, so accepting it hides a
fail-open state.

The gateway.sh dev script guarded proxy settings with [[ -n "${VAR:-}"
]], which conflates unset with explicitly-empty and dropped the latter
before the gateway's validation could see it. Use ${VAR+x} instead so
a set-but-empty variable is written into gateway.toml and rejected at
startup by validate_proxy_config rather than silently discarded.

Addresses the remaining fail-open configuration review item on #2245.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(sandbox,podman): reject upstream proxy URLs with path, query, or fragment

parse_upstream_proxy_url accepted URLs like http://proxy.corp.com:8080/some/path
and silently discarded everything after host:port, a lenience inherited
from the original supervisor parser. A forward proxy is addressed by
host:port only, so extra components indicate a misconfiguration (for
example a pasted endpoint URL) and silently truncating them violates
the present-but-invalid-is-fatal contract enforced everywhere else in
this configuration surface.

Reject a path, query, or fragment in the shared validator with a new
UnexpectedComponent error. A bare trailing slash remains accepted
because the url crate normalizes an absent http path to "/", making the
two indistinguishable. Both the Podman driver (gateway startup) and the
supervisor (sandbox startup) inherit the rule through the shared
parser, keeping their semantics identical by construction.

Addresses the proxy URL component review item on #2245.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(sandbox,podman): remove plain-HTTP upstream proxy support

The http_proxy path tunneled plain-HTTP requests through the corporate
proxy with CONNECT to port 80 and then sent origin-form requests down
the tunnel. Conventional enterprise forward proxies expect plain HTTP
as absolute-form requests sent directly over the proxy connection, and
commonly refuse CONNECT to port 80, so the setting looked supported but
failed against typical deployments. Tunneling also blinds the proxy to
the one protocol it could inspect.

Narrow the feature to TLS (CONNECT) egress only, which is the
conventional and already-correct case: plain-HTTP requests now always
dial the destination directly, and only client CONNECT tunnels chain
through the corporate proxy. Remove the http_proxy config field, the
--sandbox-http-proxy / OPENSHELL_SANDBOX_HTTP_PROXY driver surface, the
reserved OPENSHELL_UPSTREAM_HTTP_PROXY variable, and the UpstreamScheme
plumbing. The feature never shipped, so this is a clean removal; a
stray http_proxy key in gateway.toml still fails loudly through the
config's deny_unknown_fields.

Removing the plain-HTTP proxy branch also removes its host-gateway
special case; the architecture doc now documents the real host-gateway
behavior (add driver-injected host aliases to the reserved NO_PROXY
list) instead of an invariant the HTTPS path never implemented.

Plain-HTTP forwarding through a corporate proxy can return later as
absolute-form forwarding behind its own design review.

Addresses the plain-HTTP forwarding review item on #2245.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(sandbox,podman): escape generated TOML and require explicit proxy URL form

Escape backslashes, quotes, and control characters when gateway.sh writes
proxy values into gateway.toml, so a hostile or unusual environment value
cannot corrupt the config or inject extra keys.

Restrict the upstream proxy URL grammar to the documented http://host:port
form: a scheme-less value is no longer normalized to http:// and a missing
port is no longer silently defaulted to 80. Docs, README, and CLI help now
state the explicit-form requirement consistently.

Also fix a test-only call of handle_tcp_connection that was missing the
upstream_proxy argument added in an earlier commit.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(sandbox,podman): preserve tunneled bytes read with the CONNECT response

The CONNECT handshake reads the corporate proxy's response in chunks, so
the read that completes the header block can also contain the first
tunneled payload bytes. Those bytes were discarded, silently corrupting
the start of the tunnel for server-speaks-first destinations or proxies
that coalesce writes.

connect_via now returns a PrefixedStream that replays any bytes received
past the response terminator before reading from the socket again; writes
pass through unchanged. Direct dials wrap the stream with an empty prefix
so downstream relay and TLS paths keep a single stream type, and
tls_connect_upstream is generalized to any AsyncRead + AsyncWrite stream.

Adds regression coverage for a combined response/payload read and for
prefix replay ordering.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(sandbox,podman): gate cleartext proxy Basic auth behind an explicit opt-in

Proxy-Authorization: Basic is base64 over the plain-TCP connection to the
http:// corporate proxy, so anyone on the network path between the sandbox
host and the proxy can recover the credential. Sending it is now an
explicit operator decision instead of an implicit side effect of
configuring proxy_auth_file.

Add a proxy_auth_allow_insecure driver setting (CLI
--sandbox-proxy-auth-allow-insecure, env
OPENSHELL_SANDBOX_PROXY_AUTH_ALLOW_INSECURE), delivered to the supervisor
as the reserved OPENSHELL_UPSTREAM_PROXY_AUTH_ALLOW_INSECURE variable.
Fail-closed pairing on both sides: an auth file without the
acknowledgement is rejected at gateway startup and at supervisor startup,
as is the acknowledgement without an auth file or any value other than
'true'. gateway.sh writes the key only as a TOML boolean; a non-boolean
value is emitted as a quoted string so the gateway rejects it at startup
instead of risking injection.

Documents the exposure prominently in the gateway config reference, the
driver README, and the sandbox architecture doc.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(sandbox,podman): honor port qualifiers and resolved addresses in NO_PROXY

Two divergences from the documented NO_PROXY contract:

- A port-qualified entry (internal.corp:8443) was silently stripped to its
  hostname and bypassed the proxy for every port, excluding traffic the
  operator never listed. Entries now keep the optional :port qualifier
  (also on IP and CIDR entries) and only apply to that destination port; a
  trailing qualifier that is not a valid port stays part of the pattern
  instead of widening the entry.

- IP and CIDR entries only matched IP-literal hosts, so a bypass like
  10.0.0.0/8 never applied to hostnames resolving into that range. NO_PROXY
  evaluation now sees the validated resolved addresses: an IP/CIDR entry
  matching through resolution authorizes a direct dial of only the
  addresses it contains, so a bypass scoped to an internal range cannot
  widen into a direct dial of addresses outside it. Hostname-level matches
  (loopback, wildcard, domain entries, IP-literal hosts) keep authorizing
  all validated addresses.

proxy_for is replaced by decision(host, port, resolved) returning either
the proxy endpoint or the permitted direct-dial subset, and dial_upstream
restricts the direct connect to that subset.

Adds regression coverage for port-scoped bypasses on domain, IP, and CIDR
entries, invalid port qualifiers, resolved-address matching, and
split-resolution subset dialing.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(sandbox,podman): bind proxied CONNECT tunnels to validated addresses

The CONNECT request sent to the corporate proxy carried the destination
hostname, so the proxy resolved the name itself and the addresses that had
passed SSRF and allowed_ips validation were discarded. Split-horizon DNS
or rebinding at the proxy could then reach internal or otherwise
unapproved destinations through a tunnel the supervisor logged as
validated, and IP-range policy could not be enforced at all on proxied
dials.

CONNECT now targets a validated resolved address by default: the proxy
performs no DNS resolution and the tunnel stays bound to the answer the
supervisor checked. The hostname still travels inside the tunnel (TLS SNI,
application Host), so destination servers behave normally. In
split-horizon networks, operators point the gateway host at the corporate
resolver so internal names validate to their internal addresses.

For proxies whose ACLs filter on hostnames and reject IP CONNECT targets,
a new proxy_connect_by_hostname opt-in (CLI
--sandbox-proxy-connect-by-hostname, env
OPENSHELL_SANDBOX_PROXY_CONNECT_BY_HOSTNAME, reserved
OPENSHELL_UPSTREAM_PROXY_CONNECT_BY_HOSTNAME) restores hostname CONNECT,
documented as re-opening proxy-side resolution and making the proxy's ACLs
the effective egress control. Fail-closed pairing on both sides: the
opt-in without a proxy, or any value other than 'true', is fatal.

Adds regression coverage for the IP CONNECT request line (including IPv6
bracketing and hostname non-leakage), the hostname opt-in, and the
config pairing rules in driver and supervisor.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(sandbox,podman): reject empty port after bracketed IPv6 proxy host

http://[fd00::1]: passed the explicit-port check because the bracketed
branch only tested for a colon after the bracket, then fell back to port
80 — violating the fail-closed http://host:port contract and potentially
sending configured Basic credentials to an unintended service. Require a
non-empty suffix after ]:, matching the unbracketed branch, and cover the
case in the shared parser tests.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(sandbox,podman): fall back across validated addresses in proxied CONNECT

The direct path hands TcpStream::connect the whole validated address list
and it falls back across them, but the validated-IP CONNECT path attempted
only the first address, so a dual-stack destination could fail through the
corporate proxy even when a later validated address was reachable.

connect_via_validated tries each validated address in order under one
aggregate CONNECT_HANDSHAKE_TIMEOUT budget, returning the first success;
when every attempt fails the error names the attempt count and carries the
last failure. An empty address list is rejected up front.

Adds regressions for first-fails/second-succeeds fallback (asserting both
CONNECT request lines), the aggregate all-addresses failure message, and
the empty-list rejection.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(sandbox,podman): strip new reserved proxy vars and complete docs/tests

Follow-ups from review:

- Add OPENSHELL_UPSTREAM_PROXY_AUTH_ALLOW_INSECURE and
  OPENSHELL_UPSTREAM_PROXY_CONNECT_BY_HOSTNAME to the supervisor-only
  strip list so workload child processes never inherit them, matching the
  documented contract for the other reserved proxy variables, and cover
  both in the supervisor-only variable test.
- Document all five OPENSHELL_SANDBOX_* proxy variables in the
  mise run gateway help text, marked Podman-only and stating the
  auth-file/acknowledgement pairing, and complete the gateway-key list in
  the Podman README.
- Add NO_PROXY composition coverage for bracketed IPv6 entries with port
  qualifiers, bare IPv6 entries, and IPv6 CIDR matching against an
  IPv6-literal host and against a hostname's resolved addresses,
  including the port-qualified CIDR form.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(sandbox,podman): cap each proxied CONNECT attempt within the shared budget

A proxy that accepted the first CONNECT request but never responded
consumed the entire aggregate handshake timeout, so later validated
addresses were never tried and the hang defeated the multi-address
fallback.

Each attempt is now time-boxed to its fair share of the time remaining
before the shared deadline (remaining / attempts_left): a hanging attempt
is cut off with enough budget left for every remaining address, while
time a fast failure does not use rolls over to later attempts and the
total never exceeds CONNECT_HANDSHAKE_TIMEOUT. A timed-out attempt is
recorded like any other failure, and the aggregate error distinguishes
all-attempted from budget-exhausted runs.

Adds a first-hangs/second-succeeds regression driven through a
test-visible budget parameter so it runs in about a second instead of a
real 30s window.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(sandbox,podman): deliver corporate proxy config on the supervisor argv

The proxy settings were injected as reserved OPENSHELL_UPSTREAM_* container
environment variables. The driver only wrote names the operator configured,
but container runtimes layer the spec environment over ENV values baked
into the sandbox image, so an image could supply NO_PROXY=*, enable
hostname CONNECT, or point an unconfigured deployment at an
attacker-controlled proxy whenever the operator left a field unset.

The settings now travel as supervisor command-line arguments
(--upstream-proxy, --upstream-no-proxy, --upstream-proxy-auth-file,
--upstream-proxy-auth-allow-insecure,
--upstream-proxy-connect-by-hostname) built by the driver from operator
config. The driver sets the container entrypoint and command explicitly,
so neither sandbox spec/template environment nor image ENV can influence
argv, and an omitted flag genuinely means unconfigured — in every
supervisor topology, since the supervisor no longer consults its
environment for these settings at all. Credentials stay on the root-only
secret mount; only the mount path appears on argv.

The reserved environment names, their strip-list entries, and the
env-based validation surface are removed. UpstreamProxyConfig::from_args
replaces from_env, reusing the same shared fail-closed validation and
pairing rules keyed by the CLI flag names.

Driver tests now assert the argv contract, including that sandbox-supplied
environment cannot add, remove, or redirect proxy flags; supervisor tests
cover from_args mapping and its pairing rules.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* docs(sandbox): align proxy comments with the argv transport

The argv migration left comments describing the configuration as reserved
environment variables ("reserved value", "present-but-empty variable",
"reserved upstream proxy variables"). Rephrase them as driver-supplied
arguments and operator settings so the documented trust boundary matches
the implementation. Comment-only change.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(sandbox): parse bracketed IPv6 authorities in client CONNECT targets

parse_target split the CONNECT authority at the first colon, so an
IPv6-literal target like [2001:db8::1]:443 always failed port parsing and
IPv6-literal clients could never reach policy evaluation; a regression
test even locked in that failure. Parse the RFC 3986 bracketed form and
return the host bracket-free, matching what DNS resolution, SSRF
validation, NO_PROXY matching, and the upstream CONNECT builder expect.
Unclosed brackets, a missing or empty port after the bracket, and
non-numeric ports are rejected; unbracketed behavior is unchanged.

Replaces the failure-locking test with success coverage for bracketed
targets and adds malformed-bracket rejection cases.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* test(podman): cover proxy-auth secret cleanup across lifecycle failures

The per-sandbox proxy-auth credential secret is staged before the
container is created and removed on cleanup, but no test proved the
cleanup paths actually issue the secret removal. Add Podman-stub tests
that drive create_sandbox to a container-create failure and to a
start failure, and delete_sandbox for an out-of-band deletion, asserting
each path issues the DELETE for the per-sandbox proxy-auth secret so a
credential can never outlive the sandbox that owned it.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* docs: list corporate proxy keys in the Podman compute-driver overview

The Fern Podman driver section enumerated its gateway.toml keys but
omitted the corporate egress proxy settings. Add https_proxy, no_proxy,
proxy_auth_file, proxy_auth_allow_insecure, and proxy_connect_by_hostname
with a pointer to the gateway configuration reference for the full
contract. No navigation change: the reference folder already includes the
gateway configuration page.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* test(sandbox): cover the SSRF-to-TLS composition across the proxy tunnel

Existing tests exercised validated-IP CONNECT and the upstream-TLS helper
independently, but not the full boundary. Add an end-to-end regression
that stands up a fake corporate proxy tunneling to a fake TLS server and
drives the real path: connect_via_validated CONNECTs to the validated
address, the proxy splices the tunnel, and tls_connect_upstream verifies
the upstream certificate against the original hostname carried in SNI.

It asserts the CONNECT authority is the validated IP and never the
hostname, that verification succeeds for the matching hostname, and that a
mismatched hostname is rejected — proving a rebinding or split-horizon
substitution behind the proxy cannot pass certificate verification.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(sandbox,podman): bound proxy-auth reads, reject port 0, fix stale comment

Three review findings:

- CWE-400: the proxy-auth credential file was read with an unbounded
  read_to_string on both the driver (sandbox-create) and supervisor
  (startup) paths, so a huge file or a special file such as /dev/zero
  could exhaust memory. Add a shared bounded reader in openshell-core that
  rejects non-regular files, caps the size at 4 KiB, and reads at most that
  many bytes; the driver runs it via spawn_blocking. Covers oversized,
  special-file, and missing-path cases on both sides.

- Reject an upstream proxy URL with port 0: it passed the explicit-port
  check and startup validation but is not a connectable TCP port, so every
  proxied dial would fail later. Add a typed ZeroPort error with
  shared-validator and Podman-config tests.

- Reword a driver-config comment that still described a 'reserved
  variable' to match the argv transport.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(sandbox): open proxy-auth file non-blocking to reject FIFOs promptly

read_upstream_proxy_credential_file opened the path with a blocking
File::open before the regular-file check, so a configured FIFO with no
writer would block open() indefinitely — hanging sandbox creation on the
driver and supervisor startup. Open with O_NONBLOCK on Unix so the open
returns immediately, then reject the non-regular file as before;
O_NONBLOCK has no effect on the later read of a regular file. Adds a
mkfifo regression asserting the reader returns promptly with a
non-regular-file error instead of hanging.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* test(podman): cover corporate proxy egress across driver and supervisor

The existing corporate-proxy tests construct config structs or call CONNECT
helpers directly, so none of them detect a break in the wiring between
layers: gateway TOML deserialization, the Podman argv and secret-mount
semantics, supervisor CLI parsing, or policy denial before proxy contact.

Add a Podman e2e that drives the whole chain against a fake authenticated
forward proxy and asserts that an approved TLS request traverses it with a
validated-IP CONNECT, a policy-denied destination is refused with 403
without ever reaching the proxy, credentials arrive through the mounted
per-sandbox secret, and deleting the sandbox removes that secret.

SupportContainer is a new harness fixture. Unlike ContainerHttpServer it
probes readiness with a TCP connect rather than an HTTP GET, so it can host
a forward proxy and TLS servers, and it exposes container logs and network
IP for assertions.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* test(podman): restart the gateway on the proxy-config panic path

The panic cleanup for the temporary corporate proxy configuration
restored the gateway TOML but left the gateway process running with the
temporary configuration still loaded, which could poison later test
binaries in the same run. Nothing restarted it: the only ManagedGateway
is the short-lived one inside restart_gateway, and its Drop only calls
start, which does not reload config for an already-running gateway.

Restore and synchronously stop/start the gateway in Drop, and set
restored only after the normal restore and restart both succeed so a
failed restart no longer suppresses the fallback.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(supervisor-network): derive the loopback proxy bypass from resolved IPs

The automatic bypass treated the host string "localhost" as proof that
the destination was loopback and returned every resolved address. A
sandbox controls its own /etc/hosts and resolve_socket_addrs consults it
before DNS, so a workload could map localhost to any policy-allowed
address and dial it directly, escaping the operator proxy and the
inspection and audit boundary it exists to provide.

Check the resolved addresses instead: the name bypasses only when the
resolution is non-empty and every address is loopback. A mixed answer is
not partially honored, and an IP literal is still authoritative for
itself. A spoofed localhost falls through to the entries below, so an
explicit operator NO_PROXY entry is still honored.

This matches the trust model detect_trusted_host_gateway already applies
to the same hosts file, which validates the mapped address rather than
the alias.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(podman): clean up proxy-auth secret when container already deleted

The delete_sandbox early-return path cleaned up the token secret but
skipped the proxy-auth secret, leaking it on disk. Also update tests
for recent API changes (Optional socket_path, workspace field,
list-based container lookup).

Signed-off-by: Philippe Martin <philippe@openshell.dev>
Signed-off-by: Philippe Martin <phmartin@redhat.com>

---------

Signed-off-by: Philippe Martin <phmartin@redhat.com>
Signed-off-by: Philippe Martin <philippe@openshell.dev>
2026-07-24 16:51:52 +00:00
Drew Newberry 59f7839f6b fix(auth): report gateway authentication status (#2435)
* fix(auth): report gateway authentication status

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(auth): reuse gateway info for status probe

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-07-23 20:01:28 +00:00
Mesut Oezdil 744a65d52e fix(driver-podman): avoid panic when HOME is unset on macOS (#2327)
* fix(driver-podman): resolve Podman socket via auto-detection

Align Podman socket selection with the existing Docker model: explicit
config wins, otherwise probe openshell-core for a responsive socket.
This also fixes the original HOME-unset panic, since resolution no
longer hardcodes a per-OS default path.

- add detect_podman_socket() in openshell-core, mirroring
  detect_docker_socket
- PodmanComputeConfig.socket_path is now Option<PathBuf>, no default
- remove default_socket_path() (podman driver) and podman_socket_path()
  (vm driver), both replaced by the shared detector
- update server env override and CLI for the new Option type
- add tests: responsive-candidate detection in openshell-core, and
  config-error (not panic) when no socket is configured or reachable

* fix(driver-podman): address review feedback on socket resolution

Extract socket resolution into resolve_socket_path, taking the
detector as a parameter so tests do not depend on real env vars or
the host's actual Podman state. Replace the flaky env-mutating test
with three deterministic cases: explicit wins, detected is used when
absent, and neither source errors.

Fix the socket_path doc comment to describe it from a config user's
point of view, matching DockerComputeConfig's docstring. Drop a
comment that only made sense next to the Docker driver code.

Update the Podman README and gateway docs: they described a fixed
per-OS default path that no longer exists, replace with the actual
probe-then-fail behavior.

* docs(driver-podman): simplify socket default description

Previous wording was self-contradictory (says auto-detect on unset,
then lists the same var as a probed candidate) and omitted the Linux
/run/user/uid/podman/podman.sock candidate.
2026-07-21 14:01:19 +00:00
Mesut Oezdil 80987e91c8 docs: fix broken links and small inconsistencies (#2329)
- README: fix github-sandbox tutorial link missing get-started segment
- README: replace dead community-sandboxes doc link with the actual repo
- README: match supported host list to support-matrix.mdx
- architecture/README: list the missing google-vertex-ai-provider doc
- SECURITY.md: fix a mis-indented list item
- standardize on NVIDIA/OpenShell-Community casing for repo links
2026-07-20 17:46:31 +00:00
Piotr Mlocek d556748771 feat(supervisor-middleware): add network egress middleware (#2027)
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-07-16 17:47:49 -07:00
krishicks cf4deccd37 fix(gateway): probe Docker socket during driver auto-detection (#2303)
Previously, Docker was auto-detected when the CLI was installed or a candidate
Unix socket existed. Neither check verified that the Docker API was responsive.
A similar check was done when auto-detecting Podman in the past, but was
replaced in 1f07bf04 with a probe of candidate Podman sockets instead.

This change applies the functional API probing approach introduced for Podman
in 1f07bf04 to Docker. It also makes Docker driver initialization use the same
socket-selection mechanism as Docker auto-detection instead of Bollard’s local
defaults. This means the previously auto-detectable Docker socket paths
$HOME/.docker/run/docker.sock and $XDG_RUNTIME_DIR/docker.sock will actually be
usable.

When no working compute driver can be auto-detected, the gateway exits early
with a message saying as much:

> configuration error: no compute driver configured and auto-detection found no
> suitable driver; set --drivers or OPENSHELL_DRIVERS to kubernetes, podman,
> docker, or vm

This makes for a better user experience when installing OpenShell without an
available supported compute driver.

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-07-16 10:02:24 -07:00
Drew Newberry 83003e80fc feat(interceptors): initial gateway interceptor implementation and reference example (#2005)
* feat(gateway): add descriptor-driven interceptors

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(gateway): add service-reflected interceptors

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* wip

* fix(gateway): harden interceptor evaluation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(interceptors): label metrics and harden governance smoke

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* remove on_error: ignore

* feat(gateway-interceptors): emit log annotations

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(examples): govern provider profiles in interceptor

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(gateway): preserve update config annotations

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(providers): support interceptor profile catalogs

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* wip

* feat(governance-interceptor): sign provider profiles

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(providers): use configured profile sources for refresh updates

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(gateway-interceptors): add phase-specific evaluation payloads

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(providers): compose provider profile sources

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(gateway-interceptors): preserve committed responses

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(gateway-interceptors): reject ambiguous protobuf oneofs

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(gateway-interceptors): validate patch candidates per binding

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(gateway-interceptors): use reflected protobuf codec

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(governance-example): canonicalize signed protobuf hashes

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(gateway): commit policy provenance atomically

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(gateway): close signed governance bypasses

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(gateway): isolate interceptor secrets and authority

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(gateway): snapshot provider profiles per request

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* chore(gateway): resolve server clippy warnings

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(server): satisfy provider source clippy lint

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(server): initialize policy test annotations

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(gateway-interceptors): require explicit route allowlist

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(proto): clarify update annotation semantics

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-07-14 20:58:36 -07:00
mjamiv 614c8c164d feat(kubernetes): support PVC subPath driver config (#2034)
* feat(kubernetes): support PVC subPath driver config

Signed-off-by: mjamiv <michael.commack@gmail.com>

* test(kubernetes): cover writable PVC driver config

Signed-off-by: mjamiv <michael.commack@gmail.com>

* fix(kubernetes): address PVC subPath review feedback

Signed-off-by: mjamiv <michael.commack@gmail.com>

* fix(kubernetes): address PVC config review follow-up

Signed-off-by: mjamiv <michael.commack@gmail.com>

* fix(kubernetes): address PVC review follow-ups

---------

Signed-off-by: mjamiv <michael.commack@gmail.com>
2026-07-10 20:17:37 +00:00
Taylor MutchandSeth Jennings 8eacb4779f feat(kubernetes): add sidecar supervisor topology (#2076)
* feat(kubernetes): add sidecar supervisor topology

Add the Kubernetes sidecar supervisor topology, its Helm/Skaffold configuration, topology documentation, and sidecar e2e matrix coverage. Skip root-only sandbox identity rewriting when process enforcement is network-only so the low-permission sidecar process container can start successfully.

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(supervisor): avoid similar process id names

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(supervisor): avoid similar process id names

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(sandbox): avoid similar proxy id names

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* docs(kubernetes): clarify sidecar topology limits

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(kubernetes): keep sidecar process leaf capless

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(kubernetes): refresh sidecar provider env snapshots

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* test(supervisor): align hot-swap identity regression

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(kubernetes): stage sidecar mtls files before proxy chown

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(kubernetes): simplify sidecar supervisor topology

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* chore(helm): reuse sidecar skaffold values

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(supervisor): avoid similar iptables helper names

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(e2e): harden kube gateway wrapper setup

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(supervisor): avoid nft batch rollback on OCP

Run nftables setup as individual commands so optional conntrack and log expressions can fail without rolling back required table, chain, and reject rules.

Signed-off-by: Seth Jennings <sjenning@redhat.com>
Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(kubernetes): preserve process identity in sidecar topology

Render sidecar pods with a shared process namespace, keep binary-aware network policy enabled, and move Kubernetes sidecar settings under the nested sidecar config table.

Also apply unprivileged Landlock/seccomp setup in NetworkOnly supervisor mode so sidecar topology keeps sandbox child hardening without privileged process setup.

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* refactor(kubernetes): replace sidecar snapshots with control socket

Coordinate sidecar policy and provider bootstrap over a local Unix socket so the process leaf no longer reads policy/provider snapshot files.

Report entrypoint startup through the control channel and keep gateway credentials confined to the network sidecar.

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* feat(kubernetes): support relaxed sidecar network identity

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(sandbox): satisfy sidecar clippy lint

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* refactor(kubernetes): standardize topology naming

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(sandbox): satisfy linux clippy timeout import

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(kubernetes): support kata sidecar on ipv4 pods

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(kubernetes): satisfy linux clippy for sidecar fallback

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* chore(kubernetes): remove stale supervisor topology references

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(kubernetes): enable sidecar binary policy inspection

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(kubernetes): harden sidecar control boundary

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(kubernetes): couple sidecar supervisor lifecycles

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

---------

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
Signed-off-by: Seth Jennings <sjenning@redhat.com>
Co-authored-by: Seth Jennings <sjenning@redhat.com>
2026-07-10 13:01:39 -07:00
Taylor Mutch bebf440b25 fix(helm): propagate supervisor image overrides (#2216)
Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
2026-07-10 12:31:11 -07:00
Florent BENOIT 10702133a3 fix(core): pin supervisor image tag to gateway version for all drivers (#2070)
* fix(core): pin supervisor image tag to gateway version for all drivers

The Podman and Kubernetes drivers defaulted the supervisor image to
`:latest` via DEFAULT_SUPERVISOR_IMAGE, while the Docker driver already
resolved a version-pinned tag. Extract the tag resolution logic into
openshell-core so all three drivers use the same
OPENSHELL_IMAGE_TAG > IMAGE_TAG > CARGO_PKG_VERSION priority chain.

Closes #2068

Signed-off-by: Florent Benoit <fbenoit@redhat.com>

* refactor(core): simplify supervisor image tag resolver to slice-based API

Remove the Docker driver's wrapper functions and call
openshell_core::config::default_supervisor_image() directly.
Simplify resolve_supervisor_image_tag to accept &[&str] instead
of three separate parameters.

Signed-off-by: Florent Benoit <fbenoit@nvidia.com>
Signed-off-by: Florent Benoit <fbenoit@redhat.com>

---------

Signed-off-by: Florent Benoit <fbenoit@redhat.com>
Signed-off-by: Florent Benoit <fbenoit@nvidia.com>
2026-07-10 11:26:49 -07:00
Mesut Oezdil ed8ce8208f docs: fix Docker version format from 28.04 to 28.0 (#2136) 2026-07-08 16:57:07 -07:00
Evan Lezar 6252aa17c8 rfc-0006: add driver config passthrough proposal (#1589)
* docs(rfc): add driver config passthrough proposal

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* docs(rfc): link driver config proposal PR

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* docs(rfc): clarify driver config scope

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* docs(rfc): clarify driver-local config schemas

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* docs(rfc): clarify driver config extension path

* docs(rfc): update driver config baseline

* docs(drivers): document bind-mount selinux_label and whitespace rules

---------

Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-07-06 09:56:34 +02:00
Seth Jennings 6461677c32 feat(policy): accept numeric UIDs for sandbox process identity (#1973)
* feat(policy): accept numeric UIDs in sandbox process identity validation

Allow run_as_user and run_as_group to be either the literal 'sandbox'
or a numeric UID/GID within [1000, 2_000_000_000]. This removes the
hard dependency on a baked-in 'sandbox' user in container images,
enabling compute drivers to inject resolved UIDs at sandbox creation.

Phase 1 of #1959.

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* feat(supervisor): accept numeric UIDs for process identity dropping

Allow run_as_user and run_as_group to be numeric UIDs/GIDs, removing
the hard dependency on a baked-in 'sandbox' user in container images.

Changes:
- validate_sandbox_user(): accepts numeric UIDs without passwd lookup
  (logs OCSF event); keeps passwd check for "sandbox" name; rejects
  non-numeric non-sandbox strings that fail passwd lookup
- prepare_filesystem(): passes numeric UIDs/GIDs directly to chown()
  instead of requiring a passwd entry
- drop_privileges(): resolves numeric UIDs/GIDs directly via UID::from_raw
  / Gid::from_raw; skips initgroups when target uid matches current euid;
  uses guard conditions before setgid/setuid calls
- session_user_and_home(): falls back to ("{uid}", "/sandbox") for
  numeric UIDs, avoiding a passwd lookup that will fail

Re-exports MIN_SANDBOX_UID and MAX_SANDBOX_UID from openshell-policy
so callers have consistent range constants.

Phase 2 of #1959.

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* feat(driver-kubernetes): resolve sandbox UID/GID from config or OpenShift SCC annotations

Phase 3 of the numeric-UID plan: allow operators to specify explicit
sandbox_uid/sandbox_gid in Kubernetes driver config, auto-detect from
OpenShift SCC namespace annotations, and propagate resolved values to
supervisor container env vars and PVC init container securityContext.

Changes:
- Add sandbox_uid/sandbox_gid fields to KubernetesComputeConfig
- Add SANDBOX_UID/SANDBOX_GID env var constants to openshell-core
- Implement resolve_sandbox_identity() to fetch namespace annotations
  and auto-detect OpenShift SCC UID ranges (sa.scc.uid-range)
- Pass resolved UID/GID through SandboxPodParams to pod spec builder
- Inject SANDBOX_UID/SANDBOX_GID env vars into supervisor container
- Update PVC init container securityContext with resolved UID/GID
  instead of hard-coded root
- Add comprehensive unit tests for resolution logic and annotation
  parsing (resolve_sandbox_uid, resolve_sandbox_gid, OpenShift SCC
  annotation parsing)

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* feat(driver-vm): add configurable sandbox UID/GID and update docs/examples

Phase 4 of the numeric-UID plan: replace hardcoded SANDBOX_UID (10001)
in VM rootfs preparation with configurable sandbox_uid/sandbox_gid fields.

Changes:
- Add sandbox_uid/sandbox_gid to VmDriverConfig with serde derives
- Pass resolved UID/GID through prepare_sandbox_rootfs_from_image_root
  to ensure_sandbox_guest_user which writes /etc/passwd/group/gshadow
- Update BYOC Dockerfile: remove groupadd/useradd, document runtime UID
  injection and the ability to skip baked-in sandbox user
- Update gateway-config.mdx: document sandbox_uid/sandbox_gid for both
  Kubernetes (with OpenShift SCC autodetection) and VM drivers
- Update sandbox-compute-drivers.mdx: add Sandbox User Identity section
  explaining numeric UID support across all compute drivers
- Update rootfs tests to use non-default UIDs, verify config passthrough

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* code review changes

* fix(supervisor): harden tests for restricted CI container environments

Guard tests against CI-specific constraints: root without CAP_SETPCAP,
UIDs with no /etc/passwd entry, and restricted /proc access.

Signed-off-by: Seth Jennings <sjennings@nvidia.com>
Signed-off-by: Seth Jennings <sjenning@redhat.com>

---------

Signed-off-by: Seth Jennings <sjenning@redhat.com>
Signed-off-by: Seth Jennings <sjennings@nvidia.com>
2026-07-02 13:31:06 -07:00
Florian Bergmann 43bb030266 feat(docker,podman): add SELinux label support for bind mounts (#2092)
* feat(docker,podman): add SELinux label support for bind mounts

The Docker Engine structured Mount API does not support SELinux
relabelling (:z / :Z). Move user-supplied bind mounts from the
structured `mounts` field to the legacy string-format `binds` field,
which does support these options.

Add a shared `SelinuxLabel` enum (shared/private) to openshell-core so
both Docker and Podman drivers accept an optional `selinux_label` field
on bind mount configs. For Docker, labels are appended to the bind
string; for Podman, they are pushed to the mount options vec.

Signed-off-by: Florian Bergmann <fbergman@redhat.com>

* fix(docker): reject missing bind source paths on legacy binds

Moving user bind mounts from the structured Mount API to the legacy
Binds field changed Docker's behavior for missing source directories:
the legacy path silently creates them as empty root-owned dirs instead
of erroring. Add an explicit Path::exists() check to preserve the
fail-fast behavior operators expect.

Signed-off-by: Florian Bergmann <fbergman@redhat.com>

---------

Signed-off-by: Florian Bergmann <fbergman@redhat.com>
2026-07-02 08:16:52 -07:00
Taylor Mutch 914da339b4 feat(kubernetes): add combined topology config surface (#2074)
* feat(kubernetes): add combined topology config surface

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* docs(kubernetes): clarify topology defaults

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

---------

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
2026-06-30 15:07:55 -07:00
Shiju 5477e2f21d docs(mcp): fix granular policy lifecycle examples (#2066)
Signed-off-by: Shiju <shiju@nvidia.com>
2026-06-30 13:36:37 -07:00
krishicksandddurst 7bce1223dc feat(policy): add JSON-RPC and MCP L7 policies (#1865)
Add policy schema, proto, provider profile, OPA, and L7 proxy support for
`protocol: json-rpc` and `protocol: mcp`. Generic JSON-RPC endpoints match
exact method names only, with `method: "*"` as the all-method sentinel;
wildcard/glob methods and params matchers are rejected.

Parse JSON-RPC request bodies and batches in the forward proxy, deny
response-shaped client frames, limit receive-stream GET allowance to MCP
endpoints, and redact params in decision logs. Preserve L7 rule params on the
proto load path so MCP `tools/call` tool filters behave like YAML-loaded
policies.

Add MCP conformance coverage, JSON-RPC L7 e2e coverage, and docs for the new
protocols and current matcher limitations.

Signed-off-by: Kris Hicks <khicks@nvidia.com>
Co-authored-by: ddurst <267424412+ddurst-nvidia@users.noreply.github.com>
2026-06-26 15:52:16 -07:00
Taylor Mutch ba21bb32a2 feat(kubernetes): support agent-sandbox v1beta1 (#2009)
* feat(kubernetes): support agent-sandbox v1beta1

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* ci(kubernetes): test agent-sandbox api versions

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(kubernetes): retry agent-sandbox raw 404s

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(kubernetes): harden agent-sandbox api setup

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(kubernetes): cache agent-sandbox api version

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* docs(kubernetes): document agent-sandbox upgrade behavior

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* refactor(kubernetes): collapse agent-sandbox api selection

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

---------

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
2026-06-26 11:44:52 -07:00
Jesse JaggarsandRussell Bryant f569a0ade6 feat(sandbox): proxy-side AWS SigV4 credential signing for CONNECT tunnels (#1638)
Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>
Co-authored-by: Russell Bryant <russell.bryant@gmail.com>
2026-06-26 10:52:24 -07:00
st-grandEvan Lezar 75a317ea45 feat(server): support out-of-tree compute drivers via --compute-driver-socket (#1703)
* feat(server): support remote compute driver endpoints

Add named remote compute driver endpoint support to the gateway.

Remote drivers are selected by a non-reserved compute driver name and either a CLI/env socket endpoint or [openshell.drivers.<name>].socket_path. The VM driver now enters ComputeRuntime through the same acquired remote endpoint path, while Docker, Podman, and Kubernetes retain their in-process drivers.

Require --drivers/OPENSHELL_DRIVERS when pairing an ad-hoc socket endpoint so the socket does not imply a magic driver name, and keep reserved in-tree names unavailable for unmanaged socket endpoints.

Co-authored-by: Evan Lezar <elezar@nvidia.com>

Signed-off-by: st-gr <38470677+st-gr@users.noreply.github.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(server): cover remote compute driver UDS lifecycle

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: st-gr <38470677+st-gr@users.noreply.github.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>
Co-authored-by: Evan Lezar <elezar@nvidia.com>
2026-06-26 12:32:20 +02:00
Evan Lezar d93293ad12 fix(e2e): stabilize local Docker smoke test (#1935)
* fix(docker): honor configured supervisor image

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(cli): isolate ssh from host linker environment

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-06-25 16:51:27 +02:00