mirror of
https://github.com/NVIDIA/OpenShell.git
synced 2026-10-04 00:23:53 +08:00
codex/mxc-https-test-socket-owner
174
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
4c5fce6e59 |
refactor(compute): support external driver parity (#2744)
* refactor(compute): negotiate external driver behavior Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(compute): revert external driver documentation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(compute): negotiate gateway-managed lifecycle Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(compute): remove driver feature negotiation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(compute): let drivers declare gateway lifecycle Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(compute): clarify lifecycle ownership Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
7909fb5d0f |
refactor(compute): unify gateway restart reconciliation (#2743)
* refactor(compute): unify gateway restart reconciliation Remove the Docker-specific gateway shutdown cleanup and reconcile persisted running intent through ComputeDriver::StartSandbox for Docker, Podman, and VM drivers. Explicitly stopped sandboxes remain stopped. Refs #2417 Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(compute): stop local sandboxes on shutdown Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(e2e): match managed Podman containers Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(compute): synchronize lifecycle sweeps Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
6e90f3d5a0 |
feat(providers): store refresh credentials in credential drivers (#2801)
* feat(providers): store refresh credentials in credential drivers Signed-off-by: Mrunal Patel <mrunalp@gmail.com> * fix(providers): harden refresh credential lifecycle Signed-off-by: Mrunal Patel <mrunalp@gmail.com> * fix(providers): migrate legacy refresh secrets before skip Signed-off-by: Mrunal Patel <mrunalp@gmail.com> * refactor(providers): defer credential migration Signed-off-by: Mrunal Patel <mrunalp@gmail.com> * fix(providers): make refresh configuration atomic Signed-off-by: Mrunal Patel <mrunalp@gmail.com> --------- Signed-off-by: Mrunal Patel <mrunalp@gmail.com> |
||
|
|
998db04780 |
feat(policy): allow non-root sandbox identities (#2785)
Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
2eb0880a00 |
feat(cli): support OIDC device authorization grant for headless login (#2795)
* feat(cli): support OIDC device authorization grant for headless login Closes #2793 Add OAuth 2.0 Device Authorization Grant (RFC 8628) support to the OpenShell CLI's OIDC login flow. When running in a headless environment (OPENSHELL_NO_BROWSER=1) without a client secret configured, the CLI now uses the device code flow instead of the browser-based PKCE flow. The device code flow: - Requests a device code and user code from the IdP's device authorization endpoint - Displays a verification URL and user code to the user - Polls the token endpoint until the user completes authorization or the code expires - Supports slow_down responses per RFC 8628 by increasing the polling interval This implementation: - Extends OidcDiscovery to optionally capture device_authorization_endpoint - Adds oidc_device_code_flow function with proper error handling for all RFC 8628 error codes - Updates gateway add and gateway login to dispatch to device flow when browser is suppressed - Adds comprehensive unit tests for device flow structs and response parsing - Updates gateway authentication documentation to describe the device code fallback Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(cli): validate OIDC device token responses Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(cli): add PKCE to OIDC device flow Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * docs(cli): document PKCE device flow Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> --------- Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> |
||
|
|
3a16012dbe |
fix(cli): prompt for fresh OIDC login after logout (#2773)
Signed-off-by: Gordon Sim <gsim@redhat.com> |
||
|
|
0d708d6d51 |
fix(policy): gate uninspected credentialed endpoints (#2493)
* fix(policy): gate uninspected credentialed endpoints Signed-off-by: Adrien Langou <alangou@nvidia.com> * refactor(cli): extract allowed-ip option parsing Signed-off-by: Adrien Langou <alangou@nvidia.com> * fix(policy): gate endpointless credential bindings Signed-off-by: Adrien Langou <alangou@nvidia.com> --------- Signed-off-by: Adrien Langou <alangou@nvidia.com> |
||
|
|
d51a653f9c |
feat(driver-podman): add userns config (#2562)
* refactor(driver): extract shared supervisor binary helpers Move supervisor binary extraction, caching, and validation helpers from the Docker driver into openshell-core::driver_utils so both Docker and Podman drivers can reuse them. Moved helpers: extract_first_tar_entry, write_cache_binary_atomic, supervisor_cache_path, temp_extract_container_name, and validate_linux_elf_binary. The shared extract_first_tar_entry gains entry-type and empty-payload checks that the Docker-local version lacked. supervisor_cache_path takes a driver_subdir parameter so each driver caches under its own namespace (docker-supervisor vs podman-supervisor). Signed-off-by: Giuseppe Scrivano <gscrivan@redhat.com> * feat(driver-podman): add userns config Add a `userns` option to the Podman compute driver that maps to Podman's user namespace modes. The mode string is split on the first colon into the API's `nsmode` and `value` fields so parameterized values like `auto:size=65536` and `keep-id:uid=1000,gid=1000` are forwarded correctly. When the mode is `auto`, the container spec also sets `idmappings.AutoUserNs = true` as required by the API. An allowlist validates the mode at startup: `auto` and `keep-id` accept optional parameters; `host`, `private`, and `nomap` reject them; everything else is an error. Podman image volumes use overlay mounts internally and the kernel does not support idmapped mounts on overlay (`mount_setattr` returns EINVAL). When userns is configured (any mode except `host`), the driver extracts the supervisor binary from the image to a host-side cache and bind-mounts it instead of using an image volume. Configurable via TOML `userns = "auto"`, CLI `--userns`, or environment variable `OPENSHELL_PODMAN_USERNS`. Signed-off-by: Giuseppe Scrivano <gscrivan@redhat.com> --------- Signed-off-by: Giuseppe Scrivano <gscrivan@redhat.com> |
||
|
|
44bf0df485 |
feat(middleware): inspect WebSocket text messages (#2477)
* feat(middleware): inspect websocket text messages Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(middleware): address websocket review feedback Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(network): bound websocket message assembly Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(network): harden websocket upgrade lifecycle Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * refactor(middleware): unify in-process and remote transports Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * feat(middleware): support regex websocket redaction Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(middleware): bound persistent streaming sessions Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(middleware): accept websocket sequence gaps Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * refactor(middleware): refine websocket introspection contract Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(middleware): clarify websocket preflight lifecycle Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(middleware): clarify websocket coverage semantics Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(network): type websocket frame failures Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(network): return 503 when middleware admission is exhausted Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(middleware): align streaming API contract Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(middleware): clarify WebSocket event result scope Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(rfc): simplify middleware revision history Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * feat(examples): add WebSocket content guard support Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(middleware): unify binding payload limits Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * refactor(middleware): align payload limit terminology Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(middleware): address websocket review feedback Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(network): address websocket review findings Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * test(network): allow Linux handler setup in preflight regression Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(network): harden websocket relay finalization Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(network): inspect compressed websocket messages Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * test(network): stabilize compressed websocket regressions Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(go-sdk): regenerate middleware protobuf binding Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(middleware): clarify websocket skip lifecycle Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(middleware): address WebSocket review feedback Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> --------- Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> |
||
|
|
59479f492a |
feat(k8s): add namespace-per-workspace support (RFC 0011 Phase 3) (#2656)
* feat(k8s): add namespace-per-workspace support (RFC 0011 Phase 3) Implement three workspace namespace modes for the Kubernetes compute driver: shared (default, preserves current single-namespace behavior), managed (auto-creates/deletes namespaces per workspace), and operator (pre-provisioned namespaces with dynamic discovery via label selector or drop-in allowlist file). Key changes: - WorkspaceMode enum and namespace resolution in driver config - Managed namespace lifecycle with ServiceAccount and OpenShift SCC annotation propagation - Cluster-wide sandbox CR watchers for managed/operator modes - NamespaceValidator (Exact/Prefix/Allowlist) for SA token auth - Workspace-aware credential secret storage - Helm ClusterRole for multi-namespace RBAC - Gateway config, architecture, and reference docs Signed-off-by: Derek Carr <decarr@redhat.com> * test(k8s): add e2e tests for workspace namespace modes Add end-to-end tests for managed and operator workspace modes introduced in RFC 0011 Phase 3. The managed mode tests verify namespace creation with correct labels, ServiceAccount provisioning, sandbox CR placement, and namespace survival with remaining sandboxes. The operator mode tests verify rejection of unlabeled and nonexistent namespaces. The positive operator path (sandbox in labeled namespace) is known to fail due to an RBAC gap and will be addressed separately. Also fixes Helm 4 compatibility: move SPDX license headers inside conditional guards in 8 chart templates to prevent empty comment-only documents, and fix a trailing whitespace trimmer in clusterrole.yaml that concatenated the license header with apiVersion. Adds cleanup sweep in with-kube-gateway.sh to remove managed and operator namespaces before Helm uninstall, and mise tasks for running each mode independently. Signed-off-by: Derek Carr <decarr@redhat.com> * feat(k8s): add operator namespace label watcher Spawn a background kube::runtime::watcher in the K8s driver that watches namespaces matching the configured label selector and populates the OperatorNamespaceAllowlist at runtime. The driver owns the allowlist and exposes its Arc so the server can share the same set with the SA token authenticator. create_sandbox now gates pod creation on the allowlist in operator mode — workspaces whose namespace is not yet labeled are rejected at resource render time rather than silently proceeding. Workspace lifecycle itself is unaffected; only sandbox (resource) creation is gated. Signed-off-by: Derek Carr <decarr@redhat.com> * fix(k8s): harden operator mode and address review findings Close the fail-open gap in operator mode when only operator_namespace_file is configured: the allowlist is now created unconditionally in operator mode (fail-closed from startup). Implement the namespace file watcher using the notify crate, following the TLS hot-reload pattern (parent-directory watch, 1s debounce, ConfigMap symlink-swap safe). The file format is a JSON array of namespace name strings. Additional fixes from the 10-reviewer audit: - Change allowlist rejection from InvalidArgument to FailedPrecondition so callers know the request may succeed later once the namespace is provisioned. - NamespaceValidator::Allowlist now holds the OperatorNamespaceAllowlist newtype instead of a raw Arc<RwLock<BTreeSet>>, eliminating silent denial on RwLock poison. - Verify LABEL_MANAGED_BY and LABEL_GATEWAY_ID ownership before deleting a managed namespace. - Replace fixed 5s sleep in operator e2e test with a 30s poll loop. - Add Helm validation for workspaceMode values. - Fix Helm README type column and description for operator fields. - Add insert/remove methods to OperatorNamespaceAllowlist; label watcher now uses them instead of reaching through shared(). - Reject configs with both operator_namespace_label and operator_namespace_file set. Signed-off-by: Derek Carr <decarr@redhat.com> * feat(k8s): add workspace-level compute driver RPCs and harden RBAC Decouple namespace lifecycle from sandbox lifecycle by adding EnsureWorkspace/DeleteWorkspace RPCs to the ComputeDriver service. Namespace creation now happens before credential storage and namespace deletion happens on workspace delete, fixing credential storage in managed workspace mode. - Add EnsureWorkspace and DeleteWorkspace proto RPCs with implementations across all compute drivers (K8s managed delegates to ensure_namespace/delete_namespace_if_empty; others no-op) - Wire ensure_workspace into provider create/update/refresh paths so the namespace exists before the credential driver writes secrets - Wire delete_workspace into workspace deletion for cleanup - Remove delete_namespace_if_empty from sandbox deletion path - Scope ClusterRole secrets access to non-shared workspace modes - Add TODO for TLS cert hot-reload in sandbox gRPC client - Harden e2e tests with control-plane sandbox resolution assertions - Fix docker image save --platform flag for OCI index manifests Signed-off-by: Derek Carr <decarr@redhat.com> * fix(k8s): address re-review findings and add test coverage - Use server-side apply for TLS secret sync (fixes second sandbox creation failure when TLS is enabled) - Scope gateway-ID label selector unconditionally across all workspace modes (fixes operator reads/watches/deletes seeing foreign sandboxes) - Validate operator allowlist in EnsureWorkspace and DeleteWorkspace RPCs (prevents credential writes to namespaces outside the allowlist) - Extend ClusterRole secrets patch+delete to all non-shared modes with credential driver enabled (fixes operator credential storage RBAC) - Validate namespace ownership on 409 conflict in ensure_namespace (prevents adopting unowned namespaces in managed mode) - Replace delete_namespace_if_empty with unconditional delete_namespace letting Kubernetes cascade cleanup (fixes stuck terminating CRs) - Strengthen NetworkPolicy TODO to cover both managed and operator modes - Extract selector and ownership logic into testable free functions - Add unit tests for gateway-ID selectors and namespace ownership - Add Helm ClusterRole RBAC tests for operator credential driver Signed-off-by: Derek Carr <decarr@redhat.com> * ci(k8s): add workspace managed and operator mode e2e to CI Wire the existing e2e:kubernetes:workspace-managed and e2e:kubernetes:workspace-operator mise tasks into the branch-e2e workflow so they run alongside the other core Kubernetes e2e suites. Both are gated by run_core_e2e and included in the Core E2E result gate. Signed-off-by: Derek Carr <decarr@redhat.com> * test(k8s): add e2e tests for workspace namespace modes Add 7 new e2e tests covering workspace namespace lifecycle, TLS secret copying, ownership conflict detection, DNS-1123 validation, operator namespace preservation, and dynamic label watcher behavior. Fix async sandbox deletion race condition in existing tests by polling sandbox list instead of asserting immediately after delete. Signed-off-by: Derek Carr <decarr@redhat.com> * fix(k8s): grant secrets/patch unconditionally and backfill gateway-id labels Address two review findings: 1. RBAC: server-side apply (PATCH) is used for TLS secret sync in multi-namespace modes, but the ClusterRole only granted patch when the kubernetes-secrets credential driver was enabled. Grant patch unconditionally for non-shared modes since TLS sync always needs it; keep delete gated on the credential driver. 2. Upgrade safety: the new gateway-id label selector would orphan legacy Sandbox CRs that predate its introduction. Add a startup backfill in shared mode that patches any managed Sandbox CR missing the gateway-id label before the driver begins serving requests. Signed-off-by: Derek Carr <decarr@redhat.com> * fix(k8s): address workspace namespace review findings Signed-off-by: Derek Carr <decarr@redhat.com> * fix(k8s): address follow-up review findings Signed-off-by: Derek Carr <decarr@redhat.com> * fix(k8s): preserve workspace lookup after rebase Signed-off-by: Derek Carr <decarr@redhat.com> * fix(helm): allow managed secret creation Signed-off-by: Derek Carr <decarr@redhat.com> * fix(k8s): stop pods in workspace namespace Signed-off-by: Derek Carr <decarr@redhat.com> * test(k8s): scope pod deletion check to v1alpha1 Signed-off-by: Derek Carr <decarr@redhat.com> * fix(k8s): address workspace namespace review findings Signed-off-by: Derek Carr <decarr@redhat.com> --------- Signed-off-by: Derek Carr <decarr@redhat.com> |
||
|
|
bdabb54cb3 |
fix(security): authenticate extension services (#2638)
* fix(security): authenticate extension services Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * feat(extension-core): verify gateway JWTs Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * refactor(extension-core): keep inbound verification external This should become an extension SDK package. Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(security): harden the extension authentication contract Follow-up hardening on the alpha extension authentication mechanism. Claim contract: - Extension tokens carry an explicit `typ` of `openshell-ext+jwt`. They share a signing key with sandbox-to-gateway admission tokens and were otherwise separated by audience alone, so a verifier that neglects to check `aud` could accept a gateway credential. The header is a second, independent discriminator. - Publish OIDC-shaped discovery at `/.well-known/openid-configuration` so a service configured with only the gateway URL can learn the exact expected issuer and the JWKS location. It is shaped, not compliant: `issuer` is the gateway identity, not the serving URL. Audience agreement: - `MiddlewareManifest` and `InterceptorManifest` gain `expected_audience`. The audience is otherwise configured independently on each side of the boundary, where a mismatch surfaces only as an opaque authentication failure on every call. OpenShell now compares the two and fails at startup. An empty field keeps the check off for existing services. Compatibility: - Add `allow_insecure_transport` per registration. Enabling gateway JWT signing previously made any plaintext endpoint a hard startup failure, including the endpoint form used in our own documentation. The opt-out attaches no credential, is refused by the gateway if a supervisor asks for one, and warns at every startup. - Make the transport requirement kind-aware. A middleware endpoint must be reachable from every sandbox supervisor, so only interceptors may use a gateway-local Unix socket. Credential lifecycle: - Replace the process-global slot map with a supervisor-owned `ExtensionCredentialStore` shared explicitly across the gateway connections the supervisor opens, removing test-order coupling. - Rotate only when a credential is missing or has passed four fifths of its lifetime. Configuration polling ran every ten seconds against fifteen-minute credentials, so each poll re-ran gateway effective-policy resolution and re-minted the gateway token. - Bound credential minting per sandbox, since each request resolves the caller's effective policy. Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs: record alpha extension authentication in RFC appendices Restore the RFC 0009 and 0010 bodies to their accepted text and move every extension-authentication update into appendices instead. An RFC records a decision at a point in time; superseding detail belongs alongside it rather than rewritten into it. RFC 0009's appendix carries the shared contract: claims, authorization, key distribution, the `allow_insecure_transport` replacement for the body's `allow_insecure`, and residual risks. RFC 0010's records only what differs for interceptors and links to it. The existing protocol-extensions appendix, which parked the phase 2 transport question, now points forward to what was built. Also document the audience handshake, the discovery endpoint, the `typ` requirement, and `jti` replay guidance in the extensibility and gateway configuration pages, and correct the middleware transport guidance: middleware endpoints must be reachable from sandbox supervisors, so Unix sockets are not an option there. Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * refactor(extension-core): abstract extension server trust Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(core): update middleware manifest example Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(extension-auth): preserve unsigned gateway compatibility Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(extension-auth): reject cross-domain token replay Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> --------- Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> |
||
|
|
f12f3ef8d5 |
fix(macos): restore Homebrew sandbox callbacks (#2739)
* fix(macos): restore Docker gateway callbacks Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(gateway): reuse reachable primary callback listener Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
d0c6dc3fd8 |
feat(kubernetes): support corporate upstream proxy (#2633)
* feat(kubernetes): support corporate upstream proxy Signed-off-by: loveRhythm1990 <qiuweimin@126.com> * fix(kubernetes): reject proxy_auth_secret_key values Kubernetes cannot create Gateway validation accepted proxy_auth_secret_key values that Kubernetes rejects when creating the Secret (keys longer than 253 bytes, or the reserved "."/".." names), turning an invalid deployment setting into repeated sandbox Pod-provisioning failures instead of a startup error. Reject them in validate_upstream_proxy_config so they fail closed at gateway startup. Signed-off-by: loveRhythm1990 <qiuweimin@126.com> * docs(skill): add corporate upstream proxy checks to debug-openshell-cluster Add a Kubernetes corporate upstream proxy troubleshooting section covering rendered [openshell.drivers.kubernetes] configuration, credential Secret volume events, supervisor arguments and mounts confined to the network- supervising container, and proxy reachability. Signed-off-by: loveRhythm1990 <qiuweimin@126.com> --------- Signed-off-by: loveRhythm1990 <qiuweimin@126.com> |
||
|
|
c4b500a7de |
feat(helm): cert-manager external issuer + OpenShift passthrough Route (#2468)
* fix(core): trust public root CAs alongside the sandbox mTLS CA The supervisor gRPC client only trusted the CA configured via OPENSHELL_TLS_CA, since tonic ClientTlsConfig starts with an empty root store unless with_native_roots()/with_webpki_roots() is also enabled. Deployments where the gateway server certificate is issued by a public CA (e.g. cert-manager against an ACME issuer) caused every supervisor connection to fail the TLS handshake with "UnknownCA", since the sandbox mTLS CA and the server cert issuer were no longer the same. Enable both native and webpki roots in addition to the configured CA. tonic root store is a union of all configured sources, so this does not weaken verification for existing self-signed deployments. webpki-roots (compiled in) is enabled alongside native-roots since the supervisor binary may run in minimal sandbox images without a populated system CA bundle. Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * feat(helm): support external cert-manager issuers and OpenShift Route passthrough Add certManager.serverIssuerRef/clientIssuerRef so the gateway and mTLS client certificates can be issued by a real Issuer/ClusterIssuer (e.g. ACME) instead of only the chart built-in self-signed CA. Add openshiftRoute template for exposing the gateway via a TLS passthrough Route so the gateway keeps terminating its own TLS/mTLS. The server Certificate excludes internal-only SANs (cluster-local, localhost, loopback) when an external issuer is configured, since ACME issuers reject those per CA/Browser Forum baseline requirements. A template-time fail guard catches the misconfiguration at helm install time rather than asynchronously at cert-manager issuance time. Includes Helm unittest coverage for both issuerRef overrides and Route rendering, plus a CI values overlay for lint coverage. Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * docs: document cert-manager external issuer and OpenShift Route Update managing-certificates.mdx with the serverIssuerRef workflow and install-time validation behavior. Add a production section to the OpenShift guide covering passthrough Route with a real certificate. Regenerate Helm README for new certManager and openshiftRoute values. Sync debug-openshell-cluster skill with new troubleshooting steps for ACME issuance failures and supervisor UnknownCA from mismatched CAs. Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(helm,core): address PR review feedback on cert-manager external issuer Addresses all five blocking review items from #2468: 1. Remove .with_native_roots() from supervisor gRPC client -- the supervisor runs inside the user-selected sandbox image, so the image CA bundle is not operator-controlled. Keep .with_webpki_roots() (compiled-in, not user-controlled) alongside the configured CA. 2. Fail at render time when serverIssuerRef.name is set but clientCaFromServerTlsSecret is still true. Add negative Helm test. 3. Remove clientIssuerRef -- changing only clientIssuerRef breaks both directions because trust bundles are not modeled separately. Change serverIssuerRef.kind default from ClusterIssuer to Issuer. 4. Add server.oidc.issuer and server.oidc.audience to the documented OpenShift production Helm command. Add Access Control prerequisite. 5. Fail at render time when openshiftRoute.enabled and disableTls are both true. Add negative Helm test. Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(drivers): strip GATEWAY_TLS_SERVER_NAME from Docker and Podman env Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(helm,drivers): guard default clientCaSecretName and add env-strip tests Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * feat(tls): SNI-based dual certificate for internal and external server TLS Split the gateway server certificate into two: an internal cert issued by the chart's own CA (for supervisor connections via cluster-local SANs) and an external cert issued by an operator-configured Issuer such as ACME/Let's Encrypt (for CLI and Route access via public SANs). The gateway uses SNI-based certificate selection: connections whose SNI hostname matches external_server_names receive the external cert; all others (including those with no SNI) receive the internal cert. Security improvement: remove .with_webpki_roots() from the supervisor gRPC client so supervisors trust only the chart CA, closing a MITM vector via publicly-trusted certificates in user-supplied container images. Key changes: - Add DualCertResolver with SNI-based cert selection and full test coverage - Add external_cert_path, external_key_path, external_server_names to TlsConfig - Validate partial external cert config (error on cert-without-key or vice versa) - Validate empty external_server_names when external cert is configured - Split cert-manager templates into internal + external Certificate resources - Add Helm guards for misconfigured external issuer (empty serverDnsNames, internal-only SANs with external issuer, conflicting clientCaFromServerTlsSecret) - Update gateway-config.mdx, managing-certificates.mdx, openshift.mdx docs - Update debug-openshell-cluster skill for dual-cert troubleshooting Signed-off-by: Pi Agent <agent@openshell.local> * fix(drivers): strip GATEWAY_TLS_SERVER_NAME in VM driver and correct comments Add the same GATEWAY_TLS_SERVER_NAME environment stripping to the VM compute driver that Docker, Podman, and Kubernetes drivers already perform. Without this, a sandbox user on the VM driver could override the TLS server name the supervisor verifies. Fix stale comments in Docker and Podman drivers that referenced 'with WebPKI roots trusted' — WebPKI roots are explicitly not trusted after the tls-webpki-roots removal. Use tls-ring instead of bare channel for tonic in openshell-core so the TLS API (ClientTlsConfig, Endpoint::tls_config) is available without pulling in any root certificate store. Signed-off-by: Pi Agent <agent@openshell.local> * fix(tls,helm): wildcard SNI matching and Route host validation Add RFC 6125 single-level wildcard matching to DualCertResolver so external_server_names entries like *.example.com correctly match SNI hostnames like gw.example.com. Previously only exact matches worked, silently falling back to the internal cert for wildcard configurations. Add a Helm fail guard in route.yaml that rejects openshiftRoute.host values not listed in certManager.serverDnsNames when an external issuer is configured — catches cert/route hostname mismatches at install time instead of at TLS connect time. Quote the host field in route.yaml for robustness. Signed-off-by: Pi Agent <agent@openshell.local> * fix(helm): address blocking review items — client-CA guard, wildcard Route, serverIssuerRef gate 1. Remove the obsolete guard rejecting serverIssuerRef + clientCaFromServerTlsSecret=true. The internal server certificate is always signed by the chart CA (the same CA that signs the client cert), so clientCaFromServerTlsSecret=true is correct — its filtered ca.crt is exactly the right trust anchor. The old workaround (mounting openshell-ca-tls directly) unnecessarily exposed the CA private key to the gateway container. Remove the client-CA overrides from docs, CI overlay, and production examples. 2. Route host validation now supports wildcard certificates per RFC 6125: single-level wildcards like *.example.com match gateway.example.com but not deep.sub.example.com. Require an explicit openshiftRoute.host when an external issuer is configured — without one, OpenShift generates a hostname absent from serverDnsNames. 3. Reject serverIssuerRef.name when certManager.enabled is false — the external certificate, its Secret mount, and the gateway TLS config all require cert-manager to be enabled. Validated on ROSA (dev.dyee.p3) with branch-built images: - Fresh install with letsencrypt-prod ClusterIssuer - SNI dual-cert: external hostname served Let's Encrypt cert - Supervisor mTLS via internal cert path: ConnectSupervisor accepted - Client CA volume: filtered ca.crt from internal server secret (no key) - CLI connected via Route + OIDC Helm tests: 81 pass across 7 suites. Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> --------- Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> Signed-off-by: Pi Agent <agent@openshell.local> |
||
|
|
0f8fad23c4 |
feat(sandbox): add stop and start operations (#2653)
* feat(sandbox): add suspend and resume operations Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(server): preserve lifecycle work after cancellation Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(server): reconcile ambiguous lifecycle outcomes Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(server): complete suspended session cleanup Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(vm): preserve suspension state on resume failure Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(server): retry retained lifecycle transitions Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(server): clean sessions after suspend reconciliation Signed-off-by: Seth Jennings <sjenning@redhat.com> * test(sandbox): cover deleting suspended sandbox Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(kubernetes): preserve progressing sandbox suspension Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(kubernetes): bound suspend status polling Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(kubernetes): detect legacy sandbox suspension Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(tui): render suspended sandbox phases Signed-off-by: Seth Jennings <sjenning@redhat.com> * refactor(sandbox): rename suspend and resume lifecycle Signed-off-by: Seth Jennings <sjenning@redhat.com> * perf(server): clean stopped sessions on transition Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(kubernetes): fail fast on rejected stop Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(compute): fence stale restart lifecycle events Signed-off-by: Seth Jennings <sjenning@redhat.com> --------- Signed-off-by: Seth Jennings <sjenning@redhat.com> |
||
|
|
0120535efc |
feat(proxy): bind static credentials to provider endpoints (#2510)
* feat(proxy): bind static credentials to provider endpoints Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * test(e2e): verify static credential endpoint isolation Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * docs(provider): explain static credential endpoint binding Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * fix(e2e): use valid endpoint isolation fixtures Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * docs(provider): explain static credential endpoint binding Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * fix(credentials): preserve binding identity across rotations Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(proxy): enforce bindings across request lifecycle Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(proxy): close credential relay gaps Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * docs(credentials): clarify binding failure behavior Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(credentials): hash selected provider profile scope Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(proxy): resolve credentials after request admission Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * docs(credentials): clarify binding failure diagnostics Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(proxy): align single-route credential denials Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(credentials): harden endpoint-bound rotation Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(credentials): enforce identity and authority binding Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(credentials): snapshot provider environment atomically Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(e2e): include authority port in query proxy requests Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(credentials): close credential revocation gaps Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * docs(proxy): explain authority mismatch diagnostics Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(credentials): enforce binding lifecycle invariants Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(provider): reject credential config collisions Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(network): capture credential scope atomically Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(network): distinguish origin and absolute targets Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(provider): isolate endpointless profile credentials Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(network): normalize IPv6 request authorities Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * docs(credentials): clarify endpointless profile isolation Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * feat(policy): bind endpointless provider credentials Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(credentials): use current GCP placeholder revision Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * docs(providers): explain policy credential bindings Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(credentials): cover endpointless fail-closed invariant Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(policy): expect ambiguity rejection at creation Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(server): authenticate rebased policy requests Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * refactor(proxy): share credential mismatch finding builder Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(credentials): cover malformed binding metadata Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(credentials): verify multi-key endpoint isolation Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(e2e): cover same-host credential path denial Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * docs(credentials): document serialized refresh contract Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * refactor(proxy): consolidate L7 log formatting Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * perf(credentials): precompile endpoint binding patterns Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * perf(credentials): share identity epoch revisions Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(proxy): require explicit request default ports Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(policy): validate SigV4 credential sources Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(credentials): preserve endpoint bindings for credential handles Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * feat(go-sdk): expose network credential bindings Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> --------- Signed-off-by: John Myers <johntmyers@users.noreply.github.com> Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> Co-authored-by: John Myers <johntmyers@users.noreply.github.com> |
||
|
|
4cb77a900e |
fix(e2e): separate Podman Machine loopback listeners (#2622)
* fix(e2e): separate Podman Machine loopback listeners Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> * test(e2e): remove shallow harness checks Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> * refactor(e2e): trim Podman listener workaround Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> * fix(e2e): bypass proxies for Podman health probe Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> --------- Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> |
||
|
|
5548405fcb |
feat(credentials): add provider credential storage drivers (#2437)
* feat(credentials): add provider credential storage drivers Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(credentials): harden credential update handling Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(credentials): harden credential driver security, correctness, and performance Address review findings from the credential storage drivers PR: - Route additional_credentials through the driver on refresh to prevent silent data loss for multi-credential providers (e.g. AWS STS) - Clean up stored credential handles on CAS failure during refresh to prevent orphaned secrets in external backends - Enforce namespace validation in the Kubernetes Secrets driver to prevent cross-namespace credential access when allow_reference_namespace is not enabled - Cache Vault Kubernetes auth tokens with 80% TTL to avoid re-authenticating on every credential operation - Parallelize resolve_credentials in all three drivers using try_join_all for faster sandbox startup - Add existingSecret support for the KEK Secret to fix helm template/GitOps workflows where lookup returns empty and regenerates the key - Document RBAC blast radius for the Kubernetes Secrets credential driver and recommend a dedicated namespace Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com> Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> * fix(credentials): add optimistic concurrency, fix thundering herd, parallelize operations Use resourceVersion optimistic concurrency with retry loop for K8s Secret ownership checks to prevent TOCTOU races. Switch Vault token cache from RwLock to Mutex with double-check pattern to prevent thundering herd on cache miss. Parallelize credential store and delete operations across independent keys using try_join_all. Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com> Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> * fix(credentials): handle partial failures, add delete retry, consolidate cleanup Replace try_join_all with join_all in credential store/delete operations to handle partial failures — successfully-stored handles are cleaned up when another key fails. Add retry loop with conflict detection to db-credstore delete_credential, matching the K8s driver pattern. Consolidate 4 manual cleanup_pre_stored_provider_credentials call sites into a single error handler using an async block. Remove inconsistent .trim() from db-credstore validate_handle_owner. Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com> Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> * fix(credentials): fix retry loop guard and remove unprotected validation Remove attempt-count guard from 409/Aborted match arms in retry loops so the post-loop Status::aborted error is reachable after exhausting retries. Previously, last-attempt conflicts fell through to the catch-all error arm, producing misleading Status::unavailable errors. Remove duplicate validation calls that ran after prepare_provider_credential_update but outside the cleanup-protected async block, which would leak pre-stored handles on failure. Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com> Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> * fix(credentials): add workspace/provider UUID to credential backend paths Include workspace and provider ID in credential backend object paths to ensure cross-workspace uniqueness and prevent credential collision (GATOR-1806c9be-01). - Updated credential driver proto to include workspace and provider_id fields - Modified Vault driver to include workspace/provider_id in managed_secret_path - Modified Kubernetes Secrets driver to include workspace/provider_id in credential_owner_id and managed_secret_name - Updated all credential runtime calls to pass workspace/provider_id - Updated tests to use the new signatures This prevents two workspaces sharing the same external credential store from colliding on provider names, which was a critical security issue (CWE-639). * fix(credentials): preserve provider-level expiration for handle-backed credentials Compute effective expiration from both provider and driver values using the earliest non-zero timestamp and skip expired values before insertion (GATOR-1806c9be-02). - Modified resolve_provider_handles to check provider credential_expires_at_ms - Skip expired credentials during resolution instead of returning them - Use effective expiration (min of provider and driver) in resolution results - Fix inference.rs to preserve earliest expiration when merging This ensures handle-backed credentials respect the same expiration semantics as inline credentials. * fix(credentials): stage refresh changes under new handles before validation Stage credential replacements under new immutable handles instead of reusing existing handles to prevent overwriting committed values before validation/CAS (GATOR-1806c9be-03). - Stage credentials with empty existing_handles map to force new handle creation - Validate and CAS before the new values are committed to backend storage - Delete old handles only after successful CAS - On CAS failure, delete only the newly staged handles - This prevents CWE-362/CWE-367 race conditions where failed refreshes could still modify or delete the active credential The fix ensures that a rejected refresh cannot modify the backend object still referenced by the committed provider record. * fix(credentials): add timeouts to credential driver RPCs Apply configured timeouts to both startup capability negotiation and runtime RPCs to prevent indefinite hangs (GATOR-1806c9be-05). - Add DEFAULT_CREDENTIAL_DRIVER_RPC_TIMEOUT_SECS constant (30s) - Apply timeout to GetCapabilities during startup connection - Apply timeout to all runtime RPCs (store, delete, resolve) - Use tokio::time::timeout to bound the entire GetCapabilities operation during startup, not just the socket connection - Return contextual deadline errors on timeout This prevents a faulty or overloaded driver from hanging gateway operations indefinitely. * fix(credentials): fix test to use consistent workspace/provider identity The Kubernetes auth Vault resolve test was constructing a managed path with test-workspace/test-provider-id but sending default/prov-123 in the request, causing validation to reject the request (GATOR-18e32351-01). - Update test to use test-workspace and test-provider-id in the request to match the logical_path construction - This ensures the test exercises the intended code path and validates Kubernetes auth resolution properly The test now passes and correctly validates identity enforcement. * fix(credentials): use unique staging ID for refresh to avoid overwrites Stage refresh replacements under genuinely distinct immutable handles using a unique staging ID to prevent overwriting committed values (GATOR-1806c9be-03). - Generate a unique staging ID using UUID for each refresh operation - Use this staging ID when storing credentials instead of the real provider ID - Pass the same staging ID during cleanup on failure to delete only staged objects - This ensures deterministic paths (Vault) and object names (K8s) don't collide with the committed provider's credentials The fix prevents failed refreshes from silently replacing active credentials or breaking providers by deleting still-referenced backend objects. * fix(credentials): wrap credential driver RPCs in local timeouts Add local tokio::time::timeout wrappers around credential driver RPCs to bound non-compliant or stalled UDS peers (GATOR-1806c9be-05). - Wrap StoreCredential, DeleteCredential, and ResolveCredentials in local timeouts - Return contextual deadline_exceeded errors when timeouts occur - Keep existing gRPC timeout metadata for compliant implementations - GetCapabilities during startup was already wrapped in previous commit This ensures a faulty local driver cannot hang gateway operations indefinitely, even if it accepts the connection but never responds to the RPC. * fix(credentials): preserve ownership for staged refreshes Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(credentials): bound startup capability probe Signed-off-by: Seth Jennings <sjenning@redhat.com> * test(provider): authenticate credential handler requests Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(ci): grant actions read to credential driver e2e Signed-off-by: Seth Jennings <sjenning@redhat.com> --------- Signed-off-by: Taylor Mutch <taylormutch@gmail.com> Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com> Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> Signed-off-by: Seth Jennings <sjenning@redhat.com> Co-authored-by: Taylor Mutch <taylormutch@gmail.com> Co-authored-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> |
||
|
|
537805568d |
feat(sandbox): honor OCI image working directories (#2530)
* feat(sandbox): honor Docker OCI working directories Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> * fix(sandbox): honor effective workspace access Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> * test(sandbox): cover enforced workspace denial Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> * docs(docker): explain effective workdir checks Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> * fix(sandbox): validate effective workspace writes Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> * fix(sandbox): reserve supervisor control roots Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> * refactor(sandbox): centralize control paths Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> * fix(sandbox): reserve OCI runtime mount roots Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> --------- Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> |
||
|
|
8328412959 |
feat(vm): export driver traces over OTLP (#2564)
Continue distributed traces across the gateway-to-driver process boundary and export VM driver spans to the same OTLP/gRPC collector. The driver reports as the distinct openshell-driver-vm service. Updated the gateway architecture and configuration reference with a generic external-driver forwarding contract. Instrumented: - Every RemoteComputeDriver RPC injects the active W3C trace context into tonic metadata. Managed VM readiness and runtime initialization give startup capability probes stable parent operations rather than isolated root spans. - A tonic service layer creates fixed, low-cardinality server spans for every ComputeDriver RPC. New handlers inherit tracing automatically; failures record OpenTelemetry error status and the gRPC status code. - Background provisioning remains attached to CreateSandbox after the RPC returns without extending the RPC span lifetime. - Provisioning records image preparation, bootstrap image resolution, overlay preparation, lifecycle configuration, pre-launch hooks, guest preparation, and launcher spawn as child spans. - VM startup reconciliation roots one trace for the persisted-sandbox scan, with per-sandbox restore and provision operations beneath it. The root remains open until all spawned restore tasks finish. - Delete cleanup records its own child operation. Design notes: - The gateway forwards its configured OTLP endpoint to managed external drivers. SDK `OTEL_*` variables continue to own sampling, batching, limits, headers, and transport tuning. - The VM driver has its own tracer provider and service resource so trace backends preserve the service boundary. - RPC operation names come from an explicit method mapping, keeping cardinality bounded without parsing the protobuf descriptor set at runtime. - Propagation uses a remote SpanContext for spawned provisioning. This keeps one trace while allowing the CreateSandbox server span to finish when the RPC response is sent. - Startup restoration is independent of gateway requests. It begins at the VM driver reconciliation span rather than attaching to an unrelated RPC. - Existing tracing events remain on the logging path. The OpenTelemetry layer exports spans only and excludes the SDK exporter callsites to avoid recursive traces. - Export configuration failures do not prevent the driver from serving, and buffered spans are drained during graceful shutdown. - Trace fields identify drivers, sandboxes, images, lifecycle phases, and gRPC outcomes without recording credentials, sandbox tokens, or request query parameters. Refs #2507 Signed-off-by: Kris Hicks <khicks@nvidia.com> |
||
|
|
905b554c7c |
refactor(network): consolidate proxy egress pipeline (#2373)
* refactor(network): introduce shared egress pipeline Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(network): cover shared proxy egress paths Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * refactor(network): make destination authorization explicit Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * refactor(network): pin proxy relay policy context Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(network): lock relay generation contracts Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(network): establish phase zero compatibility baseline Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * feat(policy): detect ambiguous network endpoints Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * feat(sandbox): fail closed on invalid policy updates Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * refactor(network): invalidate relays on policy changes Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * docs(policy): document validation failure posture Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(network): cover validation and middleware egress Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(config): move policy failure mode to gateway toml Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(network): name proxy contracts by behavior Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(network): align overlap validation with endpoint selection Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(network): expect hard loopback denial Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(network): match declared endpoint denial Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(policy): preserve path-specific endpoint overrides Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(network): respect hard-blocked host gateways Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(network): reconcile proxy refactor with main Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(network): preserve CONNECT policy generation Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(policy): cover runtime endpoint glob semantics Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(server): reject ambiguous policies before persistence Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * docs(policy): explain ambiguity preflight behavior Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(policy): compare body limits within protocol Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(proxy): avoid global tracing capture race Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * chore(server): format rebased provider tests Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(server): authenticate rebased policy requests Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(sandbox): retain runtime on middleware outage Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(server): preflight provider composition activation Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(sandbox): distinguish runtime failure transitions Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> --------- Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> Signed-off-by: John Myers <johntmyers@users.noreply.github.com> Co-authored-by: John Myers <johntmyers@users.noreply.github.com> |
||
|
|
d220d89468 |
feat(compute): negotiate gateway callback listeners (#2492)
* feat(compute): query gateway listener requirements Signed-off-by: Evan Lezar <elezar@nvidia.com> * feat(compute): add Podman listener requirements Signed-off-by: Evan Lezar <elezar@nvidia.com> * test(docker): use default gateway bind address Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(gateway): avoid wildcard primary listener Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(podman): validate callback listener discovery Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(server): support split dual-stack listeners Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(podman): support legacy rootless listener discovery Signed-off-by: Evan Lezar <elezar@nvidia.com> * test(e2e): accept loopback plaintext rejection Signed-off-by: Evan Lezar <elezar@nvidia.com> * docs(agent): add callback listener diagnostics Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(server): restrict compute callback listeners Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(podman): validate local callback port Signed-off-by: Evan Lezar <elezar@nvidia.com> * test(server): clarify callback listener contract Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(podman): require pasta for local callbacks Signed-off-by: Evan Lezar <elezar@nvidia.com> * docs(gateway): document RPM listener default Signed-off-by: Evan Lezar <elezar@nvidia.com> * refactor(server): keep listener provenance diagnostic-only Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(compute): preserve callback listener isolation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(e2e): remove Podman callback relay Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(packaging): preserve Podman callback loopback Signed-off-by: Drew Newberry <anewberry@nvidia.com> * ci(e2e): run VM smoke on nested-virt runner Signed-off-by: Drew Newberry <anewberry@nvidia.com> * ci(e2e): gate VM smoke on usable KVM Signed-off-by: Drew Newberry <anewberry@nvidia.com> * ci(e2e): probe KVM through VM driver Signed-off-by: Drew Newberry <anewberry@nvidia.com> * ci(e2e): tolerate hosted KVM denial Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(server): close traced futures before assertions Signed-off-by: Drew Newberry <anewberry@nvidia.com> * revert: remove tracing test stabilization Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Evan Lezar <elezar@nvidia.com> Signed-off-by: Drew Newberry <anewberry@nvidia.com> Co-authored-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
fa24299092 |
feat(gateway): export traces over OTLP (#2534)
Add an opt-in OTLP/gRPC trace exporter to the gateway. Export is enabled
by the presence of an `[openshell.gateway.otlp]` table with an endpoint;
there is no separate toggle.
Instrumented:
- Inbound request server spans, named for the RPC (`$service/$method`) or
`{method} {path}` for plain HTTP. They continue valid W3C `traceparent`
context when present and start a new trace otherwise. gRPC spans also
carry `rpc.system`, `rpc.service`, `rpc.method`, and trailer-derived
`rpc.grpc.status_code`.
- Compute driver calls (create, delete, list, get, validate, watch) as
client spans anchored on the `ComputeDriver` contract.
- Store reads and writes as children of the current request or loop span.
- Work with no inbound request: compute driver initialization, the sandbox
reconcile sweep, provider credential refresh tick, and driver watch events.
Each roots one operation trace so its child work does not arrive as anonymous
single-span traces.
This is deliberately not exhaustive. Auth, policy evaluation, and
middleware remain uninstrumented, as do store lifecycle calls (`ping`,
`close`) that a readiness poll would turn into a span per tick. The aim is
a useful trace tree at a reviewable size; coverage can grow against real
traces.
Design notes:
- The TOML table owns whether and where to export. The SDK `OTEL_*`
variables own how; sampling, batching, and limits are not mirrored into
gateway config.
- The OpenTelemetry layer exports spans only. Existing `tracing` events
remain on the stdout and sandbox-log paths and are not copied into trace
payloads.
- Telemetry never blocks the gateway. A malformed endpoint logs an error
and disables export rather than failing startup, and buffered spans are
drained during graceful shutdown.
- Failed spans carry error status without a separate `error.type` attribute.
Request spans use HTTP status and gRPC response trailers; driver spans use
the returned gRPC status; autonomous loop spans record failed results
explicitly. Store spans exempt `UniqueViolation` and `Conflict`, because
those errors report expected contention such as a held lease or an
optimistic-concurrency retry.
- The compute driver is reachable only through `TracedDriver::call`, so a
call cannot skip its span. This is the client half of a client/server pair
and the single place to inject context if drivers move out of process.
- Tests share one process-wide subscriber and in-memory exporter because
`tracing` caches callsite interest globally.
Inbound W3C trace context is propagated into gateway request spans. Context
is not yet injected into outbound driver calls, so a future out-of-process
driver would still need propagation at the `TracedDriver` seam.
The Helm chart is intentionally unchanged, so OTLP export cannot yet be
enabled on a chart-deployed gateway.
Refs #2507
Signed-off-by: Kris Hicks <khicks@nvidia.com>
|
||
|
|
9c019a93f5 |
Wire authorization into workspace model (#2445)
* feat(auth): implement RFC 0011 Phase 2 workspace authorization Signed-off-by: Derek Carr <decarr@redhat.com> * fix(auth): address PR review feedback on workspace authorization - Docker e2e: add --health-port and switch readiness probe from `openshell status` to `curl /healthz`, fixing a false-positive readiness check in OIDC mode where the CLI exited 0 without actually contacting the gateway - ListWorkspaces: move membership filtering from post-query N+1 lookups into a SQL EXISTS subquery so pagination applies to the visible set, not the global ordering. Add generic list_with_membership to the persistence layer. - Descriptor validator: reject role/scope fields on unauthenticated and sandbox auth modes, and allow-list workspace_role as user/admin and global_role as platform_admin to catch typos at startup Signed-off-by: Derek Carr <decarr@redhat.com> * fix(server): use authed request in delete telemetry test The workspace authorization added by the Phase 2 auth changes requires a Principal on every delete request. The delete-telemetry test was still using a bare Request::new, so extract_principal failed before the handler could acquire the delete gate, causing a 5-second timeout flake. Signed-off-by: Derek Carr <decarr@redhat.com> * fix(auth): address gator review findings for workspace authorization - Inject unauthenticated-local-dev principal in no-auth gateway mode so handlers that call extract_principal() always find one. - Cap label-selector membership query at MAX_PAGE_SIZE instead of u32::MAX to bound the in-memory read. - Authorize workspace membership before resolving workspace existence in all sandbox RPCs to prevent workspace-name enumeration by non-members. - Remove dead_code allow on AuthorizedWorkspace.workspace now that callers use the normalized name from the authz result. Signed-off-by: Derek Carr <decarr@redhat.com> * fix(auth): close workspace-name oracle and label-selector truncation Swap authorize-before-resolve ordering in 27 handlers across provider.rs, service.rs, policy.rs, and workspace.rs to prevent CWE-203 workspace-name enumeration by non-members. Add combined membership+label SQL query (list_with_membership_and_selector) to both persistence backends so ListWorkspaces with label selectors no longer silently drops results beyond the first page of membership matches. Signed-off-by: Derek Carr <decarr@redhat.com> * test(auth): add non-member rejection and membership+label persistence tests Add comprehensive test coverage for workspace authorization changes: - Non-member rejection tests across all 44 workspace-scoped handlers (sandbox, provider, service, policy, workspace, inference) verifying PERMISSION_DENIED is returned instead of NOT_FOUND to prevent CWE-203 workspace-name oracle - Persistence test for list_with_membership_and_selector verifying SQL-level membership EXISTS + label filtering, multiple predicates, no-match cases, and pagination Signed-off-by: Derek Carr <decarr@redhat.com> * fix(auth): format merged import line in sandbox tests Signed-off-by: Derek Carr <decarr@redhat.com> * fix(auth): address gator re-review findings on workspace authorization - Fix TUI unconditionally setting providers_v2_enabled after provider refresh; read the actual gateway setting via GetGatewayConfig at startup instead - Fix SQLite json_extract with dotted label keys (e.g. example.com/env) by quoting the key in the JSON path - Add authed_request wrappers to upstream OCI identity tests that were missing a principal after rebase - Add test proving GetGatewayConfig is accessible without Platform Admin - Add test for dotted/prefixed Kubernetes-style label key filtering Signed-off-by: Derek Carr <decarr@redhat.com> * fix(auth): address second gator re-review findings - Loosen GetGatewayConfig from platform_admin to scope-only so workspace users can discover providers_v2_enabled during sandbox creation with inferred-provider commands; update proto descriptor, descriptor validation, and RFC 0011 access table - Add validate_label_selector to handle_list_workspaces and escape single quotes in SQLite json_extract interpolation (CWE-89 defense-in-depth) - Re-fetch providers_v2_enabled after TUI gateway switch so the new gateway's capability is reflected - Add e2e test for workspace user with inferred-provider command - Add persistence test for adversarial label keys with SQL injection attempts - Add handler test for invalid label selector rejection in ListWorkspaces Signed-off-by: Derek Carr <decarr@redhat.com> * fix(auth): address third gator review findings - Cap label selector pairs at 64 (CWE-400) to bound SQLite dynamic SQL - Add SCOPE_ONLY_METHODS allowlist for scope-without-role RPCs (CWE-863) - Normalize ID-based data-plane handlers to return NOT_FOUND for unauthorized sandboxes, closing the cross-workspace oracle (CWE-203) - Fix TUI provider profile cache lookup key mismatch for legacy providers with empty profile_workspace - Add whoami to CLI skill reference command tree - Update TUI skill doc with workspace, provider, and settings coverage - Document scope/workspace orthogonality on GetGatewayConfig proto Signed-off-by: Derek Carr <decarr@redhat.com> * fix(auth): extend CWE-203 normalization to policy.rs sandbox handlers GetSandboxConfig and GetSandboxLogs in policy.rs had the same fetch-before-authorize pattern that leaked cross-workspace sandbox existence. Promote fetch_and_authorize_sandbox to pub(super) and use it from both sandbox.rs and policy.rs handlers. Signed-off-by: Derek Carr <decarr@redhat.com> * test(auth): update assertions for CWE-203 sandbox ID normalization Cross-workspace sandbox access via ID-based handlers now returns NOT_FOUND instead of PERMISSION_DENIED to prevent existence inference. Update the unit test and OIDC e2e assertion to match. Signed-off-by: Derek Carr <decarr@redhat.com> * fix(auth): narrow CWE-203 error mapping and correct whoami output formats Only remap PERMISSION_DENIED to NOT_FOUND in fetch_and_authorize_sandbox and RevokeSshSession, letting INTERNAL and UNAUTHENTICATED propagate as-is. Fix whoami --output format values in cli-reference.md to match the actual CLI (table/json/yaml, not text/json). Signed-off-by: Derek Carr <decarr@redhat.com> * fix(ci): share network namespace with Keycloak in containerized CI In GitHub Actions job containers, Docker port publishing lands on the host, not inside the job container. Detect this environment and attach Keycloak to the job container's network namespace instead, with hardened defaults (cap-drop ALL, no-new-privileges, loopback-only listener). Signed-off-by: Derek Carr <decarr@redhat.com> --------- Signed-off-by: Derek Carr <decarr@redhat.com> |
||
|
|
bc14018cad |
feat(sandbox): use policy-first OCI image identity (#2509)
* feat(sandbox): use policy-first OCI image identity Closes #2331 Preserve per-field policy omission, derive Docker and Podman fallbacks from the inspected immutable image, and resolve the final numeric identity before starting agent children. Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> * fix(sandbox): preserve declared process identities Keep explicit policy values and OCI-declared names intact, defer passwd lookup until a primary GID is required, and refresh stale policy examples. Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> * fix(supervisor): reuse resolved OCI identity Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> * fix(supervisor): allow Linux pre-exec arguments Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> * fix(kubernetes): protect resolved sandbox identity Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> * fix(sandbox): prepare workspace for OCI identity Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> * refactor(sandbox): own only workspace root Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> * fix(sandbox): harden partial identity drops Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> * test(sandbox): scope OCI image e2e to Docker Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> * fix(sandbox): narrow OCI identity fallback scope Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> * test(podman): cover OCI identity launch Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> * fix(podman): exercise OCI fallback in E2E Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> --------- Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> |
||
|
|
7955c8309b |
feat(k8s): support configuring workspace PVC storageClassName (#2463)
The Kubernetes driver's default workspace PVC never set storageClassName, so on clusters with no default StorageClass the PVC stayed Pending and sandbox creation failed. Add a workspace_storage_class option to KubernetesComputeConfig, wired through SandboxPodParams into the generated volumeClaimTemplates. When non-empty it sets storageClassName; empty preserves the current behavior of relying on the cluster default StorageClass. Expose it via the OPENSHELL_K8S_WORKSPACE_STORAGE_CLASS env var on both the standalone driver and the embedded gateway runtime defaults, and via the server.workspaceStorageClass Helm value. Closes #2442 Signed-off-by: lr90 <qiuweimin@matrixorigin.cn> |
||
|
|
f00ad23a26 |
fix(podman): tolerate shutdown transport closes (#2498)
Signed-off-by: Evan Lezar <elezar@nvidia.com> |
||
|
|
deced8716b |
refactor(policy): extract shared L7 endpoint validation (#2389)
Move L7 endpoint semantic checks into the openshell-policy crate so both
profile lint and the runtime validator share one implementation. This
eliminates drift between the two validation paths.
The shared validator covers 9 checks: unknown protocol, rules/access
mutual exclusivity, JSON-RPC family access rejection, json-rpc requires
rules, non-JSON-RPC protocol requires rules or access, MCP requires
rules when allow_all is false, rules-would-deny-all detection,
deny_rules require protocol, and deny_rules require base allow set.
Changes rules/deny_rules fields to Option<Vec<...>> so absent vs empty
is distinguishable at lint time. Adds is_effectively_empty() to
L7AllowProfile for deny-all detection of allow: {} objects. Makes
rules_would_deny_all MCP-aware by checking tool/params.name selectors
before classifying a rule as deny-all. Adds params field to
L7AllowProfile so MCP tool selectors survive proto round-trip.
Signed-off-by: Grace Smith <gsmith@redhat.com>
Signed-off-by: Grace Smith <grasmith@redhat.com>
|
||
|
|
77e5c32217 |
feat(sandbox,gateway): route sandbox egress through corporate HTTP proxy (#2245)
* feat(sandbox,gateway): route sandbox egress through corporate HTTP proxy - Chain sandbox egress through a corporate HTTP proxy so outbound traffic from within the sandbox respects the host proxy settings - Forward sandbox proxy environment variables to the generated Podman config so the proxy is applied consistently to Podman-managed workloads Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(sandbox,podman): make corporate proxy routing operator-owned The operator-configured corporate egress proxy was injected under the conventional HTTPS_PROXY/HTTP_PROXY/NO_PROXY names as defaults beneath sandbox spec/template environment, so a sandbox creator could redirect egress at an arbitrary proxy or disable proxying with NO_PROXY=*. Route the boundary through reserved, supervisor-only variables (OPENSHELL_UPSTREAM_HTTPS_PROXY/HTTP_PROXY/NO_PROXY) written in the Podman driver's required-variable tier. Any sandbox-supplied value under a reserved name is stripped before the operator value is applied, so the supervisor never observes a reserved proxy variable the operator did not set. The supervisor now reads only the reserved names and ignores the conventional proxy variables the sandbox controls. Add the reserved proxy variables to the supervisor-only child-environment denylist so the corporate proxy URL and any embedded credentials are not inherited by the sandbox workload, which reaches egress through the local policy proxy and never needs them. Signed-off-by: Philippe Martin <phmartin@redhat.com> * feat(sandbox,podman): deliver corporate proxy credentials via secret file Proxy credentials were embedded inline in the proxy URL, so they were stored in gateway.toml and exposed in container metadata via 'podman inspect'. Reject inline 'user:pass@' credentials in https_proxy/http_proxy at startup (parsed with the url crate rather than hand-splitting), and add a proxy_auth_file option pointing at a 'user:pass' file. The driver stages that file as a per-sandbox root-only Podman secret, mounts it at a fixed path, and exports only the path in the reserved OPENSHELL_UPSTREAM_PROXY_AUTH_FILE variable, so the credential never appears in config, environment, or container metadata. Reading the file fails closed on a missing, empty, or control-character-bearing value. The supervisor reads the credential from the mounted file and builds the Proxy-Authorization: Basic header, rejecting control characters, and no longer derives credentials from URL userinfo. The auth-file path is added to the child-environment strip list, and the generated gateway.toml is written owner-only (mode 0600). Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(sandbox,podman): fail closed on invalid upstream proxy configuration The reserved OPENSHELL_UPSTREAM_* variables are an operator-owned egress boundary, but the supervisor treated present-but-invalid values as unset: an unsupported or malformed proxy URL was ignored with a warning, an unreadable auth file proceeded without credentials, and a malformed credential silently became unauthenticated. Any of these could quietly downgrade the corporate proxy boundary to direct dialing or unauthenticated proxy access. Make every configured-but-invalid proxy or auth setting fatal to supervisor proxy startup, emit an OCSF ConfigStateChange failure event before refusing, and share URL validation semantics between the Podman driver and the supervisor through a single validator in openshell-core (parse_upstream_proxy_url), so a value accepted at sandbox-create time can never be rejected in-container or vice versa. Inline user:pass@ URL credentials are now fatal in the supervisor too (previously warn-and-strip), matching the driver. Unset or empty variables still mean no proxy; only present-but-invalid values fail. Error paths never include credential content. Addresses the fail-closed review item on #2245. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(sandbox,podman): finish the fail-closed upstream proxy credential contract The upstream proxy URL already had a single shared validator, but the credential did not: the Podman driver rejected only CR/LF/NUL while the supervisor rejected every control character, so a credential accepted at sandbox-create time (e.g. one containing a tab) could still be rejected in-container. Present-but-whitespace reserved OPENSHELL_UPSTREAM_* values were also silently treated as unset, quietly downgrading the operator's egress boundary to direct dialing. Add parse_upstream_proxy_credential to openshell-core as the single source of truth for the documented user:pass credential form (non-empty user, no control characters, trimmed) and use it in both the Podman driver's secret staging and the supervisor's Proxy-Authorization header construction. Error variants carry no payload so credential content can never leak into messages. Make a present-but-empty reserved variable fatal to supervisor proxy startup instead of meaning "unset"; only fully unset variables disable the proxy. The driver correspondingly rejects an empty no_proxy at config time so it can never inject a value the supervisor refuses. Addresses the remaining fail-closed credential/config review item on Signed-off-by: Philippe Martin <phmartin@redhat.com> #2245. * fix(sandbox,podman): close remaining fail-open upstream proxy config paths Two configuration paths could still silently run without the proxy boundary the operator believed was in effect. A no_proxy bypass list configured without any https_proxy/http_proxy was accepted by both the driver and the supervisor and simply meant "dial everything directly". Reject it on both sides, exactly like the existing proxy_auth_file-without-proxy rule: an operator who wrote a bypass list assumed proxying was active, so accepting it hides a fail-open state. The gateway.sh dev script guarded proxy settings with [[ -n "${VAR:-}" ]], which conflates unset with explicitly-empty and dropped the latter before the gateway's validation could see it. Use ${VAR+x} instead so a set-but-empty variable is written into gateway.toml and rejected at startup by validate_proxy_config rather than silently discarded. Addresses the remaining fail-open configuration review item on #2245. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(sandbox,podman): reject upstream proxy URLs with path, query, or fragment parse_upstream_proxy_url accepted URLs like http://proxy.corp.com:8080/some/path and silently discarded everything after host:port, a lenience inherited from the original supervisor parser. A forward proxy is addressed by host:port only, so extra components indicate a misconfiguration (for example a pasted endpoint URL) and silently truncating them violates the present-but-invalid-is-fatal contract enforced everywhere else in this configuration surface. Reject a path, query, or fragment in the shared validator with a new UnexpectedComponent error. A bare trailing slash remains accepted because the url crate normalizes an absent http path to "/", making the two indistinguishable. Both the Podman driver (gateway startup) and the supervisor (sandbox startup) inherit the rule through the shared parser, keeping their semantics identical by construction. Addresses the proxy URL component review item on #2245. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(sandbox,podman): remove plain-HTTP upstream proxy support The http_proxy path tunneled plain-HTTP requests through the corporate proxy with CONNECT to port 80 and then sent origin-form requests down the tunnel. Conventional enterprise forward proxies expect plain HTTP as absolute-form requests sent directly over the proxy connection, and commonly refuse CONNECT to port 80, so the setting looked supported but failed against typical deployments. Tunneling also blinds the proxy to the one protocol it could inspect. Narrow the feature to TLS (CONNECT) egress only, which is the conventional and already-correct case: plain-HTTP requests now always dial the destination directly, and only client CONNECT tunnels chain through the corporate proxy. Remove the http_proxy config field, the --sandbox-http-proxy / OPENSHELL_SANDBOX_HTTP_PROXY driver surface, the reserved OPENSHELL_UPSTREAM_HTTP_PROXY variable, and the UpstreamScheme plumbing. The feature never shipped, so this is a clean removal; a stray http_proxy key in gateway.toml still fails loudly through the config's deny_unknown_fields. Removing the plain-HTTP proxy branch also removes its host-gateway special case; the architecture doc now documents the real host-gateway behavior (add driver-injected host aliases to the reserved NO_PROXY list) instead of an invariant the HTTPS path never implemented. Plain-HTTP forwarding through a corporate proxy can return later as absolute-form forwarding behind its own design review. Addresses the plain-HTTP forwarding review item on #2245. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(sandbox,podman): escape generated TOML and require explicit proxy URL form Escape backslashes, quotes, and control characters when gateway.sh writes proxy values into gateway.toml, so a hostile or unusual environment value cannot corrupt the config or inject extra keys. Restrict the upstream proxy URL grammar to the documented http://host:port form: a scheme-less value is no longer normalized to http:// and a missing port is no longer silently defaulted to 80. Docs, README, and CLI help now state the explicit-form requirement consistently. Also fix a test-only call of handle_tcp_connection that was missing the upstream_proxy argument added in an earlier commit. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(sandbox,podman): preserve tunneled bytes read with the CONNECT response The CONNECT handshake reads the corporate proxy's response in chunks, so the read that completes the header block can also contain the first tunneled payload bytes. Those bytes were discarded, silently corrupting the start of the tunnel for server-speaks-first destinations or proxies that coalesce writes. connect_via now returns a PrefixedStream that replays any bytes received past the response terminator before reading from the socket again; writes pass through unchanged. Direct dials wrap the stream with an empty prefix so downstream relay and TLS paths keep a single stream type, and tls_connect_upstream is generalized to any AsyncRead + AsyncWrite stream. Adds regression coverage for a combined response/payload read and for prefix replay ordering. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(sandbox,podman): gate cleartext proxy Basic auth behind an explicit opt-in Proxy-Authorization: Basic is base64 over the plain-TCP connection to the http:// corporate proxy, so anyone on the network path between the sandbox host and the proxy can recover the credential. Sending it is now an explicit operator decision instead of an implicit side effect of configuring proxy_auth_file. Add a proxy_auth_allow_insecure driver setting (CLI --sandbox-proxy-auth-allow-insecure, env OPENSHELL_SANDBOX_PROXY_AUTH_ALLOW_INSECURE), delivered to the supervisor as the reserved OPENSHELL_UPSTREAM_PROXY_AUTH_ALLOW_INSECURE variable. Fail-closed pairing on both sides: an auth file without the acknowledgement is rejected at gateway startup and at supervisor startup, as is the acknowledgement without an auth file or any value other than 'true'. gateway.sh writes the key only as a TOML boolean; a non-boolean value is emitted as a quoted string so the gateway rejects it at startup instead of risking injection. Documents the exposure prominently in the gateway config reference, the driver README, and the sandbox architecture doc. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(sandbox,podman): honor port qualifiers and resolved addresses in NO_PROXY Two divergences from the documented NO_PROXY contract: - A port-qualified entry (internal.corp:8443) was silently stripped to its hostname and bypassed the proxy for every port, excluding traffic the operator never listed. Entries now keep the optional :port qualifier (also on IP and CIDR entries) and only apply to that destination port; a trailing qualifier that is not a valid port stays part of the pattern instead of widening the entry. - IP and CIDR entries only matched IP-literal hosts, so a bypass like 10.0.0.0/8 never applied to hostnames resolving into that range. NO_PROXY evaluation now sees the validated resolved addresses: an IP/CIDR entry matching through resolution authorizes a direct dial of only the addresses it contains, so a bypass scoped to an internal range cannot widen into a direct dial of addresses outside it. Hostname-level matches (loopback, wildcard, domain entries, IP-literal hosts) keep authorizing all validated addresses. proxy_for is replaced by decision(host, port, resolved) returning either the proxy endpoint or the permitted direct-dial subset, and dial_upstream restricts the direct connect to that subset. Adds regression coverage for port-scoped bypasses on domain, IP, and CIDR entries, invalid port qualifiers, resolved-address matching, and split-resolution subset dialing. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(sandbox,podman): bind proxied CONNECT tunnels to validated addresses The CONNECT request sent to the corporate proxy carried the destination hostname, so the proxy resolved the name itself and the addresses that had passed SSRF and allowed_ips validation were discarded. Split-horizon DNS or rebinding at the proxy could then reach internal or otherwise unapproved destinations through a tunnel the supervisor logged as validated, and IP-range policy could not be enforced at all on proxied dials. CONNECT now targets a validated resolved address by default: the proxy performs no DNS resolution and the tunnel stays bound to the answer the supervisor checked. The hostname still travels inside the tunnel (TLS SNI, application Host), so destination servers behave normally. In split-horizon networks, operators point the gateway host at the corporate resolver so internal names validate to their internal addresses. For proxies whose ACLs filter on hostnames and reject IP CONNECT targets, a new proxy_connect_by_hostname opt-in (CLI --sandbox-proxy-connect-by-hostname, env OPENSHELL_SANDBOX_PROXY_CONNECT_BY_HOSTNAME, reserved OPENSHELL_UPSTREAM_PROXY_CONNECT_BY_HOSTNAME) restores hostname CONNECT, documented as re-opening proxy-side resolution and making the proxy's ACLs the effective egress control. Fail-closed pairing on both sides: the opt-in without a proxy, or any value other than 'true', is fatal. Adds regression coverage for the IP CONNECT request line (including IPv6 bracketing and hostname non-leakage), the hostname opt-in, and the config pairing rules in driver and supervisor. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(sandbox,podman): reject empty port after bracketed IPv6 proxy host http://[fd00::1]: passed the explicit-port check because the bracketed branch only tested for a colon after the bracket, then fell back to port 80 — violating the fail-closed http://host:port contract and potentially sending configured Basic credentials to an unintended service. Require a non-empty suffix after ]:, matching the unbracketed branch, and cover the case in the shared parser tests. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(sandbox,podman): fall back across validated addresses in proxied CONNECT The direct path hands TcpStream::connect the whole validated address list and it falls back across them, but the validated-IP CONNECT path attempted only the first address, so a dual-stack destination could fail through the corporate proxy even when a later validated address was reachable. connect_via_validated tries each validated address in order under one aggregate CONNECT_HANDSHAKE_TIMEOUT budget, returning the first success; when every attempt fails the error names the attempt count and carries the last failure. An empty address list is rejected up front. Adds regressions for first-fails/second-succeeds fallback (asserting both CONNECT request lines), the aggregate all-addresses failure message, and the empty-list rejection. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(sandbox,podman): strip new reserved proxy vars and complete docs/tests Follow-ups from review: - Add OPENSHELL_UPSTREAM_PROXY_AUTH_ALLOW_INSECURE and OPENSHELL_UPSTREAM_PROXY_CONNECT_BY_HOSTNAME to the supervisor-only strip list so workload child processes never inherit them, matching the documented contract for the other reserved proxy variables, and cover both in the supervisor-only variable test. - Document all five OPENSHELL_SANDBOX_* proxy variables in the mise run gateway help text, marked Podman-only and stating the auth-file/acknowledgement pairing, and complete the gateway-key list in the Podman README. - Add NO_PROXY composition coverage for bracketed IPv6 entries with port qualifiers, bare IPv6 entries, and IPv6 CIDR matching against an IPv6-literal host and against a hostname's resolved addresses, including the port-qualified CIDR form. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(sandbox,podman): cap each proxied CONNECT attempt within the shared budget A proxy that accepted the first CONNECT request but never responded consumed the entire aggregate handshake timeout, so later validated addresses were never tried and the hang defeated the multi-address fallback. Each attempt is now time-boxed to its fair share of the time remaining before the shared deadline (remaining / attempts_left): a hanging attempt is cut off with enough budget left for every remaining address, while time a fast failure does not use rolls over to later attempts and the total never exceeds CONNECT_HANDSHAKE_TIMEOUT. A timed-out attempt is recorded like any other failure, and the aggregate error distinguishes all-attempted from budget-exhausted runs. Adds a first-hangs/second-succeeds regression driven through a test-visible budget parameter so it runs in about a second instead of a real 30s window. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(sandbox,podman): deliver corporate proxy config on the supervisor argv The proxy settings were injected as reserved OPENSHELL_UPSTREAM_* container environment variables. The driver only wrote names the operator configured, but container runtimes layer the spec environment over ENV values baked into the sandbox image, so an image could supply NO_PROXY=*, enable hostname CONNECT, or point an unconfigured deployment at an attacker-controlled proxy whenever the operator left a field unset. The settings now travel as supervisor command-line arguments (--upstream-proxy, --upstream-no-proxy, --upstream-proxy-auth-file, --upstream-proxy-auth-allow-insecure, --upstream-proxy-connect-by-hostname) built by the driver from operator config. The driver sets the container entrypoint and command explicitly, so neither sandbox spec/template environment nor image ENV can influence argv, and an omitted flag genuinely means unconfigured — in every supervisor topology, since the supervisor no longer consults its environment for these settings at all. Credentials stay on the root-only secret mount; only the mount path appears on argv. The reserved environment names, their strip-list entries, and the env-based validation surface are removed. UpstreamProxyConfig::from_args replaces from_env, reusing the same shared fail-closed validation and pairing rules keyed by the CLI flag names. Driver tests now assert the argv contract, including that sandbox-supplied environment cannot add, remove, or redirect proxy flags; supervisor tests cover from_args mapping and its pairing rules. Signed-off-by: Philippe Martin <phmartin@redhat.com> * docs(sandbox): align proxy comments with the argv transport The argv migration left comments describing the configuration as reserved environment variables ("reserved value", "present-but-empty variable", "reserved upstream proxy variables"). Rephrase them as driver-supplied arguments and operator settings so the documented trust boundary matches the implementation. Comment-only change. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(sandbox): parse bracketed IPv6 authorities in client CONNECT targets parse_target split the CONNECT authority at the first colon, so an IPv6-literal target like [2001:db8::1]:443 always failed port parsing and IPv6-literal clients could never reach policy evaluation; a regression test even locked in that failure. Parse the RFC 3986 bracketed form and return the host bracket-free, matching what DNS resolution, SSRF validation, NO_PROXY matching, and the upstream CONNECT builder expect. Unclosed brackets, a missing or empty port after the bracket, and non-numeric ports are rejected; unbracketed behavior is unchanged. Replaces the failure-locking test with success coverage for bracketed targets and adds malformed-bracket rejection cases. Signed-off-by: Philippe Martin <phmartin@redhat.com> * test(podman): cover proxy-auth secret cleanup across lifecycle failures The per-sandbox proxy-auth credential secret is staged before the container is created and removed on cleanup, but no test proved the cleanup paths actually issue the secret removal. Add Podman-stub tests that drive create_sandbox to a container-create failure and to a start failure, and delete_sandbox for an out-of-band deletion, asserting each path issues the DELETE for the per-sandbox proxy-auth secret so a credential can never outlive the sandbox that owned it. Signed-off-by: Philippe Martin <phmartin@redhat.com> * docs: list corporate proxy keys in the Podman compute-driver overview The Fern Podman driver section enumerated its gateway.toml keys but omitted the corporate egress proxy settings. Add https_proxy, no_proxy, proxy_auth_file, proxy_auth_allow_insecure, and proxy_connect_by_hostname with a pointer to the gateway configuration reference for the full contract. No navigation change: the reference folder already includes the gateway configuration page. Signed-off-by: Philippe Martin <phmartin@redhat.com> * test(sandbox): cover the SSRF-to-TLS composition across the proxy tunnel Existing tests exercised validated-IP CONNECT and the upstream-TLS helper independently, but not the full boundary. Add an end-to-end regression that stands up a fake corporate proxy tunneling to a fake TLS server and drives the real path: connect_via_validated CONNECTs to the validated address, the proxy splices the tunnel, and tls_connect_upstream verifies the upstream certificate against the original hostname carried in SNI. It asserts the CONNECT authority is the validated IP and never the hostname, that verification succeeds for the matching hostname, and that a mismatched hostname is rejected — proving a rebinding or split-horizon substitution behind the proxy cannot pass certificate verification. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(sandbox,podman): bound proxy-auth reads, reject port 0, fix stale comment Three review findings: - CWE-400: the proxy-auth credential file was read with an unbounded read_to_string on both the driver (sandbox-create) and supervisor (startup) paths, so a huge file or a special file such as /dev/zero could exhaust memory. Add a shared bounded reader in openshell-core that rejects non-regular files, caps the size at 4 KiB, and reads at most that many bytes; the driver runs it via spawn_blocking. Covers oversized, special-file, and missing-path cases on both sides. - Reject an upstream proxy URL with port 0: it passed the explicit-port check and startup validation but is not a connectable TCP port, so every proxied dial would fail later. Add a typed ZeroPort error with shared-validator and Podman-config tests. - Reword a driver-config comment that still described a 'reserved variable' to match the argv transport. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(sandbox): open proxy-auth file non-blocking to reject FIFOs promptly read_upstream_proxy_credential_file opened the path with a blocking File::open before the regular-file check, so a configured FIFO with no writer would block open() indefinitely — hanging sandbox creation on the driver and supervisor startup. Open with O_NONBLOCK on Unix so the open returns immediately, then reject the non-regular file as before; O_NONBLOCK has no effect on the later read of a regular file. Adds a mkfifo regression asserting the reader returns promptly with a non-regular-file error instead of hanging. Signed-off-by: Philippe Martin <phmartin@redhat.com> * test(podman): cover corporate proxy egress across driver and supervisor The existing corporate-proxy tests construct config structs or call CONNECT helpers directly, so none of them detect a break in the wiring between layers: gateway TOML deserialization, the Podman argv and secret-mount semantics, supervisor CLI parsing, or policy denial before proxy contact. Add a Podman e2e that drives the whole chain against a fake authenticated forward proxy and asserts that an approved TLS request traverses it with a validated-IP CONNECT, a policy-denied destination is refused with 403 without ever reaching the proxy, credentials arrive through the mounted per-sandbox secret, and deleting the sandbox removes that secret. SupportContainer is a new harness fixture. Unlike ContainerHttpServer it probes readiness with a TCP connect rather than an HTTP GET, so it can host a forward proxy and TLS servers, and it exposes container logs and network IP for assertions. Signed-off-by: Philippe Martin <phmartin@redhat.com> * test(podman): restart the gateway on the proxy-config panic path The panic cleanup for the temporary corporate proxy configuration restored the gateway TOML but left the gateway process running with the temporary configuration still loaded, which could poison later test binaries in the same run. Nothing restarted it: the only ManagedGateway is the short-lived one inside restart_gateway, and its Drop only calls start, which does not reload config for an already-running gateway. Restore and synchronously stop/start the gateway in Drop, and set restored only after the normal restore and restart both succeed so a failed restart no longer suppresses the fallback. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(supervisor-network): derive the loopback proxy bypass from resolved IPs The automatic bypass treated the host string "localhost" as proof that the destination was loopback and returned every resolved address. A sandbox controls its own /etc/hosts and resolve_socket_addrs consults it before DNS, so a workload could map localhost to any policy-allowed address and dial it directly, escaping the operator proxy and the inspection and audit boundary it exists to provide. Check the resolved addresses instead: the name bypasses only when the resolution is non-empty and every address is loopback. A mixed answer is not partially honored, and an IP literal is still authoritative for itself. A spoofed localhost falls through to the entries below, so an explicit operator NO_PROXY entry is still honored. This matches the trust model detect_trusted_host_gateway already applies to the same hosts file, which validates the mapped address rather than the alias. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(podman): clean up proxy-auth secret when container already deleted The delete_sandbox early-return path cleaned up the token secret but skipped the proxy-auth secret, leaking it on disk. Also update tests for recent API changes (Optional socket_path, workspace field, list-based container lookup). Signed-off-by: Philippe Martin <philippe@openshell.dev> Signed-off-by: Philippe Martin <phmartin@redhat.com> --------- Signed-off-by: Philippe Martin <phmartin@redhat.com> Signed-off-by: Philippe Martin <philippe@openshell.dev> |
||
|
|
59f7839f6b |
fix(auth): report gateway authentication status (#2435)
* fix(auth): report gateway authentication status Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(auth): reuse gateway info for status probe Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
744a65d52e |
fix(driver-podman): avoid panic when HOME is unset on macOS (#2327)
* fix(driver-podman): resolve Podman socket via auto-detection Align Podman socket selection with the existing Docker model: explicit config wins, otherwise probe openshell-core for a responsive socket. This also fixes the original HOME-unset panic, since resolution no longer hardcodes a per-OS default path. - add detect_podman_socket() in openshell-core, mirroring detect_docker_socket - PodmanComputeConfig.socket_path is now Option<PathBuf>, no default - remove default_socket_path() (podman driver) and podman_socket_path() (vm driver), both replaced by the shared detector - update server env override and CLI for the new Option type - add tests: responsive-candidate detection in openshell-core, and config-error (not panic) when no socket is configured or reachable * fix(driver-podman): address review feedback on socket resolution Extract socket resolution into resolve_socket_path, taking the detector as a parameter so tests do not depend on real env vars or the host's actual Podman state. Replace the flaky env-mutating test with three deterministic cases: explicit wins, detected is used when absent, and neither source errors. Fix the socket_path doc comment to describe it from a config user's point of view, matching DockerComputeConfig's docstring. Drop a comment that only made sense next to the Docker driver code. Update the Podman README and gateway docs: they described a fixed per-OS default path that no longer exists, replace with the actual probe-then-fail behavior. * docs(driver-podman): simplify socket default description Previous wording was self-contradictory (says auto-detect on unset, then lists the same var as a probed candidate) and omitted the Linux /run/user/uid/podman/podman.sock candidate. |
||
|
|
80987e91c8 |
docs: fix broken links and small inconsistencies (#2329)
- README: fix github-sandbox tutorial link missing get-started segment - README: replace dead community-sandboxes doc link with the actual repo - README: match supported host list to support-matrix.mdx - architecture/README: list the missing google-vertex-ai-provider doc - SECURITY.md: fix a mis-indented list item - standardize on NVIDIA/OpenShell-Community casing for repo links |
||
|
|
d556748771 |
feat(supervisor-middleware): add network egress middleware (#2027)
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> |
||
|
|
cf4deccd37 |
fix(gateway): probe Docker socket during driver auto-detection (#2303)
Previously, Docker was auto-detected when the CLI was installed or a candidate Unix socket existed. Neither check verified that the Docker API was responsive. A similar check was done when auto-detecting Podman in the past, but was replaced in |
||
|
|
83003e80fc |
feat(interceptors): initial gateway interceptor implementation and reference example (#2005)
* feat(gateway): add descriptor-driven interceptors Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(gateway): add service-reflected interceptors Signed-off-by: Drew Newberry <anewberry@nvidia.com> * wip * fix(gateway): harden interceptor evaluation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(interceptors): label metrics and harden governance smoke Signed-off-by: Drew Newberry <anewberry@nvidia.com> * remove on_error: ignore * feat(gateway-interceptors): emit log annotations Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(examples): govern provider profiles in interceptor Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(gateway): preserve update config annotations Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(providers): support interceptor profile catalogs Signed-off-by: Drew Newberry <anewberry@nvidia.com> * wip * feat(governance-interceptor): sign provider profiles Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(providers): use configured profile sources for refresh updates Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(gateway-interceptors): add phase-specific evaluation payloads Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(providers): compose provider profile sources Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(gateway-interceptors): preserve committed responses Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(gateway-interceptors): reject ambiguous protobuf oneofs Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(gateway-interceptors): validate patch candidates per binding Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(gateway-interceptors): use reflected protobuf codec Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(governance-example): canonicalize signed protobuf hashes Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(gateway): commit policy provenance atomically Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(gateway): close signed governance bypasses Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(gateway): isolate interceptor secrets and authority Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(gateway): snapshot provider profiles per request Signed-off-by: Drew Newberry <anewberry@nvidia.com> * chore(gateway): resolve server clippy warnings Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(server): satisfy provider source clippy lint Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(server): initialize policy test annotations Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(gateway-interceptors): require explicit route allowlist Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(proto): clarify update annotation semantics Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
614c8c164d |
feat(kubernetes): support PVC subPath driver config (#2034)
* feat(kubernetes): support PVC subPath driver config Signed-off-by: mjamiv <michael.commack@gmail.com> * test(kubernetes): cover writable PVC driver config Signed-off-by: mjamiv <michael.commack@gmail.com> * fix(kubernetes): address PVC subPath review feedback Signed-off-by: mjamiv <michael.commack@gmail.com> * fix(kubernetes): address PVC config review follow-up Signed-off-by: mjamiv <michael.commack@gmail.com> * fix(kubernetes): address PVC review follow-ups --------- Signed-off-by: mjamiv <michael.commack@gmail.com> |
||
|
|
8eacb4779f |
feat(kubernetes): add sidecar supervisor topology (#2076)
* feat(kubernetes): add sidecar supervisor topology Add the Kubernetes sidecar supervisor topology, its Helm/Skaffold configuration, topology documentation, and sidecar e2e matrix coverage. Skip root-only sandbox identity rewriting when process enforcement is network-only so the low-permission sidecar process container can start successfully. Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(supervisor): avoid similar process id names Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(supervisor): avoid similar process id names Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(sandbox): avoid similar proxy id names Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * docs(kubernetes): clarify sidecar topology limits Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(kubernetes): keep sidecar process leaf capless Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(kubernetes): refresh sidecar provider env snapshots Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * test(supervisor): align hot-swap identity regression Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(kubernetes): stage sidecar mtls files before proxy chown Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(kubernetes): simplify sidecar supervisor topology Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * chore(helm): reuse sidecar skaffold values Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(supervisor): avoid similar iptables helper names Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(e2e): harden kube gateway wrapper setup Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(supervisor): avoid nft batch rollback on OCP Run nftables setup as individual commands so optional conntrack and log expressions can fail without rolling back required table, chain, and reject rules. Signed-off-by: Seth Jennings <sjenning@redhat.com> Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(kubernetes): preserve process identity in sidecar topology Render sidecar pods with a shared process namespace, keep binary-aware network policy enabled, and move Kubernetes sidecar settings under the nested sidecar config table. Also apply unprivileged Landlock/seccomp setup in NetworkOnly supervisor mode so sidecar topology keeps sandbox child hardening without privileged process setup. Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * refactor(kubernetes): replace sidecar snapshots with control socket Coordinate sidecar policy and provider bootstrap over a local Unix socket so the process leaf no longer reads policy/provider snapshot files. Report entrypoint startup through the control channel and keep gateway credentials confined to the network sidecar. Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * feat(kubernetes): support relaxed sidecar network identity Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(sandbox): satisfy sidecar clippy lint Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * refactor(kubernetes): standardize topology naming Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(sandbox): satisfy linux clippy timeout import Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(kubernetes): support kata sidecar on ipv4 pods Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(kubernetes): satisfy linux clippy for sidecar fallback Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * chore(kubernetes): remove stale supervisor topology references Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(kubernetes): enable sidecar binary policy inspection Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(kubernetes): harden sidecar control boundary Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(kubernetes): couple sidecar supervisor lifecycles Signed-off-by: Taylor Mutch <taylormutch@gmail.com> --------- Signed-off-by: Taylor Mutch <taylormutch@gmail.com> Signed-off-by: Seth Jennings <sjenning@redhat.com> Co-authored-by: Seth Jennings <sjenning@redhat.com> |
||
|
|
bebf440b25 |
fix(helm): propagate supervisor image overrides (#2216)
Signed-off-by: Taylor Mutch <taylormutch@gmail.com> |
||
|
|
10702133a3 |
fix(core): pin supervisor image tag to gateway version for all drivers (#2070)
* fix(core): pin supervisor image tag to gateway version for all drivers The Podman and Kubernetes drivers defaulted the supervisor image to `:latest` via DEFAULT_SUPERVISOR_IMAGE, while the Docker driver already resolved a version-pinned tag. Extract the tag resolution logic into openshell-core so all three drivers use the same OPENSHELL_IMAGE_TAG > IMAGE_TAG > CARGO_PKG_VERSION priority chain. Closes #2068 Signed-off-by: Florent Benoit <fbenoit@redhat.com> * refactor(core): simplify supervisor image tag resolver to slice-based API Remove the Docker driver's wrapper functions and call openshell_core::config::default_supervisor_image() directly. Simplify resolve_supervisor_image_tag to accept &[&str] instead of three separate parameters. Signed-off-by: Florent Benoit <fbenoit@nvidia.com> Signed-off-by: Florent Benoit <fbenoit@redhat.com> --------- Signed-off-by: Florent Benoit <fbenoit@redhat.com> Signed-off-by: Florent Benoit <fbenoit@nvidia.com> |
||
|
|
ed8ce8208f | docs: fix Docker version format from 28.04 to 28.0 (#2136) | ||
|
|
6252aa17c8 |
rfc-0006: add driver config passthrough proposal (#1589)
* docs(rfc): add driver config passthrough proposal Signed-off-by: Evan Lezar <elezar@nvidia.com> * docs(rfc): link driver config proposal PR Signed-off-by: Evan Lezar <elezar@nvidia.com> * docs(rfc): clarify driver config scope Signed-off-by: Evan Lezar <elezar@nvidia.com> * docs(rfc): clarify driver-local config schemas Signed-off-by: Evan Lezar <elezar@nvidia.com> * docs(rfc): clarify driver config extension path * docs(rfc): update driver config baseline * docs(drivers): document bind-mount selinux_label and whitespace rules --------- Signed-off-by: Evan Lezar <elezar@nvidia.com> |
||
|
|
6461677c32 |
feat(policy): accept numeric UIDs for sandbox process identity (#1973)
* feat(policy): accept numeric UIDs in sandbox process identity validation Allow run_as_user and run_as_group to be either the literal 'sandbox' or a numeric UID/GID within [1000, 2_000_000_000]. This removes the hard dependency on a baked-in 'sandbox' user in container images, enabling compute drivers to inject resolved UIDs at sandbox creation. Phase 1 of #1959. Signed-off-by: Seth Jennings <sjenning@redhat.com> * feat(supervisor): accept numeric UIDs for process identity dropping Allow run_as_user and run_as_group to be numeric UIDs/GIDs, removing the hard dependency on a baked-in 'sandbox' user in container images. Changes: - validate_sandbox_user(): accepts numeric UIDs without passwd lookup (logs OCSF event); keeps passwd check for "sandbox" name; rejects non-numeric non-sandbox strings that fail passwd lookup - prepare_filesystem(): passes numeric UIDs/GIDs directly to chown() instead of requiring a passwd entry - drop_privileges(): resolves numeric UIDs/GIDs directly via UID::from_raw / Gid::from_raw; skips initgroups when target uid matches current euid; uses guard conditions before setgid/setuid calls - session_user_and_home(): falls back to ("{uid}", "/sandbox") for numeric UIDs, avoiding a passwd lookup that will fail Re-exports MIN_SANDBOX_UID and MAX_SANDBOX_UID from openshell-policy so callers have consistent range constants. Phase 2 of #1959. Signed-off-by: Seth Jennings <sjenning@redhat.com> * feat(driver-kubernetes): resolve sandbox UID/GID from config or OpenShift SCC annotations Phase 3 of the numeric-UID plan: allow operators to specify explicit sandbox_uid/sandbox_gid in Kubernetes driver config, auto-detect from OpenShift SCC namespace annotations, and propagate resolved values to supervisor container env vars and PVC init container securityContext. Changes: - Add sandbox_uid/sandbox_gid fields to KubernetesComputeConfig - Add SANDBOX_UID/SANDBOX_GID env var constants to openshell-core - Implement resolve_sandbox_identity() to fetch namespace annotations and auto-detect OpenShift SCC UID ranges (sa.scc.uid-range) - Pass resolved UID/GID through SandboxPodParams to pod spec builder - Inject SANDBOX_UID/SANDBOX_GID env vars into supervisor container - Update PVC init container securityContext with resolved UID/GID instead of hard-coded root - Add comprehensive unit tests for resolution logic and annotation parsing (resolve_sandbox_uid, resolve_sandbox_gid, OpenShift SCC annotation parsing) Signed-off-by: Seth Jennings <sjenning@redhat.com> * feat(driver-vm): add configurable sandbox UID/GID and update docs/examples Phase 4 of the numeric-UID plan: replace hardcoded SANDBOX_UID (10001) in VM rootfs preparation with configurable sandbox_uid/sandbox_gid fields. Changes: - Add sandbox_uid/sandbox_gid to VmDriverConfig with serde derives - Pass resolved UID/GID through prepare_sandbox_rootfs_from_image_root to ensure_sandbox_guest_user which writes /etc/passwd/group/gshadow - Update BYOC Dockerfile: remove groupadd/useradd, document runtime UID injection and the ability to skip baked-in sandbox user - Update gateway-config.mdx: document sandbox_uid/sandbox_gid for both Kubernetes (with OpenShift SCC autodetection) and VM drivers - Update sandbox-compute-drivers.mdx: add Sandbox User Identity section explaining numeric UID support across all compute drivers - Update rootfs tests to use non-default UIDs, verify config passthrough Signed-off-by: Seth Jennings <sjenning@redhat.com> * code review changes * fix(supervisor): harden tests for restricted CI container environments Guard tests against CI-specific constraints: root without CAP_SETPCAP, UIDs with no /etc/passwd entry, and restricted /proc access. Signed-off-by: Seth Jennings <sjennings@nvidia.com> Signed-off-by: Seth Jennings <sjenning@redhat.com> --------- Signed-off-by: Seth Jennings <sjenning@redhat.com> Signed-off-by: Seth Jennings <sjennings@nvidia.com> |
||
|
|
43bb030266 |
feat(docker,podman): add SELinux label support for bind mounts (#2092)
* feat(docker,podman): add SELinux label support for bind mounts The Docker Engine structured Mount API does not support SELinux relabelling (:z / :Z). Move user-supplied bind mounts from the structured `mounts` field to the legacy string-format `binds` field, which does support these options. Add a shared `SelinuxLabel` enum (shared/private) to openshell-core so both Docker and Podman drivers accept an optional `selinux_label` field on bind mount configs. For Docker, labels are appended to the bind string; for Podman, they are pushed to the mount options vec. Signed-off-by: Florian Bergmann <fbergman@redhat.com> * fix(docker): reject missing bind source paths on legacy binds Moving user bind mounts from the structured Mount API to the legacy Binds field changed Docker's behavior for missing source directories: the legacy path silently creates them as empty root-owned dirs instead of erroring. Add an explicit Path::exists() check to preserve the fail-fast behavior operators expect. Signed-off-by: Florian Bergmann <fbergman@redhat.com> --------- Signed-off-by: Florian Bergmann <fbergman@redhat.com> |
||
|
|
914da339b4 |
feat(kubernetes): add combined topology config surface (#2074)
* feat(kubernetes): add combined topology config surface Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * docs(kubernetes): clarify topology defaults Signed-off-by: Taylor Mutch <taylormutch@gmail.com> --------- Signed-off-by: Taylor Mutch <taylormutch@gmail.com> |
||
|
|
5477e2f21d |
docs(mcp): fix granular policy lifecycle examples (#2066)
Signed-off-by: Shiju <shiju@nvidia.com> |
||
|
|
7bce1223dc |
feat(policy): add JSON-RPC and MCP L7 policies (#1865)
Add policy schema, proto, provider profile, OPA, and L7 proxy support for `protocol: json-rpc` and `protocol: mcp`. Generic JSON-RPC endpoints match exact method names only, with `method: "*"` as the all-method sentinel; wildcard/glob methods and params matchers are rejected. Parse JSON-RPC request bodies and batches in the forward proxy, deny response-shaped client frames, limit receive-stream GET allowance to MCP endpoints, and redact params in decision logs. Preserve L7 rule params on the proto load path so MCP `tools/call` tool filters behave like YAML-loaded policies. Add MCP conformance coverage, JSON-RPC L7 e2e coverage, and docs for the new protocols and current matcher limitations. Signed-off-by: Kris Hicks <khicks@nvidia.com> Co-authored-by: ddurst <267424412+ddurst-nvidia@users.noreply.github.com> |
||
|
|
ba21bb32a2 |
feat(kubernetes): support agent-sandbox v1beta1 (#2009)
* feat(kubernetes): support agent-sandbox v1beta1 Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * ci(kubernetes): test agent-sandbox api versions Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(kubernetes): retry agent-sandbox raw 404s Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(kubernetes): harden agent-sandbox api setup Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(kubernetes): cache agent-sandbox api version Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * docs(kubernetes): document agent-sandbox upgrade behavior Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * refactor(kubernetes): collapse agent-sandbox api selection Signed-off-by: Taylor Mutch <taylormutch@gmail.com> --------- Signed-off-by: Taylor Mutch <taylormutch@gmail.com> |
||
|
|
f569a0ade6 |
feat(sandbox): proxy-side AWS SigV4 credential signing for CONNECT tunnels (#1638)
Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> Co-authored-by: Russell Bryant <russell.bryant@gmail.com> |
||
|
|
75a317ea45 |
feat(server): support out-of-tree compute drivers via --compute-driver-socket (#1703)
* feat(server): support remote compute driver endpoints Add named remote compute driver endpoint support to the gateway. Remote drivers are selected by a non-reserved compute driver name and either a CLI/env socket endpoint or [openshell.drivers.<name>].socket_path. The VM driver now enters ComputeRuntime through the same acquired remote endpoint path, while Docker, Podman, and Kubernetes retain their in-process drivers. Require --drivers/OPENSHELL_DRIVERS when pairing an ad-hoc socket endpoint so the socket does not imply a magic driver name, and keep reserved in-tree names unavailable for unmanaged socket endpoints. Co-authored-by: Evan Lezar <elezar@nvidia.com> Signed-off-by: st-gr <38470677+st-gr@users.noreply.github.com> Signed-off-by: Evan Lezar <elezar@nvidia.com> * test(server): cover remote compute driver UDS lifecycle Signed-off-by: Evan Lezar <elezar@nvidia.com> --------- Signed-off-by: st-gr <38470677+st-gr@users.noreply.github.com> Signed-off-by: Evan Lezar <elezar@nvidia.com> Co-authored-by: Evan Lezar <elezar@nvidia.com> |
||
|
|
d93293ad12 |
fix(e2e): stabilize local Docker smoke test (#1935)
* fix(docker): honor configured supervisor image Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(cli): isolate ssh from host linker environment Signed-off-by: Evan Lezar <elezar@nvidia.com> --------- Signed-off-by: Evan Lezar <elezar@nvidia.com> |