428 Commits
Author SHA1 Message Date
Artem Lytvyn eb15e1a4c9 feat(sandbox): add --no-login-shell to skip shell startup files on exec (#2852)
* feat(sandbox): add --no-login-shell to skip shell startup files on exec

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* chore(sdk/go): regenerate proto bindings for no_login_shell

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* test(supervisor-process): cover login-shell flag selection

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* feat(sandbox): gate --no-login-shell on supervisor SSH banner

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

---------

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>
2026-08-31 16:17:58 +00:00
Drew Newberry 69a05ebb3b fix(sandbox): complete successful main processes (#2884) 2026-08-28 18:21:20 -07:00
krishicks f68867b869 feat(gateway): identify gateways in exported traces (#2647)
* feat(gateway): add installation name configuration

Add a first-class operator-assigned gateway name with TOML, CLI, environment,
and Helm configuration surfaces. Local gateways default to openshell, while
Helm defaults to the chart fullname; operators sharing a collector across
namespaces or clusters can set a globally distinct name.

Signed-off-by: Kris Hicks <khicks@nvidia.com>

* feat(gateway): identify gateways in exported traces

Attach the configured gateway installation name and compute driver to the
gateway OpenTelemetry resource so operators can filter traces from multiple
installations that share a collector. Forward the gateway name and OTLP
endpoint to managed external drivers so their distinct service resources carry
the same installation identity.

Keep service.name stable per process type, omit blank resource values, and
leave per-span operation names and request attributes unchanged.

Refs #2507

Signed-off-by: Kris Hicks <khicks@nvidia.com>

---------

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-08-27 17:27:34 +00:00
alangou 37072ee81c feat(build): publish OCI SBOM and provenance attestations (#2836)
* feat(build): embed auditable Rust dependency metadata

Signed-off-by: Adrien Langou <alangou@nvidia.com>

* feat(build): publish OCI SBOM and provenance attestations

Signed-off-by: Adrien Langou <alangou@nvidia.com>

---------

Signed-off-by: Adrien Langou <alangou@nvidia.com>
2026-08-27 14:53:33 +00:00
krishicks d0dfb22baf feat(kubernetes): export driver traces over OTLP (#2958)
Mirror the VM, Podman, and Docker driver tracing setup for Kubernetes.
Export standalone driver spans through OTLP/gRPC as the distinct
openshell-driver-kubernetes service, preserve gateway trace context, record
lifecycle operations and gRPC failures, and flush spans on shutdown.

Kubernetes currently runs in-process when selected as a built-in gateway
driver. Use the temporary server-boundary shim shared with Podman and Docker
so traces retain the shape they will have when Kubernetes moves to a
separate process. Move the common ComputeDriver RPC tracing layer into
openshell-otel to keep all drivers aligned.

Propagate the active W3C context through the controller-reserved Sandbox
annotation and enable Agent Sandbox OTLP export in the local k3s workflow.
This connects asynchronous controller reconciliation spans to the originating
OpenShell create trace.

Expose gateway OTLP configuration through Helm and add an Aspire collector
to the local k3s workflow. Extend helm:k3s:forward with OTLP ingest and trace
UI forwarding for Kubernetes and local container gateway development.

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-08-26 21:22:51 +00:00
Evan Lezar 4e992093f8 fix(docker): trace standalone driver over OTLP (#2923)
Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-08-26 06:27:21 +00:00
Vyncint Ng 0e79653a7f feat(tui): show persisted guidance on rejected policy chunks (#2908)
* feat(tui): show persisted guidance on rejected policy chunks

The reviewer's rejection note is stored in PolicyChunk.rejection_reason and
already reaches the TUI in GetDraftPolicyResponse.chunks, but openshell-tui
never read the field. The note was dropped at the last step, so a reviewer
had no way to recall why a chunk had been rejected.

Render it in two places, following the truncate-in-list / full-in-popup
convention in the TUI development guide: a truncated, dimmed suffix on the
list row, and a "Guidance:" line in the detail popup.

Gate the accessor on status == "rejected" rather than on the field alone.
Approving a chunk passes None for the reason and the gateway writes the
field only when Some, so a chunk that was rejected and later approved still
carries the old note. Reading the field unconditionally would surface a
stale rejection on an approved rule.

Part of #1098, Definition of Done item "Rejected chunks show persisted
guidance". The other items on that issue are untouched.

Signed-off-by: Vyncint Ng <115854244+vyncint@users.noreply.github.com>

* fix(tui): scroll the draft detail popup so long guidance stays reachable

A rejection reason has no server-side length cap, so a paragraph-length one
overflowed the fixed 22-row detail popup: the tail was clipped and the later
fields and both action hints were pushed off screen. The same overflow already
affected the unbounded rationale and security_notes fields, so fix the popup
rather than special-case the guidance line.

Split the popup's inner area into a scrolling body and a pinned hint row,
following the pattern in ui/create_provider.rs, and drive it with j/k, the
arrow keys, PageUp, PageDown, g and G. The approve and close controls now stay
on screen at every scroll position, and the bottom border carries the scroll
position.

Wrap free-form values explicitly instead of relying on Paragraph's Wrap, so the
rendered row count is exactly lines.len() and the scroll clamp cannot under-run
the content. Paragraph::line_count would answer the same question but sits
behind ratatui's unstable-rendered-line-info feature.

Add deterministic TestBackend coverage at 80x24 with a 2,000-character reason,
covering the head and tail, the pinned hints, over-scroll clamping, and the
pre-existing long-rationale case. Document the reviewer-facing guidance in the
policy advisor page.

Part of #1098.

Signed-off-by: Vyncint Ng <115854244+vyncint@users.noreply.github.com>

* fix(tui): wrap the denial row and hint footer on a narrow popup

Making the scroll clamp exact meant dropping Paragraph's Wrap, which also
removed wrapping from the lines that do not go through push_wrapped. At 70
columns the denial row clipped mid-value and lost the last-seen timestamp, and
at 60 a pending chunk with scrollable content pushed [Esc] Close past the right
edge of the single hint row. The key still worked, but that is the same
controls-not-visible failure one axis over.

Keep the denial row on one line while it fits and wrap it onto the label indent
when it does not, so the common 80-column layout is unchanged. Pack the footer
hints into as many rows as they need without splitting a hint, and derive the
footer height from that: adding a hint row shrinks the body and can itself
change whether the content scrolls, so the two settle together.

The earlier tests were all 80x24, which is why they missed this. Cover the
denial timestamps and the approve, reject and close hints at 60, 70 and 80
columns, plus the packing helper directly.

Part of #1098.

Signed-off-by: Vyncint Ng <115854244+vyncint@users.noreply.github.com>

* fix(tui): measure popup wrapping in display columns, not chars

wrap_value packed rows by chars().count(), and a CJK glyph is one char but two
terminal columns. A double-width rejection reason therefore produced rows about
twice the popup's width: the right half of every row was clipped, and unlike
the vertical case there is no horizontal scroll to recover it.

Measure width with Span::width(), the same measurement ratatui applies when it
lays cells out, so the wrap agrees with the renderer by construction and no new
dependency is needed. This covers word packing, the hard-break path for a word
wider than the line, the label indent, the footer hint packing, and the denial
row's fits-on-one-line check.

The list row's shortened copy had the same defect through truncate_str, which
counts chars. Leave truncate_str alone for its two existing callers and add
truncate_display for the guidance row.

Cover it with an 80x24 render regression that interleaves markers through the
CJK text: a trailing marker lands on its own short row in both the broken and
fixed layouts, so it would not detect this. The two unit tests measure with an
independent column oracle rather than the helper under test, which would make
them tautological.

Part of #1098.

Signed-off-by: Vyncint Ng <115854244+vyncint@users.noreply.github.com>

---------

Signed-off-by: Vyncint Ng <115854244+vyncint@users.noreply.github.com>
2026-08-25 18:37:36 +00:00
Mrunal Patel 18ce13b9b1 feat(providers): expose actionable OAuth refresh failures (#2887)
* fix(providers): classify OAuth refresh failures

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* test(providers): add Keycloak refresh e2e lane

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(providers): harden OAuth refresh recovery

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(providers): classify post-mint refresh failures

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(providers): tolerate malformed OAuth subtypes

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

---------

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
2026-08-25 18:00:00 +00:00
Seth Jennings 5206bc51b2 feat(sdk): add OAuth Client Credentials support to SDKs (#2907)
* feat(sdk): add renewable client credentials auth

Implement lazy OAuth client-credentials acquisition and renewal for the Python, TypeScript, and Go SDK clients, with shared security conformance coverage and service-account documentation.\n\nCloses #2803

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(sdk): address client credentials review

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(sdk): address client credentials review

Signed-off-by: Seth Jennings <sjenning@redhat.com>

---------

Signed-off-by: Seth Jennings <sjenning@redhat.com>
2026-08-25 17:53:47 +00:00
Simon Scatton 455883905a fix(python): remove CLI from wheel (#2321)
The Maturin-based wheel packaging was a historical remnant from when the local gateway launch path and OpenShell CLI were coupled in one binary. The gateway and CLI now ship as standalone artifacts, so the Python distribution should contain only the SDK.

Build a single platform-independent setuptools wheel, verify that it cannot contain native code or an openshell entry point, and simplify the release jobs and documentation for SDK-only PyPI installs.

Signed-off-by: Simon Scatton <sscatton@nvidia.com>
2026-08-25 13:20:42 +00:00
Russell Bryant aa848f164d fix(core): fall back to podman CLI when no API socket is found (#1858)
* fix(core): fall back to podman CLI when no API socket responds

Auto-detection only checked well-known Podman socket paths, so a Podman
machine exposing its API socket at a non-standard location went
undetected. The symlink at a well-known path is not always present; it
varies by Podman version, machine provider, and platform.

Extend detect_podman_socket() to fall back to podman CLI discovery when
no well-known candidate responds: podman info --format json determines
whether the service is local or remote, and podman machine inspect
resolves the host-side forwarded socket for VM-backed machines. All
existing callers (driver auto-detection, the Podman driver, and the VM
driver's container-engine fallback) pick this up without change.

Select the machine backing the active Podman connection instead of the
first entry from podman machine inspect, honoring podman's connection
precedence: CONTAINER_CONNECTION, then CONTAINER_HOST (mapped to a
connection by URI), then the containers.conf default. An explicit
endpoint that maps to no known machine is left unresolved rather than
guessing an unrelated machine. When CONTAINER_HOST is an explicit unix://
socket, use that path directly since podman info connects through it.

Update the gateway config reference and the Podman driver README, which
described probe-only detection.

Signed-off-by: Russell Bryant <rbryant@redhat.com>

* fix(core): inspect podman machine by name and bound discovery probes

Address two startup defects in Podman socket auto-detection:

- Remote discovery ran `podman machine inspect` with no arguments, which
  inspects only `podman-machine-default`. On a host whose default
  connection is a different machine, selection fell back to the first
  (wrong) entry, so the gateway could operate against the wrong backend.
  Resolve the active machine first and inspect it by name, trying the
  `-root`-stripped machine for rootful connections and returning None
  when no machine can be mapped instead of substituting another.

- `podman info`, `podman machine inspect`, and `podman system connection
  list` used unbounded `Command::output()`, so a stalled machine, SSH
  connection, or helper could hang gateway startup indefinitely. Route
  all three through a bounded runner that kills and reaps the child on a
  documented deadline and returns None so detection can continue.

Signed-off-by: Russell Bryant <rbryant@redhat.com>

* fix(core): bound podman probe stdout drainage and kill descendants

The previous timeout only bounded waiting for the direct child to exit.
Draining stdout still called read_to_end, which waits for pipe EOF — a
Podman probe can leave a daemonized descendant (SSH multiplexer, gvproxy)
that inherited the stdout pipe, so drainage could wait indefinitely even
after the direct child exited, defeating the deadline.

Run each probe as its own process-group leader, cap stdout drainage by
the same deadline via a channel, and on expiry kill the whole process
group so descendants holding the pipe are terminated and EOF is reached.
Add a regression test where the direct child exits but a backgrounded
descendant keeps stdout open.

Signed-off-by: Russell Bryant <rbryant@redhat.com>

* fix(core): make podman probe deadline absolute against escaped descendants

The prior drainage timeout still ended in an untimed rx.recv() after
killing the process group. A descendant that escaped the probe's process
group (e.g. via setsid) while holding the stdout pipe would not be killed,
so read_to_end never saw EOF, the reader never sent, and that recv() could
block gateway startup indefinitely.

On drainage timeout, best-effort kill the process group and return None
immediately, abandoning the reader thread instead of waiting on it again.
The deadline now bounds the whole call regardless of what descendants do.
Add a regression test whose stdout-holding descendant escapes the process
group (`set -m`) and assert the call still returns promptly.

Signed-off-by: Russell Bryant <rbryant@redhat.com>

---------

Signed-off-by: Russell Bryant <rbryant@redhat.com>
2026-08-25 11:49:22 +00:00
Philippe Martin 0a1f246587 feat(sandbox,podman): trust corporate CA for https:// proxies and intercepted TLS (#2512)
* feat(sandbox,podman): trust corporate CA for https:// proxies and intercepted TLS

The corporate proxy chaining only accepted plain http:// proxy URLs, so
operators whose forward proxy terminates TLS with a private corporate CA
had no way to reach it, and TLS-intercepting proxies (mitmproxy, squid
ssl-bump) that re-sign tunneled server certificates broke every upstream
handshake after CONNECT.

The supervisor now accepts https:// proxy URLs: it wraps the connection to
the proxy in TLS before the CONNECT handshake, verifying the proxy
certificate against the built-in Mozilla roots, the system CA bundle, and an
optional operator corporate CA bundle. The upstream dial returns a
Plain/Tls stream enum consumed generically by the relay paths.

The corporate CA is delivered as a driver-supplied command-line argument
(--upstream-proxy-ca-bundle), never an environment variable, matching the
hardened proxy-config model where a sandbox image cannot influence the
operator's egress boundary. It is folded into the sandbox combined trust
bundle (write_ca_files) and the L7 upstream verification store
(build_upstream_client_config) at startup, so intercepted upstream
handshakes succeed and sandbox workloads trust the re-signed certificates.
Configuration is fail-closed: a CA bundle set without a proxy, or an
unreadable or certificate-free file, is fatal rather than silently
weakening the trust boundary.

The shared parse_upstream_proxy_url validator accepts https:// (recording
the scheme so the driver and supervisor agree), keeping the explicit-port
requirement. The Podman driver gains a proxy_ca_bundle operator setting
(TOML, --sandbox-proxy-ca-bundle, OPENSHELL_SANDBOX_PROXY_CA_BUNDLE) that
bind-mounts the host PEM read-only into the sandbox (a CA certificate is not
secret) and points --upstream-proxy-ca-bundle at it, with a create-time
readability check.

The standalone dev gateway task passes OPENSHELL_SANDBOX_PROXY_CA_BUNDLE
through to the generated podman config, so a local gateway can be pointed at
a TLS-intercepting proxy without hand-editing the regenerated TOML.

Refs #1792

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(sandbox): reject CA bundles with valid PEM framing but invalid X.509 DER

The proxy CA bundle validation counted PEM blocks that base64-decoded
successfully, but did not verify the decoded bytes were accepted as
trust anchors by RootCertStore. A bundle with syntactically valid PEM
framing but invalid DER would pass the startup check while contributing
zero usable anchors, causing opaque TLS failures at runtime instead of
a fail-closed startup error.

Validate decoded certificates through RootCertStore::add_parsable_certificates
and reject the bundle unless at least one is accepted.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(sandbox,podman): allow proxy auth without insecure acknowledgement for https:// proxies

For an https:// proxy the Proxy-Authorization credential travels inside
the verified TLS session, so the proxy_auth_allow_insecure
acknowledgement is unnecessary. Previously both http:// and https://
proxies required it, producing a misleading cleartext-risk diagnostic
for a path that is already encrypted.

Skip the requirement when the proxy URL uses https://; the
acknowledgement is still tolerated if set. Updated in both the
supervisor and Podman driver validation paths, with docs and tests.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix: format

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(kubernetes): use truly unsupported scheme in proxy validation test

https:// is now a supported proxy scheme after a13c4dce, so the
unsupported-scheme test must use a genuinely unsupported scheme.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(e2e): sign the https proxy fixture listener cert with a CA

The corporate-proxy E2E fixture served a single `openssl req -x509`
certificate as its TLS listener identity. OpenSSL marks that certificate
`basicConstraints: critical, CA:TRUE`, and rustls refuses a CA
certificate presented as an end-entity certificate (CaUsedAsEndEntity).
The supervisor's TLS handshake with the proxy therefore failed, the
upstream dial errored, and the workload's CONNECT was dropped without a
response, so podman_corporate_proxy_trusts_ca_bundle_for_https_proxy
failed on the approved destination while policy denial still worked.

Generate a corporate CA and a separate listener leaf signed by it, serve
the leaf chain, and publish only the CA as the bundle the supervisor
trusts. This is what an intercepting proxy actually presents, and it
exercises the corporate-CA trust path rather than pinning the listener
certificate itself.

Refs #1792

Signed-off-by: Philippe Martin <phmartin@redhat.com>

---------

Signed-off-by: Philippe Martin <phmartin@redhat.com>
2026-08-24 21:10:25 +00:00
krishicks e457974a52 feat(docker): export driver traces over OTLP (#2851)
Mirror the VM and Podman driver tracing setup for Docker. Export Docker
driver spans through OTLP/gRPC as the distinct openshell-driver-docker
service, preserve gateway trace context, record lifecycle and asynchronous
provisioning operations, and report gRPC failures.

Docker currently runs in-process when selected as a built-in gateway driver.
Add the same temporary server-boundary shim used by Podman so traces retain
the shape they will have when Docker moves to a separate process. Generalize
the gateway provider selection for both in-process drivers and share the OTLP
collector fixture across Docker, Podman, and VM tracing tests.

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-08-24 18:39:21 +00:00
grs 40d1b48666 feat(provider): support for SPIFFE backed token exchange (#1970)
* feat(provider): add ability to request token exchange instead of client credentials as OAuth grant_type

Signed-off-by: Gordon Sim <gsim@redhat.com>

* test(proxy): add further tests for token exchange

Signed-off-by: Gordon Sim <gsim@redhat.com>

* test(provider): add runnable example for token exchange

Signed-off-by: Gordon Sim <gsim@redhat.com>

* test(e2e): cover Podman token exchange grants

Signed-off-by: Gordon Sim <gsim@redhat.com>

* refactor(oauth): extract duplicated functionality from server and supervisor

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(provider): evict nearest-to-expiry entry from intermediate token cache

Signed-off-by: Gordon Sim <gsim@redhat.com>

* doc(supervisor): add podman example for token exchange

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(provider): withhold token-exchange subject credentials

Signed-off-by: Gordon Sim <gsim@redhat.com>

---------

Signed-off-by: Gordon Sim <gsim@redhat.com>
2026-08-24 05:42:30 +00:00
John T. MyersandJohn Myers 679fe4c334 fix(policy): validate the applicable advisor candidate (#2850)
* fix(policy): bind reviews to applicable candidates

Build and validate the exact effective-policy candidate before approval, bind review to live policy/provider/credential inputs, and preserve inspected endpoint contracts during mechanistic expansion.

Closes #2821

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* fix(policy): canonicalize advisor review inputs

Serialize nested protobuf maps in stable key order for proposal review tokens and effective-policy hashes. Narrow reused multi-port endpoint contracts to the denied port so advisor proposals cannot widen binary access. Add regressions for both cases.

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* test(e2e): keep advisor sandbox running

Create the issue 2821 regression sandbox detached with a durable canonical main process so policy denial, approval, and hot-reload checks run before lifecycle exit.

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* fix(policy): apply reviewed draft batches atomically

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <johntmyers@users.noreply.github.com>
2026-08-21 19:05:53 +00:00
Artem Lytvyn 7adc05af7a feat(supervisor): expose sandbox name to middleware request context (#2771)
* feat(supervisor): expose sandbox name to middleware request context

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* feat(supervisor): add workspace to middleware request context

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

---------

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>
2026-08-21 18:50:13 +00:00
Evan Lezar 3be2cd8a29 fix(helm): preflight Agent Sandbox APIs (#2867)
* fix(helm): preflight Agent Sandbox APIs

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(kubernetes): share Agent Sandbox setup

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(e2e): wait for Agent Sandbox CRD status

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* ci(canary): sparse-checkout sandbox helper

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-08-21 14:32:19 +00:00
Drew Newberry ef296806f5 feat(sandbox): add canonical main process (#2726)
* feat(sandbox): add canonical main process

Closes #2710

Persist and supervise one canonical workload per sandbox, attach sandbox connect to its retained session, and make every unexpected main-process exit terminal.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(sandbox): simplify canonical main process contract

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): preserve legacy VM main compatibility

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): preserve main status across driver updates

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): satisfy macOS process lint

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): gate Linux exit acknowledgement publisher

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(sandbox): simplify main process plumbing

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(supervisor): make controlling tty ioctl portable

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(supervisor): initialize canonical process environment

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* perf(supervisor): optimize retained main session

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(sandbox): detach main session on ctrl-c

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): use explicit main detach keys

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(dev): atomically stage Docker supervisor

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(test): align Docker main environment assertion

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(sdk): expose canonical main process fields

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-08-20 22:18:23 +00:00
John T. Myers b2ea81822b feat(network): enable Docker and Podman policy DNS and transparent TCP (#2723)
* feat(network): enable Docker transparent TCP egress

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(e2e): cover Docker transparent TCP egress

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* feat(network): correlate transparent TCP audit events

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(examples): add transparent TCP Redis demo

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(examples): demonstrate blocked TCP connections

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(examples): focus Redis demo audit output

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): close transparent TCP policy bypasses

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(sandbox): reject unsupported TCP policy reloads

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(ci): satisfy Linux transparent TCP lints

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* feat(podman): enable transparent TCP egress

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(podman): permit policy DNS port binding

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(e2e): use qualified transparent TCP hostname

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(podman): preserve exact policy DNS names

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(podman): route policy DNS over TCP

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(dns): serve multiple TCP queries per connection

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(network): explain native DNS and TCP egress

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(sandbox): reconcile runtime reload with upstream

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): harden transparent DNS capture

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(podman): preserve resolver behavior for native tcp

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(network): clarify native tcp runtime constraints

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): remove unused transparent tcp pin

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): admit redirected transparent tcp

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): restore podman transparent networking

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(podman): permit alpine busybox binaries

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(podman): use portable alpine keepalive

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(podman): build musl networking fixture

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(podman): isolate musl DNS probe

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(podman): keep privileged port capability dropped

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): preserve transparent TCP port 53

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): report synthetic pool pressure by family

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(podman): bind tcp fixtures before readiness

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(podman): grant fixture low-port bind

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-08-20 19:26:53 +00:00
John T. Myers 4d7f402ce2 feat(policy): establish direct TCP egress foundation (#2711)
* feat(policy): accept explicit tcp endpoint protocol

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* refactor(network): snapshot authoritative egress decisions

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(policy): document explicit tcp protocol

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(policy): defer transparent TCP release guidance

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* chore(go): regenerate sandbox protobuf bindings

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): complete tcp egress foundation

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(policy): document explicit tcp contract

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): fail closed on authorization errors

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(podman): fence delayed exit events before restart

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(policy): validate network endpoint destinations

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(providers): opt in tcp credential fixture

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(policy): require dns host for transparent tcp

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-08-20 16:43:09 +00:00
Drew Newberry 4c5fce6e59 refactor(compute): support external driver parity (#2744)
* refactor(compute): negotiate external driver behavior

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(compute): revert external driver documentation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(compute): negotiate gateway-managed lifecycle

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(compute): remove driver feature negotiation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(compute): let drivers declare gateway lifecycle

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(compute): clarify lifecycle ownership

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-08-20 15:56:29 +00:00
Drew Newberry 7909fb5d0f refactor(compute): unify gateway restart reconciliation (#2743)
* refactor(compute): unify gateway restart reconciliation

Remove the Docker-specific gateway shutdown cleanup and reconcile persisted running intent through ComputeDriver::StartSandbox for Docker, Podman, and VM drivers. Explicitly stopped sandboxes remain stopped.

Refs #2417

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(compute): stop local sandboxes on shutdown

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(e2e): match managed Podman containers

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(compute): synchronize lifecycle sweeps

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-08-20 10:12:41 +00:00
Mrunal Patel 6e90f3d5a0 feat(providers): store refresh credentials in credential drivers (#2801)
* feat(providers): store refresh credentials in credential drivers

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(providers): harden refresh credential lifecycle

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(providers): migrate legacy refresh secrets before skip

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* refactor(providers): defer credential migration

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(providers): make refresh configuration atomic

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

---------

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
2026-08-19 21:37:24 +00:00
Drew Newberry 998db04780 feat(policy): allow non-root sandbox identities (#2785)
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-08-19 21:20:14 +00:00
Jesse Jaggars 2eb0880a00 feat(cli): support OIDC device authorization grant for headless login (#2795)
* feat(cli): support OIDC device authorization grant for headless login

Closes #2793

Add OAuth 2.0 Device Authorization Grant (RFC 8628) support to the OpenShell CLI's OIDC login flow. When running in a headless environment (OPENSHELL_NO_BROWSER=1) without a client secret configured, the CLI now uses the device code flow instead of the browser-based PKCE flow.

The device code flow:
- Requests a device code and user code from the IdP's device authorization endpoint
- Displays a verification URL and user code to the user
- Polls the token endpoint until the user completes authorization or the code expires
- Supports slow_down responses per RFC 8628 by increasing the polling interval

This implementation:
- Extends OidcDiscovery to optionally capture device_authorization_endpoint
- Adds oidc_device_code_flow function with proper error handling for all RFC 8628 error codes
- Updates gateway add and gateway login to dispatch to device flow when browser is suppressed
- Adds comprehensive unit tests for device flow structs and response parsing
- Updates gateway authentication documentation to describe the device code fallback

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(cli): validate OIDC device token responses

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(cli): add PKCE to OIDC device flow

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* docs(cli): document PKCE device flow

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

---------

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>
2026-08-19 17:44:46 +00:00
grs 3a16012dbe fix(cli): prompt for fresh OIDC login after logout (#2773)
Signed-off-by: Gordon Sim <gsim@redhat.com>
2026-08-19 15:46:27 +00:00
alangou 0d708d6d51 fix(policy): gate uninspected credentialed endpoints (#2493)
* fix(policy): gate uninspected credentialed endpoints

Signed-off-by: Adrien Langou <alangou@nvidia.com>

* refactor(cli): extract allowed-ip option parsing

Signed-off-by: Adrien Langou <alangou@nvidia.com>

* fix(policy): gate endpointless credential bindings

Signed-off-by: Adrien Langou <alangou@nvidia.com>

---------

Signed-off-by: Adrien Langou <alangou@nvidia.com>
2026-08-19 15:18:02 +00:00
Mrunal Patel 8d67250a5d fix(providers): keep refresh credential handles stable (#2780)
* fix(providers): keep refresh credential handles stable

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(providers): protect refresh-owned credentials

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

---------

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
2026-08-18 20:45:36 +00:00
Adel ZaaloukandClaude Opus 4.6 600bbae845 feat(ocsf): emit AI inference events via ai_operation profile on ApiActivity (#2664)
Apply the official OCSF ai_operation profile (introduced in v1.8.0) to
ApiActivity [6003] events when the inference proxy routes a model call
through inference.local. Attaches an ai_model object (name, ai_provider)
and puts token counts and latency in unmapped fields.

ApiActivity [6003] is the schema-correct class for the ai_operation
profile in v1.8.0 (HttpActivity only gets it in v1.9.0). In Splunk CIM,
ApiActivity maps to the "Change" data model, naturally separating
inference events from regular HTTP proxy traffic.

Changes:
- Add AiModel object and ai_model field on BaseEventData
- Add ApiActivityEvent struct and ApiActivityBuilder
- Add emit_ai_inference in proxy.rs using ApiActivity with ai_operation
- Vendor OCSF v1.8.0 schemas including api_activity class, ai_model
  object, and ai_operation profile definitions
- Bump OCSF_VERSION to 1.8.0
- Update schema validation to skip profile-gated required fields

Shorthand: API:INFERENCE [INFO] claude-3-haiku via anthropic 701ms [POST /v1/messages]

Splunk/SIEM backward compatibility (v1.1/v1.3 CIM mapping) is tracked
separately in #2662 as a configurable serialization concern.

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-08-18 16:30:03 +00:00
Giuseppe Scrivano d51a653f9c feat(driver-podman): add userns config (#2562)
* refactor(driver): extract shared supervisor binary helpers

Move supervisor binary extraction, caching, and validation helpers from
the Docker driver into openshell-core::driver_utils so both Docker and
Podman drivers can reuse them.

Moved helpers: extract_first_tar_entry, write_cache_binary_atomic,
supervisor_cache_path, temp_extract_container_name, and
validate_linux_elf_binary.

The shared extract_first_tar_entry gains entry-type and empty-payload
checks that the Docker-local version lacked.  supervisor_cache_path
takes a driver_subdir parameter so each driver caches under its own
namespace (docker-supervisor vs podman-supervisor).

Signed-off-by: Giuseppe Scrivano <gscrivan@redhat.com>

* feat(driver-podman): add userns config

Add a `userns` option to the Podman compute driver that maps to
Podman's user namespace modes.  The mode string is split on the first
colon into the API's `nsmode` and `value` fields so parameterized
values like `auto:size=65536` and `keep-id:uid=1000,gid=1000` are
forwarded correctly.  When the mode is `auto`, the container spec
also sets `idmappings.AutoUserNs = true` as required by the API.

An allowlist validates the mode at startup: `auto` and `keep-id`
accept optional parameters; `host`, `private`, and `nomap` reject
them; everything else is an error.

Podman image volumes use overlay mounts internally and the kernel
does not support idmapped mounts on overlay (`mount_setattr` returns
EINVAL).  When userns is configured (any mode except `host`), the
driver extracts the supervisor binary from the image to a host-side
cache and bind-mounts it instead of using an image volume.

Configurable via TOML `userns = "auto"`, CLI `--userns`, or
environment variable `OPENSHELL_PODMAN_USERNS`.

Signed-off-by: Giuseppe Scrivano <gscrivan@redhat.com>

---------

Signed-off-by: Giuseppe Scrivano <gscrivan@redhat.com>
2026-08-15 17:10:47 +00:00
Piotr Mlocek 44bf0df485 feat(middleware): inspect WebSocket text messages (#2477)
* feat(middleware): inspect websocket text messages

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): address websocket review feedback

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(network): bound websocket message assembly

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(network): harden websocket upgrade lifecycle

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* refactor(middleware): unify in-process and remote transports

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* feat(middleware): support regex websocket redaction

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): bound persistent streaming sessions

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): accept websocket sequence gaps

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* refactor(middleware): refine websocket introspection contract

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): clarify websocket preflight lifecycle

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): clarify websocket coverage semantics

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(network): type websocket frame failures

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(network): return 503 when middleware admission is exhausted

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): align streaming API contract

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): clarify WebSocket event result scope

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(rfc): simplify middleware revision history

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* feat(examples): add WebSocket content guard support

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): unify binding payload limits

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* refactor(middleware): align payload limit terminology

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): address websocket review feedback

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(network): address websocket review findings

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* test(network): allow Linux handler setup in preflight regression

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(network): harden websocket relay finalization

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(network): inspect compressed websocket messages

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* test(network): stabilize compressed websocket regressions

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(go-sdk): regenerate middleware protobuf binding

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): clarify websocket skip lifecycle

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): address WebSocket review feedback

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

---------

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-08-14 21:51:42 +00:00
Derek Carr 59479f492a feat(k8s): add namespace-per-workspace support (RFC 0011 Phase 3) (#2656)
* feat(k8s): add namespace-per-workspace support (RFC 0011 Phase 3)

Implement three workspace namespace modes for the Kubernetes compute
driver: shared (default, preserves current single-namespace behavior),
managed (auto-creates/deletes namespaces per workspace), and operator
(pre-provisioned namespaces with dynamic discovery via label selector
or drop-in allowlist file).

Key changes:
- WorkspaceMode enum and namespace resolution in driver config
- Managed namespace lifecycle with ServiceAccount and OpenShift SCC
  annotation propagation
- Cluster-wide sandbox CR watchers for managed/operator modes
- NamespaceValidator (Exact/Prefix/Allowlist) for SA token auth
- Workspace-aware credential secret storage
- Helm ClusterRole for multi-namespace RBAC
- Gateway config, architecture, and reference docs

Signed-off-by: Derek Carr <decarr@redhat.com>

* test(k8s): add e2e tests for workspace namespace modes

Add end-to-end tests for managed and operator workspace modes
introduced in RFC 0011 Phase 3. The managed mode tests verify
namespace creation with correct labels, ServiceAccount provisioning,
sandbox CR placement, and namespace survival with remaining sandboxes.
The operator mode tests verify rejection of unlabeled and nonexistent
namespaces. The positive operator path (sandbox in labeled namespace)
is known to fail due to an RBAC gap and will be addressed separately.

Also fixes Helm 4 compatibility: move SPDX license headers inside
conditional guards in 8 chart templates to prevent empty comment-only
documents, and fix a trailing whitespace trimmer in clusterrole.yaml
that concatenated the license header with apiVersion.

Adds cleanup sweep in with-kube-gateway.sh to remove managed and
operator namespaces before Helm uninstall, and mise tasks for running
each mode independently.

Signed-off-by: Derek Carr <decarr@redhat.com>

* feat(k8s): add operator namespace label watcher

Spawn a background kube::runtime::watcher in the K8s driver that
watches namespaces matching the configured label selector and populates
the OperatorNamespaceAllowlist at runtime. The driver owns the
allowlist and exposes its Arc so the server can share the same set with
the SA token authenticator.

create_sandbox now gates pod creation on the allowlist in operator
mode — workspaces whose namespace is not yet labeled are rejected at
resource render time rather than silently proceeding. Workspace
lifecycle itself is unaffected; only sandbox (resource) creation is
gated.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(k8s): harden operator mode and address review findings

Close the fail-open gap in operator mode when only
operator_namespace_file is configured: the allowlist is now created
unconditionally in operator mode (fail-closed from startup).

Implement the namespace file watcher using the notify crate, following
the TLS hot-reload pattern (parent-directory watch, 1s debounce,
ConfigMap symlink-swap safe). The file format is a JSON array of
namespace name strings.

Additional fixes from the 10-reviewer audit:
- Change allowlist rejection from InvalidArgument to FailedPrecondition
  so callers know the request may succeed later once the namespace is
  provisioned.
- NamespaceValidator::Allowlist now holds the OperatorNamespaceAllowlist
  newtype instead of a raw Arc<RwLock<BTreeSet>>, eliminating silent
  denial on RwLock poison.
- Verify LABEL_MANAGED_BY and LABEL_GATEWAY_ID ownership before
  deleting a managed namespace.
- Replace fixed 5s sleep in operator e2e test with a 30s poll loop.
- Add Helm validation for workspaceMode values.
- Fix Helm README type column and description for operator fields.
- Add insert/remove methods to OperatorNamespaceAllowlist; label
  watcher now uses them instead of reaching through shared().
- Reject configs with both operator_namespace_label and
  operator_namespace_file set.

Signed-off-by: Derek Carr <decarr@redhat.com>

* feat(k8s): add workspace-level compute driver RPCs and harden RBAC

Decouple namespace lifecycle from sandbox lifecycle by adding
EnsureWorkspace/DeleteWorkspace RPCs to the ComputeDriver service.
Namespace creation now happens before credential storage and namespace
deletion happens on workspace delete, fixing credential storage in
managed workspace mode.

- Add EnsureWorkspace and DeleteWorkspace proto RPCs with
  implementations across all compute drivers (K8s managed delegates to
  ensure_namespace/delete_namespace_if_empty; others no-op)
- Wire ensure_workspace into provider create/update/refresh paths so
  the namespace exists before the credential driver writes secrets
- Wire delete_workspace into workspace deletion for cleanup
- Remove delete_namespace_if_empty from sandbox deletion path
- Scope ClusterRole secrets access to non-shared workspace modes
- Add TODO for TLS cert hot-reload in sandbox gRPC client
- Harden e2e tests with control-plane sandbox resolution assertions
- Fix docker image save --platform flag for OCI index manifests

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(k8s): address re-review findings and add test coverage

- Use server-side apply for TLS secret sync (fixes second sandbox
  creation failure when TLS is enabled)
- Scope gateway-ID label selector unconditionally across all workspace
  modes (fixes operator reads/watches/deletes seeing foreign sandboxes)
- Validate operator allowlist in EnsureWorkspace and DeleteWorkspace
  RPCs (prevents credential writes to namespaces outside the allowlist)
- Extend ClusterRole secrets patch+delete to all non-shared modes with
  credential driver enabled (fixes operator credential storage RBAC)
- Validate namespace ownership on 409 conflict in ensure_namespace
  (prevents adopting unowned namespaces in managed mode)
- Replace delete_namespace_if_empty with unconditional delete_namespace
  letting Kubernetes cascade cleanup (fixes stuck terminating CRs)
- Strengthen NetworkPolicy TODO to cover both managed and operator modes
- Extract selector and ownership logic into testable free functions
- Add unit tests for gateway-ID selectors and namespace ownership
- Add Helm ClusterRole RBAC tests for operator credential driver

Signed-off-by: Derek Carr <decarr@redhat.com>

* ci(k8s): add workspace managed and operator mode e2e to CI

Wire the existing e2e:kubernetes:workspace-managed and
e2e:kubernetes:workspace-operator mise tasks into the branch-e2e
workflow so they run alongside the other core Kubernetes e2e suites.
Both are gated by run_core_e2e and included in the Core E2E result
gate.

Signed-off-by: Derek Carr <decarr@redhat.com>

* test(k8s): add e2e tests for workspace namespace modes

Add 7 new e2e tests covering workspace namespace lifecycle, TLS secret
copying, ownership conflict detection, DNS-1123 validation, operator
namespace preservation, and dynamic label watcher behavior. Fix async
sandbox deletion race condition in existing tests by polling sandbox
list instead of asserting immediately after delete.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(k8s): grant secrets/patch unconditionally and backfill gateway-id labels

Address two review findings:

1. RBAC: server-side apply (PATCH) is used for TLS secret sync in
   multi-namespace modes, but the ClusterRole only granted patch when
   the kubernetes-secrets credential driver was enabled. Grant patch
   unconditionally for non-shared modes since TLS sync always needs it;
   keep delete gated on the credential driver.

2. Upgrade safety: the new gateway-id label selector would orphan
   legacy Sandbox CRs that predate its introduction. Add a startup
   backfill in shared mode that patches any managed Sandbox CR missing
   the gateway-id label before the driver begins serving requests.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(k8s): address workspace namespace review findings

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(k8s): address follow-up review findings

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(k8s): preserve workspace lookup after rebase

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(helm): allow managed secret creation

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(k8s): stop pods in workspace namespace

Signed-off-by: Derek Carr <decarr@redhat.com>

* test(k8s): scope pod deletion check to v1alpha1

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(k8s): address workspace namespace review findings

Signed-off-by: Derek Carr <decarr@redhat.com>

---------

Signed-off-by: Derek Carr <decarr@redhat.com>
2026-08-14 21:22:41 +00:00
Piotr Mlocek bdabb54cb3 fix(security): authenticate extension services (#2638)
* fix(security): authenticate extension services

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* feat(extension-core): verify gateway JWTs

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* refactor(extension-core): keep inbound verification external

This should become an extension SDK package.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(security): harden the extension authentication contract

Follow-up hardening on the alpha extension authentication mechanism.

Claim contract:
- Extension tokens carry an explicit `typ` of `openshell-ext+jwt`. They
  share a signing key with sandbox-to-gateway admission tokens and were
  otherwise separated by audience alone, so a verifier that neglects to
  check `aud` could accept a gateway credential. The header is a second,
  independent discriminator.
- Publish OIDC-shaped discovery at `/.well-known/openid-configuration`
  so a service configured with only the gateway URL can learn the exact
  expected issuer and the JWKS location. It is shaped, not compliant:
  `issuer` is the gateway identity, not the serving URL.

Audience agreement:
- `MiddlewareManifest` and `InterceptorManifest` gain `expected_audience`.
  The audience is otherwise configured independently on each side of the
  boundary, where a mismatch surfaces only as an opaque authentication
  failure on every call. OpenShell now compares the two and fails at
  startup. An empty field keeps the check off for existing services.

Compatibility:
- Add `allow_insecure_transport` per registration. Enabling gateway JWT
  signing previously made any plaintext endpoint a hard startup failure,
  including the endpoint form used in our own documentation. The opt-out
  attaches no credential, is refused by the gateway if a supervisor asks
  for one, and warns at every startup.
- Make the transport requirement kind-aware. A middleware endpoint must
  be reachable from every sandbox supervisor, so only interceptors may
  use a gateway-local Unix socket.

Credential lifecycle:
- Replace the process-global slot map with a supervisor-owned
  `ExtensionCredentialStore` shared explicitly across the gateway
  connections the supervisor opens, removing test-order coupling.
- Rotate only when a credential is missing or has passed four fifths of
  its lifetime. Configuration polling ran every ten seconds against
  fifteen-minute credentials, so each poll re-ran gateway effective-policy
  resolution and re-minted the gateway token.
- Bound credential minting per sandbox, since each request resolves the
  caller's effective policy.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs: record alpha extension authentication in RFC appendices

Restore the RFC 0009 and 0010 bodies to their accepted text and move
every extension-authentication update into appendices instead. An RFC
records a decision at a point in time; superseding detail belongs
alongside it rather than rewritten into it.

RFC 0009's appendix carries the shared contract: claims, authorization,
key distribution, the `allow_insecure_transport` replacement for the
body's `allow_insecure`, and residual risks. RFC 0010's records only
what differs for interceptors and links to it. The existing
protocol-extensions appendix, which parked the phase 2 transport
question, now points forward to what was built.

Also document the audience handshake, the discovery endpoint, the
`typ` requirement, and `jti` replay guidance in the extensibility and
gateway configuration pages, and correct the middleware transport
guidance: middleware endpoints must be reachable from sandbox
supervisors, so Unix sockets are not an option there.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* refactor(extension-core): abstract extension server trust

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(core): update middleware manifest example

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(extension-auth): preserve unsigned gateway compatibility

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(extension-auth): reject cross-domain token replay

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

---------

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-08-14 20:34:10 +00:00
Drew Newberry f12f3ef8d5 fix(macos): restore Homebrew sandbox callbacks (#2739)
* fix(macos): restore Docker gateway callbacks

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(gateway): reuse reachable primary callback listener

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-08-14 18:38:20 +00:00
LR90 d0c6dc3fd8 feat(kubernetes): support corporate upstream proxy (#2633)
* feat(kubernetes): support corporate upstream proxy

Signed-off-by: loveRhythm1990 <qiuweimin@126.com>

* fix(kubernetes): reject proxy_auth_secret_key values Kubernetes cannot create

Gateway validation accepted proxy_auth_secret_key values that Kubernetes
rejects when creating the Secret (keys longer than 253 bytes, or the
reserved "."/".." names), turning an invalid deployment setting into
repeated sandbox Pod-provisioning failures instead of a startup error.
Reject them in validate_upstream_proxy_config so they fail closed at
gateway startup.

Signed-off-by: loveRhythm1990 <qiuweimin@126.com>

* docs(skill): add corporate upstream proxy checks to debug-openshell-cluster

Add a Kubernetes corporate upstream proxy troubleshooting section covering
rendered [openshell.drivers.kubernetes] configuration, credential Secret
volume events, supervisor arguments and mounts confined to the network-
supervising container, and proxy reachability.

Signed-off-by: loveRhythm1990 <qiuweimin@126.com>

---------

Signed-off-by: loveRhythm1990 <qiuweimin@126.com>
2026-08-14 16:16:43 +00:00
Jesse Jaggars c4b500a7de feat(helm): cert-manager external issuer + OpenShift passthrough Route (#2468)
* fix(core): trust public root CAs alongside the sandbox mTLS CA

The supervisor gRPC client only trusted the CA configured via
OPENSHELL_TLS_CA, since tonic ClientTlsConfig starts with an empty root
store unless with_native_roots()/with_webpki_roots() is also enabled.
Deployments where the gateway server certificate is issued by a public CA
(e.g. cert-manager against an ACME issuer) caused every supervisor
connection to fail the TLS handshake with "UnknownCA", since the sandbox
mTLS CA and the server cert issuer were no longer the same.

Enable both native and webpki roots in addition to the configured CA.
tonic root store is a union of all configured sources, so this does not
weaken verification for existing self-signed deployments. webpki-roots
(compiled in) is enabled alongside native-roots since the supervisor
binary may run in minimal sandbox images without a populated system CA
bundle.

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* feat(helm): support external cert-manager issuers and OpenShift Route passthrough

Add certManager.serverIssuerRef/clientIssuerRef so the gateway and mTLS
client certificates can be issued by a real Issuer/ClusterIssuer (e.g.
ACME) instead of only the chart built-in self-signed CA.

Add openshiftRoute template for exposing the gateway via a TLS
passthrough Route so the gateway keeps terminating its own TLS/mTLS.

The server Certificate excludes internal-only SANs (cluster-local,
localhost, loopback) when an external issuer is configured, since ACME
issuers reject those per CA/Browser Forum baseline requirements. A
template-time fail guard catches the misconfiguration at helm install
time rather than asynchronously at cert-manager issuance time.

Includes Helm unittest coverage for both issuerRef overrides and Route
rendering, plus a CI values overlay for lint coverage.

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* docs: document cert-manager external issuer and OpenShift Route

Update managing-certificates.mdx with the serverIssuerRef workflow and
install-time validation behavior. Add a production section to the
OpenShift guide covering passthrough Route with a real certificate.
Regenerate Helm README for new certManager and openshiftRoute values.
Sync debug-openshell-cluster skill with new troubleshooting steps for
ACME issuance failures and supervisor UnknownCA from mismatched CAs.

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(helm,core): address PR review feedback on cert-manager external issuer

Addresses all five blocking review items from #2468:

1. Remove .with_native_roots() from supervisor gRPC client -- the
   supervisor runs inside the user-selected sandbox image, so the
   image CA bundle is not operator-controlled. Keep .with_webpki_roots()
   (compiled-in, not user-controlled) alongside the configured CA.

2. Fail at render time when serverIssuerRef.name is set but
   clientCaFromServerTlsSecret is still true. Add negative Helm test.

3. Remove clientIssuerRef -- changing only clientIssuerRef breaks both
   directions because trust bundles are not modeled separately. Change
   serverIssuerRef.kind default from ClusterIssuer to Issuer.

4. Add server.oidc.issuer and server.oidc.audience to the documented
   OpenShift production Helm command. Add Access Control prerequisite.

5. Fail at render time when openshiftRoute.enabled and disableTls are
   both true. Add negative Helm test.

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(drivers): strip GATEWAY_TLS_SERVER_NAME from Docker and Podman env

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(helm,drivers): guard default clientCaSecretName and add env-strip tests

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* feat(tls): SNI-based dual certificate for internal and external server TLS

Split the gateway server certificate into two: an internal cert issued by
the chart's own CA (for supervisor connections via cluster-local SANs) and
an external cert issued by an operator-configured Issuer such as ACME/Let's
Encrypt (for CLI and Route access via public SANs).

The gateway uses SNI-based certificate selection: connections whose SNI
hostname matches external_server_names receive the external cert; all
others (including those with no SNI) receive the internal cert.

Security improvement: remove .with_webpki_roots() from the supervisor
gRPC client so supervisors trust only the chart CA, closing a MITM vector
via publicly-trusted certificates in user-supplied container images.

Key changes:
- Add DualCertResolver with SNI-based cert selection and full test coverage
- Add external_cert_path, external_key_path, external_server_names to TlsConfig
- Validate partial external cert config (error on cert-without-key or vice versa)
- Validate empty external_server_names when external cert is configured
- Split cert-manager templates into internal + external Certificate resources
- Add Helm guards for misconfigured external issuer (empty serverDnsNames,
  internal-only SANs with external issuer, conflicting clientCaFromServerTlsSecret)
- Update gateway-config.mdx, managing-certificates.mdx, openshift.mdx docs
- Update debug-openshell-cluster skill for dual-cert troubleshooting

Signed-off-by: Pi Agent <agent@openshell.local>

* fix(drivers): strip GATEWAY_TLS_SERVER_NAME in VM driver and correct comments

Add the same GATEWAY_TLS_SERVER_NAME environment stripping to the VM
compute driver that Docker, Podman, and Kubernetes drivers already
perform. Without this, a sandbox user on the VM driver could override
the TLS server name the supervisor verifies.

Fix stale comments in Docker and Podman drivers that referenced
'with WebPKI roots trusted' — WebPKI roots are explicitly not trusted
after the tls-webpki-roots removal.

Use tls-ring instead of bare channel for tonic in openshell-core so the
TLS API (ClientTlsConfig, Endpoint::tls_config) is available without
pulling in any root certificate store.

Signed-off-by: Pi Agent <agent@openshell.local>

* fix(tls,helm): wildcard SNI matching and Route host validation

Add RFC 6125 single-level wildcard matching to DualCertResolver so
external_server_names entries like *.example.com correctly match SNI
hostnames like gw.example.com. Previously only exact matches worked,
silently falling back to the internal cert for wildcard configurations.

Add a Helm fail guard in route.yaml that rejects openshiftRoute.host
values not listed in certManager.serverDnsNames when an external issuer
is configured — catches cert/route hostname mismatches at install time
instead of at TLS connect time.

Quote the host field in route.yaml for robustness.

Signed-off-by: Pi Agent <agent@openshell.local>

* fix(helm): address blocking review items — client-CA guard, wildcard Route, serverIssuerRef gate

1. Remove the obsolete guard rejecting serverIssuerRef + clientCaFromServerTlsSecret=true.
   The internal server certificate is always signed by the chart CA (the same
   CA that signs the client cert), so clientCaFromServerTlsSecret=true is
   correct — its filtered ca.crt is exactly the right trust anchor.  The old
   workaround (mounting openshell-ca-tls directly) unnecessarily exposed the
   CA private key to the gateway container.  Remove the client-CA overrides
   from docs, CI overlay, and production examples.

2. Route host validation now supports wildcard certificates per RFC 6125:
   single-level wildcards like *.example.com match gateway.example.com but
   not deep.sub.example.com.  Require an explicit openshiftRoute.host when
   an external issuer is configured — without one, OpenShift generates a
   hostname absent from serverDnsNames.

3. Reject serverIssuerRef.name when certManager.enabled is false — the
   external certificate, its Secret mount, and the gateway TLS config all
   require cert-manager to be enabled.

Validated on ROSA (dev.dyee.p3) with branch-built images:
- Fresh install with letsencrypt-prod ClusterIssuer
- SNI dual-cert: external hostname served Let's Encrypt cert
- Supervisor mTLS via internal cert path: ConnectSupervisor accepted
- Client CA volume: filtered ca.crt from internal server secret (no key)
- CLI connected via Route + OIDC

Helm tests: 81 pass across 7 suites.

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

---------

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>
Signed-off-by: Pi Agent <agent@openshell.local>
2026-08-14 00:12:11 +00:00
Max DubrinskyandDrew Newberry 35fb27ef14 feat(sdk): add TypeScript SDK (@nvidia/openshell-sdk) (#2122)
* feat(sdk): add TypeScript SDK (@nvidia/openshell-sdk)

First native, per-language SDK for the OpenShell gateway: a thin, idiomatic
TypeScript client over proto-generated gRPC stubs (connect-es), no FFI. Covers
the v0.1 surface — sandbox lifecycle (create/get/list/delete + waitReady/
waitDeleted), health, and streamed exec.

- sdk/typescript/: package, client/transport/errors, protoc + protoc-gen-es
  codegen (gen/ gitignored, absorbed into dist/ at build), committed lockfile.
- tasks/typescript.toml: sdk:ts install/proto/typecheck/build/ci/publish;
  sdk:ts:typecheck wired into `check`; sdk-typescript job in branch-checks
  (typecheck, build, and a --dry-run publish that validates the release path).
- Enforce SPDX headers on .ts/.tsx/.mts/.cts (skip node_modules and gen/);
  back-fill docs/_components/jsx.d.ts and fern/components/CustomFooter.tsx.
- release.py gains an npm version format; release-tag.yml publishes to
  GitHub Packages on tag, stamping the version (0.0.0 placeholder in git);
  prerelease builds publish under the `next` dist-tag, not `latest`.

Ships as @nvidia/openshell-sdk on GitHub Packages pre-GA; public npm
(@openshell/sdk) follows at GA with an unchanged public API.

Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com>

* chore(sdk): adopt TypeScript 6, tidy @types/node range

- typescript ^5.7.2 -> ^6.0.3 (6.0 is now `latest`; the old caret capped at 5.x)
- @types/node ^24.0.0 -> ^24 (same range, tidier)

No source changes; codegen, typecheck, and build pass on 6.0.3. Verified the
emitted d.ts still type-check for downstream consumers on TypeScript 5.0.4
through 5.9.3, so this does not raise the SDK's consumer TS floor.

Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com>

* refactor(sdk): group operations under a composable SandboxClient

Reshape the client from flat methods (createSandbox, listSandboxes, exec) to a
scoped SandboxClient reached as `client.sandbox.create/get/list/delete/exec`
(+ waitReady/waitDeleted), mirroring the CLI's noun-verb model and the Python
SDK's SandboxClient.

SandboxClient is also usable standalone via SandboxClient.connect();
OpenShellClient composes it over a single shared transport, so future
service/provider clients reuse one connection. health() stays top-level as a
gateway call. No behavior change; types are unchanged.

Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com>

* chore(sdk): generate TypeScript SDK stubs with buf

Replace the protoc gen.sh with `buf generate` + buf.gen.yaml. `buf`
(@bufbuild/buf) is a package devDependency and self-compiles the protos, so
the TS SDK no longer depends on the mise-pinned protoc; it drives the same
connect-es plugin. Generation stays limited to the client-surface closure
(openshell/sandbox/datamodel) via the input paths.

Output is byte-identical to the previous protoc + protoc-gen-es pipeline. Lays
the groundwork for a shared buf.yaml (lint/breaking/LSP) as a follow-up.

Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com>

* build(proto): add repo-level buf module with lint

Declare proto/ as a single buf v2 module in a root buf.yaml so buf
generate, lint, breaking, and the editor LSP resolve imports the same
way. Lint uses STANDARD with six documented exceptions for deviations
the current protos intentionally make: the flat proto/ layout with
nested packages (DIRECTORY_SAME_PACKAGE, PACKAGE_DIRECTORY_MATCH) and
the established API shape with unsuffixed services and reused
request/response messages (RPC_REQUEST_RESPONSE_UNIQUE,
RPC_REQUEST_STANDARD_NAME, RPC_RESPONSE_STANDARD_NAME, SERVICE_SUFFIX).
Every other STANDARD rule now enforces on future protos. Breaking uses
FILE.

Code generation stays package-scoped in sdk/typescript/buf.gen.yaml
since it binds to that package's connect-es plugin and output dir;
its inputs are unchanged and regeneration is byte-identical.

Wire the check in via a proto:lint mise task that runs buf from the
SDK devDependencies. It is a dependency of both sdk:ts:ci (so the
TypeScript SDK CI job enforces it) and the top-level lint aggregate
(so local pre-commit covers it).

Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com>

* chore(sdk): publish as unscoped openshell-sdk on public npm

Rename the package from @nvidia/openshell-sdk to the unscoped
openshell-sdk and target public npm (registry.npmjs.org) instead of
GitHub Packages. GitHub Packages requires a scope matching the owning
org, and the @openshell scope is blocked by an unrelated existing
package, so an unscoped name on public npm is the lowest-friction
distribution path and needs no org approval.

Rework the release-tag publish job to auth against registry.npmjs.org
with NPM_TOKEN (the job now only needs packages: read to pull the CI
image). Update the README install instructions and usage imports.

Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com>

* chore(sdk): publish @nvidia/openshell-sdk to GitHub Packages

Revert the unscoped-name switch. GitHub Packages only accepts scoped
names matching the owning org, so shipping there first (which needs no
external npm org or NPM_TOKEN, just the repo's GITHUB_TOKEN) requires
the @nvidia scope. Keeping the @nvidia/openshell-sdk name also lets a
later public-npm release use the same install specifier, so adding
public npm becomes a second publish step rather than a rename.

Restore the GitHub Packages publish auth in the release-tag job and the
scoped install instructions in the README (keeping the buf codegen
note).

Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com>

* feat(sdk-ts): add streaming exec, forward, ssh, provider, and config methods

Grow SandboxClient to the surface the first two consumers need. execStream
yields stdout/stderr chunks as they arrive and exec now drains it, keeping its
buffered ExecResult and signature unchanged. execInteractive is the TTY + stdin
transport primitive (start-first framing, output/write/resize/close/done, no
terminal glue). forward binds a local TCP listener that tunnels each accepted
connection into the sandbox for the process lifetime, minting and revoking a
per-socket SSH session token around a forwardTcp bidi. Adds createSshSession /
revokeSshSession, attach/detach/listProviders, and getConfig / setPolicy /
setSetting (sandbox-scoped, network-policy-only, with an optional wait poll).

Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com>

* build(sdk-ts): add Biome and Vitest tooling

The TypeScript SDK had no formatter or linter and no test runner. Add Biome
(format + lint, generated src/gen excluded) enforcing 2-space indent, single
quotes, semicolons, and a 120-column width, and reformat the existing
hand-written sources accordingly. Add Vitest for unit tests. Wire sdk:ts:format,
sdk:ts:lint, and sdk:ts:test mise tasks into the fmt/lint aggregates, the root
test suite, and sdk:ts:ci so they run in CI.

Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com>

* test(sdk-ts): cover the sandbox surface with in-memory transport tests

Exercise SandboxClient against an in-memory OpenShell service built with
createRouterTransport: request assembly and id resolution, u64/int64 rendered as
strings, enum lowercasing, fromConnect code mapping, the exec/execStream drain
plus a backward-compat check on exec, execInteractive start-first ordering and
done resolution, and a forward() byte relay against a loopback echo with close()
teardown.

Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com>

* docs(sdk-ts): document the new surface and connect/upload/download boundaries

Document execStream, execInteractive, forward, ssh sessions, providers, and
config/policy in the SDK README, and record the intentional boundaries:
interactive connect / PTY ownership, upload/download (no file-transfer RPC), and
detached forwards stay out of scope. Note the Biome/Vitest dev commands.

Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com>

* feat(sdk-ts): support mTLS client authentication

Add clientCert and clientKey to ConnectOptions so the SDK can
authenticate to the default local gateway, which uses mTLS user
authentication. Without a client certificate and key the SDK could
verify the server but never authenticate the caller, so it could not
connect to the standard Docker, VM, Homebrew, or Linux-package gateway.

Validate the pair as both-or-neither and pass cert and key through to
the Node TLS options for https gateways. The h2c path is unchanged.

Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com>

* chore(sdk-ts): drop the demo script and its tsx dependency

Remove src/demo.ts, the demo npm script, the tsx devDependency, and the
tsconfig build exclude for the demo. The demo was never part of the
published package, and dropping it also removes the only place that
logged part of an SSH session token.

Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com>

* fix(sdk-ts)!: harden exec streaming, waits, SSH, and forwarding

Address review feedback on the sandbox surface.

- Make the streamed command exit code observable from idiomatic
  for-await: the terminal exit is now an in-band ExecStreamEvent
  ({ type: 'exit', exitCode }) rather than the async generator return
  value, which for-await discards. A stream that ends without an exit
  event now throws instead of reporting success.
- Bound waitReady, waitDeleted, and the setPolicy wait by their timeout:
  each poll RPC carries a per-iteration deadline and the waits accept an
  AbortSignal, so a stalled call can no longer leave a wait pending
  forever. Add waitTimeoutSecs to SetPolicyOptions.
- Validate the CreateSshSession response against the proto charset and
  range contract before returning it or using its token, since the
  values feed an OpenSSH ProxyCommand.
- Respect socket backpressure when relaying forwarded responses: pause
  reading the gRPC stream when the local socket buffer is full and
  resume on drain so memory stays bounded.
- Expose create-time sandbox policy: add policy and an advanced rawSpec
  passthrough to SandboxSpec so the safety boundary is expressible at
  creation and new spec fields do not require an SDK change.

BREAKING CHANGE: execStream and the interactive exec output now yield a
terminal { type: 'exit', exitCode } event; consumers iterating the
stream must handle that arm. The exit code is no longer the async
generator return value.

Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com>

* feat(sdk-ts): export error contract, enum unions, and caller cancellation

Address Tier-1 review feedback on the TypeScript SDK public surface (PR #2122).

- Errors: export SdkError and SdkErrorCode so callers can use instanceof and
  exhaustively switch on .code. fromConnect preserves the originating
  ConnectError as .cause and its status as .connectCode, maps Aborted to a new
  'aborted' code for optimistic-concurrency conflicts, and maps Canceled and
  DeadlineExceeded to 'canceled'. errorCode() behavior is unchanged.
- Enums: replace the string-typed phase, status, scope, and policySource fields
  with lowercase literal unions (SandboxPhaseName, HealthStatus,
  SettingScopeName, PolicySourceName) backed by exhaustive Record maps. The
  unions are a hand-maintained mirror of the generated proto enums; a new drift
  test pins each literal to its generated member name.
- Cancellation: accept an optional AbortSignal on exec, execInteractive, and
  forward, threaded into both sandbox resolution and the streaming RPC. forward
  tears down its local listener on abort.

Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com>

* feat(sdk-ts): add raw escape hatch for uncurated gateway RPCs

The curated sub-clients reduce proto messages to ergonomic subsets (for
example get() drops created_at_ms, the full spec, conditions, runtime
endpoints, and current_policy_version), and not every gateway RPC has a
typed helper yet. Rather than ship methods that exist but throw, expose a
generated client for the full surface.

OpenShellClient.raw and SandboxClient.raw are generated clients covering
every gateway RPC, returning the verbatim wire messages so proto
distinctions the curated types smooth over are preserved. .transport
exposes the shared connection for building extra clients over one socket.
Generated request/response types are published at the new
@nvidia/openshell-sdk/raw subpath. Curated methods stay the default;
raw is the always-available floor.

Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com>

* fix(sdk-ts): address review feedback on exec, forward, and auth transport

Settle exec's `done` promise before yielding the exit event so a consumer
that breaks on exit no longer leaves it pending forever, and give it a lone
rejection handler plus a finally-settle so a stream error or early abandon
can never surface as an unhandled rejection or a hang.

Attach an 'error' listener to each accepted forward socket synchronously,
before forwardConnection awaits CreateSshSession; a peer reset in that
window previously emitted an unhandled 'error' and crashed the process.

Reject ambiguous or unsafe transport configs at buildTransport: oidcToken
and edgeToken together (silently OIDC-only), and any auth token sent over
plaintext http:// to a non-loopback host unless allowInsecureAuth is set.

Wrap versionPin so a non-u64 expectedResourceVersion raises
SdkError('invalid_config') instead of a raw BigInt SyntaxError, and raise
the Node engine floor to >=20.3 for AbortSignal.any().

Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com>

* fix(sdk-ts): address review feedback

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(sdk-ts): defer published sdk guide

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
2026-08-13 20:50:38 +00:00
Seth Jennings 0f8fad23c4 feat(sandbox): add stop and start operations (#2653)
* feat(sandbox): add suspend and resume operations

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(server): preserve lifecycle work after cancellation

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(server): reconcile ambiguous lifecycle outcomes

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(server): complete suspended session cleanup

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(vm): preserve suspension state on resume failure

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(server): retry retained lifecycle transitions

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(server): clean sessions after suspend reconciliation

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* test(sandbox): cover deleting suspended sandbox

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(kubernetes): preserve progressing sandbox suspension

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(kubernetes): bound suspend status polling

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(kubernetes): detect legacy sandbox suspension

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(tui): render suspended sandbox phases

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* refactor(sandbox): rename suspend and resume lifecycle

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* perf(server): clean stopped sessions on transition

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(kubernetes): fail fast on rejected stop

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(compute): fence stale restart lifecycle events

Signed-off-by: Seth Jennings <sjenning@redhat.com>

---------

Signed-off-by: Seth Jennings <sjenning@redhat.com>
2026-08-13 06:13:54 +00:00
Artem Lytvyn dd2b4e3bc0 feat(cli): warn when --env values look like credentials (#2655)
* feat(cli): add credential env match validation

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* feat(cli): warn when --env values look like credentials

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* docs(sandbox): add flag --no-credential-warnings details + polishing

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* fix(cli): match credential keywords on underscore segments

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

---------

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>
2026-08-11 21:40:55 +00:00
Matthew Grossman 170961997f chore(ci): disable telemetry in internal test runs (#2648)
* chore(ci): disable telemetry in internal test runs

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* test(ci): remove brittle telemetry wiring test

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* docs: trim CI telemetry guidance

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* test(e2e): share telemetry default with OpenShift

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

---------

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>
2026-08-10 21:07:29 +00:00
John T. MyersandJohn Myers 0120535efc feat(proxy): bind static credentials to provider endpoints (#2510)
* feat(proxy): bind static credentials to provider endpoints

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* test(e2e): verify static credential endpoint isolation

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* docs(provider): explain static credential endpoint binding

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* fix(e2e): use valid endpoint isolation fixtures

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* docs(provider): explain static credential endpoint binding

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* fix(credentials): preserve binding identity across rotations

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(proxy): enforce bindings across request lifecycle

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(proxy): close credential relay gaps

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(credentials): clarify binding failure behavior

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(credentials): hash selected provider profile scope

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(proxy): resolve credentials after request admission

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(credentials): clarify binding failure diagnostics

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(proxy): align single-route credential denials

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(credentials): harden endpoint-bound rotation

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(credentials): enforce identity and authority binding

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(credentials): snapshot provider environment atomically

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(e2e): include authority port in query proxy requests

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(credentials): close credential revocation gaps

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(proxy): explain authority mismatch diagnostics

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(credentials): enforce binding lifecycle invariants

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(provider): reject credential config collisions

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): capture credential scope atomically

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): distinguish origin and absolute targets

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(provider): isolate endpointless profile credentials

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): normalize IPv6 request authorities

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(credentials): clarify endpointless profile isolation

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* feat(policy): bind endpointless provider credentials

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(credentials): use current GCP placeholder revision

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(providers): explain policy credential bindings

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(credentials): cover endpointless fail-closed invariant

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(policy): expect ambiguity rejection at creation

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(server): authenticate rebased policy requests

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* refactor(proxy): share credential mismatch finding builder

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(credentials): cover malformed binding metadata

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(credentials): verify multi-key endpoint isolation

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(e2e): cover same-host credential path denial

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(credentials): document serialized refresh contract

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* refactor(proxy): consolidate L7 log formatting

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* perf(credentials): precompile endpoint binding patterns

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* perf(credentials): share identity epoch revisions

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(proxy): require explicit request default ports

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(policy): validate SigV4 credential sources

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(credentials): preserve endpoint bindings for credential handles

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* feat(go-sdk): expose network credential bindings

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <johntmyers@users.noreply.github.com>
2026-08-10 19:19:45 +00:00
Shiju 5e2f0d1b37 fix(policy): prevent implicit authorization inheritance (#2499)
* fix(policy): prevent implicit authorization inheritance

A network rule authorizes every listed binary to reach every listed
endpoint, so unioning an AddRule operation's binaries and endpoints
independently grants binary-by-endpoint pairs the operation never
declared. Require AddRule to declare the complete product before
merging, and reject an operation that would give one host and port two
different MCP inspection contracts. The rejection names the binaries the
operation still has to declare.

Fix proposal coverage on the same surface. An any-binary proposal was
vacuously covered by a binary-restricted loaded rule, and a complete
product split across several loaded rules was reported as uncovered.
Coverage compares merge-widened endpoint fields by containment and
exact-matches only the fields the merge never widens, so a policy the
gateway just merged always reads back as covered and the sandbox
policy.local /wait long-poll cannot spin to its deadline.

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(policy): close authorization-inheritance gaps in merge and coverage

Coverage treated an unset proposal value for a field the merge retains as a
request for the default, so a proposal that merged cleanly into an endpoint
carrying enforcement, protocol, or tls read back as uncovered and left the
policy.local /wait long poll spinning. Unset now means unspecified.

Ports are a set on the wire but each port is an independent authorization.
Coverage and the inheritance check both resolve one binary, host, path, and
port at a time, so ports spread across loaded rules resolve and a complete
declaration split across incoming endpoints is accepted.

An incoming empty binary list means any binary. It now has to declare every
merged endpoint like a new concrete path does, and once declared the promotion
is applied instead of appending an empty list and leaving the restricted scope
in place.

The endpoint-overlap fallback folded a new rule name into an existing rule
where inheritance validation then rejected it, leaving no way to grant a binary
part of a rule. Folding now keeps the requested rule name when it would widen,
and reports that it did. An MCP contract conflict still propagates because one
host and port carry a single inspection contract.

AddAllowRules and AddDenyRules select an endpoint by host and port alone and
now reject a target that resolves more than once, including two paths on one
rule. RemoveBinary rejects an any-binary rule rather than reporting a success
that leaves the binary authorized.

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(policy): gate port and MCP-contract widening independently of binary scope

The changed-port check only rejected when the operation also left an existing
binary undeclared. An operation listing every existing binary could therefore
declare one port of a multi-port endpoint and have its widened fields land on
the endpoint the merge shares across all of them, authorizing L7 rules on a
port it never named. Each changed port must now be named by the operation on
its own, independently of binary-scope coverage.

MCP contract compatibility was checked inside the endpoint fold, which only
compares endpoints agreeing on host, path, and a shared port. The sandbox
resolves one extended configuration per host and port and never consults the
path, so a second MCP endpoint under another path, or in another rule, left the
effective strict-tool-name, method-profile, and body-limit contract decided by
match order. Contract agreement is now enforced across the whole merged policy,
including provider-composed rules. A conflict already present in the baseline
is left alone so unrelated updates still apply.

An empty binary list authorizes any binary. Appending an incoming named list
made it non-empty and revoked every process the operation did not name, turning
an additive update into a silent mass revocation. An already-empty scope is now
kept and reported; only a restricted scope is replaced by an incoming
any-binary scope.

Warnings raised during a fold now name the rule that was actually modified
rather than the rule name the operation requested, which differ when the
endpoint-overlap fallback redirects the operation.

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(policy): route undeclared-port conflicts through the separate-rule fallback

The fold-only classifier decides which merge errors disappear when the incoming
authorization stays on its own rule. UndeclaredPortWouldChange was added ahead
of the existing-binary conflict but never classified, so a differently named
narrow update against a multi-port endpoint failed outright instead of landing
separately. A same-key update still returns the error, because there the
operation chose the target.

The classifier is now an exhaustive match rather than a matches! with an
implicit false. A new variant defaulting to "not fold-only" is what withdrew the
separate-rule remedy here, so adding one has to be an explicit decision.

Inspection-contract agreement now covers protocol, not only MCP options. The
sandbox resolves one extended configuration per host and port and never
consults the path, so an MCP endpoint and a REST endpoint on the same host and
port left the effective inspection protocol decided by match order. Endpoints
with no protocol carry no contract and are skipped.

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(policy): compare only MCP contracts when detecting endpoint conflicts

The post-merge conflict scan was broadened to compare inspection protocol as
well as MCP options on one host and port. The supervisor selects among matching
endpoint configs by most-specific path, so a broad REST endpoint and a narrower
GraphQL endpoint on the same host and port are unambiguous and supported. The
broader comparison rejected those updates even though the equivalent full policy
loads and serves correctly.

MCP options are not selected that way, so the scan keeps comparing them: two MCP
endpoints on one host and port still have to agree on strict tool names, method
profile, and body limit, whatever paths or rules hold them.

The policy page returns to describing the MCP-specific rule and the path-aware
selection it sits alongside.

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(policy): keep MCP off a host and port shared with other inspection

Narrowing the conflict scan back to MCP options let an MCP endpoint sit under a
path already covered by a broader REST endpoint. The supervisor picks the parser
by most-specific path, so the MCP endpoint parses the request, but
_policy_allows_l7 is existential over every endpoint matching it. A plain REST
rule on the overlapping path can therefore make allow_request true for a
JSON-RPC tool call the MCP endpoint never allowed, and the relay forwards it.

The scan now records every inspected protocol on a host and port. MCP may not
share one with a differently inspected endpoint, and two MCP endpoints there
still have to agree on one contract. Endpoints that are not inspected carry no
contract and never compete, and two non-MCP endpoints stay supported because
they share one method-and-path rule vocabulary.

Signed-off-by: Shiju <shiju@nvidia.com>

---------

Signed-off-by: Shiju <shiju@nvidia.com>
2026-08-09 23:54:04 +00:00
Matthew Grossman 4cb77a900e fix(e2e): separate Podman Machine loopback listeners (#2622)
* fix(e2e): separate Podman Machine loopback listeners

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* test(e2e): remove shallow harness checks

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* refactor(e2e): trim Podman listener workaround

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* fix(e2e): bypass proxies for Podman health probe

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

---------

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>
2026-08-07 17:39:44 +00:00
Johnny Greco 284da54de5 docs(readme): add theme-aware banner (#2619)
* docs(readme): add theme-aware banner

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(readme): exclude preview screenshot from tree

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

---------

Signed-off-by: Johnny Greco <jogreco@nvidia.com>
2026-08-05 19:01:50 +00:00
5548405fcb feat(credentials): add provider credential storage drivers (#2437)
* feat(credentials): add provider credential storage drivers

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(credentials): harden credential update handling

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(credentials): harden credential driver security, correctness, and performance

Address review findings from the credential storage drivers PR:

- Route additional_credentials through the driver on refresh to prevent
  silent data loss for multi-credential providers (e.g. AWS STS)
- Clean up stored credential handles on CAS failure during refresh to
  prevent orphaned secrets in external backends
- Enforce namespace validation in the Kubernetes Secrets driver to
  prevent cross-namespace credential access when allow_reference_namespace
  is not enabled
- Cache Vault Kubernetes auth tokens with 80% TTL to avoid re-authenticating
  on every credential operation
- Parallelize resolve_credentials in all three drivers using try_join_all
  for faster sandbox startup
- Add existingSecret support for the KEK Secret to fix helm template/GitOps
  workflows where lookup returns empty and regenerates the key
- Document RBAC blast radius for the Kubernetes Secrets credential driver
  and recommend a dedicated namespace

Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com>
Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>

* fix(credentials): add optimistic concurrency, fix thundering herd, parallelize operations

Use resourceVersion optimistic concurrency with retry loop for K8s
Secret ownership checks to prevent TOCTOU races. Switch Vault token
cache from RwLock to Mutex with double-check pattern to prevent
thundering herd on cache miss. Parallelize credential store and delete
operations across independent keys using try_join_all.

Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com>
Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>

* fix(credentials): handle partial failures, add delete retry, consolidate cleanup

Replace try_join_all with join_all in credential store/delete operations
to handle partial failures — successfully-stored handles are cleaned up
when another key fails. Add retry loop with conflict detection to
db-credstore delete_credential, matching the K8s driver pattern.
Consolidate 4 manual cleanup_pre_stored_provider_credentials call sites
into a single error handler using an async block. Remove inconsistent
.trim() from db-credstore validate_handle_owner.

Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com>
Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>

* fix(credentials): fix retry loop guard and remove unprotected validation

Remove attempt-count guard from 409/Aborted match arms in retry loops
so the post-loop Status::aborted error is reachable after exhausting
retries. Previously, last-attempt conflicts fell through to the
catch-all error arm, producing misleading Status::unavailable errors.

Remove duplicate validation calls that ran after
prepare_provider_credential_update but outside the cleanup-protected
async block, which would leak pre-stored handles on failure.

Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com>
Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>

* fix(credentials): add workspace/provider UUID to credential backend paths

Include workspace and provider ID in credential backend object paths to ensure
cross-workspace uniqueness and prevent credential collision (GATOR-1806c9be-01).

- Updated credential driver proto to include workspace and provider_id fields
- Modified Vault driver to include workspace/provider_id in managed_secret_path
- Modified Kubernetes Secrets driver to include workspace/provider_id in
  credential_owner_id and managed_secret_name
- Updated all credential runtime calls to pass workspace/provider_id
- Updated tests to use the new signatures

This prevents two workspaces sharing the same external credential store from
colliding on provider names, which was a critical security issue (CWE-639).

* fix(credentials): preserve provider-level expiration for handle-backed credentials

Compute effective expiration from both provider and driver values using the
earliest non-zero timestamp and skip expired values before insertion
(GATOR-1806c9be-02).

- Modified resolve_provider_handles to check provider credential_expires_at_ms
- Skip expired credentials during resolution instead of returning them
- Use effective expiration (min of provider and driver) in resolution results
- Fix inference.rs to preserve earliest expiration when merging

This ensures handle-backed credentials respect the same expiration semantics
as inline credentials.

* fix(credentials): stage refresh changes under new handles before validation

Stage credential replacements under new immutable handles instead of reusing
existing handles to prevent overwriting committed values before validation/CAS
(GATOR-1806c9be-03).

- Stage credentials with empty existing_handles map to force new handle creation
- Validate and CAS before the new values are committed to backend storage
- Delete old handles only after successful CAS
- On CAS failure, delete only the newly staged handles
- This prevents CWE-362/CWE-367 race conditions where failed refreshes could
  still modify or delete the active credential

The fix ensures that a rejected refresh cannot modify the backend object still
referenced by the committed provider record.

* fix(credentials): add timeouts to credential driver RPCs

Apply configured timeouts to both startup capability negotiation and runtime
RPCs to prevent indefinite hangs (GATOR-1806c9be-05).

- Add DEFAULT_CREDENTIAL_DRIVER_RPC_TIMEOUT_SECS constant (30s)
- Apply timeout to GetCapabilities during startup connection
- Apply timeout to all runtime RPCs (store, delete, resolve)
- Use tokio::time::timeout to bound the entire GetCapabilities operation
  during startup, not just the socket connection
- Return contextual deadline errors on timeout

This prevents a faulty or overloaded driver from hanging gateway operations
indefinitely.

* fix(credentials): fix test to use consistent workspace/provider identity

The Kubernetes auth Vault resolve test was constructing a managed path with
test-workspace/test-provider-id but sending default/prov-123 in the request,
causing validation to reject the request (GATOR-18e32351-01).

- Update test to use test-workspace and test-provider-id in the request to
  match the logical_path construction
- This ensures the test exercises the intended code path and validates
  Kubernetes auth resolution properly

The test now passes and correctly validates identity enforcement.

* fix(credentials): use unique staging ID for refresh to avoid overwrites

Stage refresh replacements under genuinely distinct immutable handles using
a unique staging ID to prevent overwriting committed values (GATOR-1806c9be-03).

- Generate a unique staging ID using UUID for each refresh operation
- Use this staging ID when storing credentials instead of the real provider ID
- Pass the same staging ID during cleanup on failure to delete only staged objects
- This ensures deterministic paths (Vault) and object names (K8s) don't collide
  with the committed provider's credentials

The fix prevents failed refreshes from silently replacing active credentials
or breaking providers by deleting still-referenced backend objects.

* fix(credentials): wrap credential driver RPCs in local timeouts

Add local tokio::time::timeout wrappers around credential driver RPCs to
bound non-compliant or stalled UDS peers (GATOR-1806c9be-05).

- Wrap StoreCredential, DeleteCredential, and ResolveCredentials in local timeouts
- Return contextual deadline_exceeded errors when timeouts occur
- Keep existing gRPC timeout metadata for compliant implementations
- GetCapabilities during startup was already wrapped in previous commit

This ensures a faulty local driver cannot hang gateway operations indefinitely,
even if it accepts the connection but never responds to the RPC.

* fix(credentials): preserve ownership for staged refreshes

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(credentials): bound startup capability probe

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* test(provider): authenticate credential handler requests

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(ci): grant actions read to credential driver e2e

Signed-off-by: Seth Jennings <sjenning@redhat.com>

---------

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com>
Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>
Signed-off-by: Seth Jennings <sjenning@redhat.com>
Co-authored-by: Taylor Mutch <taylormutch@gmail.com>
Co-authored-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>
2026-08-05 16:29:41 +00:00
Matthew Grossman 537805568d feat(sandbox): honor OCI image working directories (#2530)
* feat(sandbox): honor Docker OCI working directories

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* fix(sandbox): honor effective workspace access

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* test(sandbox): cover enforced workspace denial

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* docs(docker): explain effective workdir checks

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* fix(sandbox): validate effective workspace writes

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* fix(sandbox): reserve supervisor control roots

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* refactor(sandbox): centralize control paths

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* fix(sandbox): reserve OCI runtime mount roots

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

---------

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>
2026-08-04 17:38:51 +00:00
krishicks 8328412959 feat(vm): export driver traces over OTLP (#2564)
Continue distributed traces across the gateway-to-driver process boundary and
export VM driver spans to the same OTLP/gRPC collector. The driver reports as
the distinct openshell-driver-vm service.

Updated the gateway architecture and configuration reference with a generic
external-driver forwarding contract.

Instrumented:

- Every RemoteComputeDriver RPC injects the active W3C trace context into
  tonic metadata. Managed VM readiness and runtime initialization give startup
  capability probes stable parent operations rather than isolated root spans.
- A tonic service layer creates fixed, low-cardinality server spans for every
  ComputeDriver RPC. New handlers inherit tracing automatically; failures record
  OpenTelemetry error status and the gRPC status code.
- Background provisioning remains attached to CreateSandbox after the RPC
  returns without extending the RPC span lifetime.
- Provisioning records image preparation, bootstrap image resolution, overlay
  preparation, lifecycle configuration, pre-launch hooks, guest preparation,
  and launcher spawn as child spans.
- VM startup reconciliation roots one trace for the persisted-sandbox scan,
  with per-sandbox restore and provision operations beneath it. The root remains
  open until all spawned restore tasks finish.
- Delete cleanup records its own child operation.

Design notes:

- The gateway forwards its configured OTLP endpoint to managed external
  drivers. SDK `OTEL_*` variables continue to own sampling, batching, limits,
  headers, and transport tuning.
- The VM driver has its own tracer provider and service resource so trace
  backends preserve the service boundary.
- RPC operation names come from an explicit method mapping, keeping cardinality
  bounded without parsing the protobuf descriptor set at runtime.
- Propagation uses a remote SpanContext for spawned provisioning. This keeps
  one trace while allowing the CreateSandbox server span to finish when the
  RPC response is sent.
- Startup restoration is independent of gateway requests. It begins at the VM
  driver reconciliation span rather than attaching to an unrelated RPC.
- Existing tracing events remain on the logging path. The OpenTelemetry layer
  exports spans only and excludes the SDK exporter callsites to avoid recursive
  traces.
- Export configuration failures do not prevent the driver from serving, and
  buffered spans are drained during graceful shutdown.
- Trace fields identify drivers, sandboxes, images, lifecycle phases, and gRPC
  outcomes without recording credentials, sandbox tokens, or request query
  parameters.

Refs #2507

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-08-03 20:52:22 +00:00
John T. MyersandJohn Myers 905b554c7c refactor(network): consolidate proxy egress pipeline (#2373)
* refactor(network): introduce shared egress pipeline

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(network): cover shared proxy egress paths

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* refactor(network): make destination authorization explicit

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* refactor(network): pin proxy relay policy context

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(network): lock relay generation contracts

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(network): establish phase zero compatibility baseline

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* feat(policy): detect ambiguous network endpoints

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* feat(sandbox): fail closed on invalid policy updates

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* refactor(network): invalidate relays on policy changes

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(policy): document validation failure posture

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(network): cover validation and middleware egress

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(config): move policy failure mode to gateway toml

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(network): name proxy contracts by behavior

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): align overlap validation with endpoint selection

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(network): expect hard loopback denial

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(network): match declared endpoint denial

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(policy): preserve path-specific endpoint overrides

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(network): respect hard-blocked host gateways

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): reconcile proxy refactor with main

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): preserve CONNECT policy generation

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(policy): cover runtime endpoint glob semantics

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(server): reject ambiguous policies before persistence

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(policy): explain ambiguity preflight behavior

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(policy): compare body limits within protocol

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(proxy): avoid global tracing capture race

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* chore(server): format rebased provider tests

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(server): authenticate rebased policy requests

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(sandbox): retain runtime on middleware outage

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(server): preflight provider composition activation

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(sandbox): distinguish runtime failure transitions

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
Signed-off-by: John Myers <johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <johntmyers@users.noreply.github.com>
2026-07-31 20:59:00 +00:00
Evan LezarandDrew Newberry d220d89468 feat(compute): negotiate gateway callback listeners (#2492)
* feat(compute): query gateway listener requirements

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* feat(compute): add Podman listener requirements

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(docker): use default gateway bind address

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(gateway): avoid wildcard primary listener

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(podman): validate callback listener discovery

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(server): support split dual-stack listeners

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(podman): support legacy rootless listener discovery

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(e2e): accept loopback plaintext rejection

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* docs(agent): add callback listener diagnostics

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(server): restrict compute callback listeners

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(podman): validate local callback port

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(server): clarify callback listener contract

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(podman): require pasta for local callbacks

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* docs(gateway): document RPM listener default

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* refactor(server): keep listener provenance diagnostic-only

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(compute): preserve callback listener isolation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(e2e): remove Podman callback relay

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(packaging): preserve Podman callback loopback

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci(e2e): run VM smoke on nested-virt runner

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci(e2e): gate VM smoke on usable KVM

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci(e2e): probe KVM through VM driver

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci(e2e): tolerate hosted KVM denial

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(server): close traced futures before assertions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* revert: remove tracing test stabilization

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Evan Lezar <elezar@nvidia.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
2026-07-31 16:41:06 +00:00
krishicks fa24299092 feat(gateway): export traces over OTLP (#2534)
Add an opt-in OTLP/gRPC trace exporter to the gateway. Export is enabled
by the presence of an `[openshell.gateway.otlp]` table with an endpoint;
there is no separate toggle.

Instrumented:

- Inbound request server spans, named for the RPC (`$service/$method`) or
  `{method} {path}` for plain HTTP. They continue valid W3C `traceparent`
  context when present and start a new trace otherwise. gRPC spans also
  carry `rpc.system`, `rpc.service`, `rpc.method`, and trailer-derived
  `rpc.grpc.status_code`.
- Compute driver calls (create, delete, list, get, validate, watch) as
  client spans anchored on the `ComputeDriver` contract.
- Store reads and writes as children of the current request or loop span.
- Work with no inbound request: compute driver initialization, the sandbox
  reconcile sweep, provider credential refresh tick, and driver watch events.
  Each roots one operation trace so its child work does not arrive as anonymous
  single-span traces.

This is deliberately not exhaustive. Auth, policy evaluation, and
middleware remain uninstrumented, as do store lifecycle calls (`ping`,
`close`) that a readiness poll would turn into a span per tick. The aim is
a useful trace tree at a reviewable size; coverage can grow against real
traces.

Design notes:

- The TOML table owns whether and where to export. The SDK `OTEL_*`
  variables own how; sampling, batching, and limits are not mirrored into
  gateway config.
- The OpenTelemetry layer exports spans only. Existing `tracing` events
  remain on the stdout and sandbox-log paths and are not copied into trace
  payloads.
- Telemetry never blocks the gateway. A malformed endpoint logs an error
  and disables export rather than failing startup, and buffered spans are
  drained during graceful shutdown.
- Failed spans carry error status without a separate `error.type` attribute.
  Request spans use HTTP status and gRPC response trailers; driver spans use
  the returned gRPC status; autonomous loop spans record failed results
  explicitly. Store spans exempt `UniqueViolation` and `Conflict`, because
  those errors report expected contention such as a held lease or an
  optimistic-concurrency retry.
- The compute driver is reachable only through `TracedDriver::call`, so a
  call cannot skip its span. This is the client half of a client/server pair
  and the single place to inject context if drivers move out of process.
- Tests share one process-wide subscriber and in-memory exporter because
  `tracing` caches callsite interest globally.

Inbound W3C trace context is propagated into gateway request spans. Context
is not yet injected into outbound driver calls, so a future out-of-process
driver would still need propagation at the `TracedDriver` seam.

The Helm chart is intentionally unchanged, so OTLP export cannot yet be
enabled on a chart-deployed gateway.

Refs #2507

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-07-30 20:29:42 +00:00