mirror of
https://github.com/NVIDIA/OpenShell.git
synced 2026-10-02 07:34:45 +08:00
main
428
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
eb15e1a4c9 |
feat(sandbox): add --no-login-shell to skip shell startup files on exec (#2852)
* feat(sandbox): add --no-login-shell to skip shell startup files on exec Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * chore(sdk/go): regenerate proto bindings for no_login_shell Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * test(supervisor-process): cover login-shell flag selection Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * feat(sandbox): gate --no-login-shell on supervisor SSH banner Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> --------- Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> |
||
|
|
69a05ebb3b | fix(sandbox): complete successful main processes (#2884) | ||
|
|
f68867b869 |
feat(gateway): identify gateways in exported traces (#2647)
* feat(gateway): add installation name configuration Add a first-class operator-assigned gateway name with TOML, CLI, environment, and Helm configuration surfaces. Local gateways default to openshell, while Helm defaults to the chart fullname; operators sharing a collector across namespaces or clusters can set a globally distinct name. Signed-off-by: Kris Hicks <khicks@nvidia.com> * feat(gateway): identify gateways in exported traces Attach the configured gateway installation name and compute driver to the gateway OpenTelemetry resource so operators can filter traces from multiple installations that share a collector. Forward the gateway name and OTLP endpoint to managed external drivers so their distinct service resources carry the same installation identity. Keep service.name stable per process type, omit blank resource values, and leave per-span operation names and request attributes unchanged. Refs #2507 Signed-off-by: Kris Hicks <khicks@nvidia.com> --------- Signed-off-by: Kris Hicks <khicks@nvidia.com> |
||
|
|
37072ee81c |
feat(build): publish OCI SBOM and provenance attestations (#2836)
* feat(build): embed auditable Rust dependency metadata Signed-off-by: Adrien Langou <alangou@nvidia.com> * feat(build): publish OCI SBOM and provenance attestations Signed-off-by: Adrien Langou <alangou@nvidia.com> --------- Signed-off-by: Adrien Langou <alangou@nvidia.com> |
||
|
|
d0dfb22baf |
feat(kubernetes): export driver traces over OTLP (#2958)
Mirror the VM, Podman, and Docker driver tracing setup for Kubernetes. Export standalone driver spans through OTLP/gRPC as the distinct openshell-driver-kubernetes service, preserve gateway trace context, record lifecycle operations and gRPC failures, and flush spans on shutdown. Kubernetes currently runs in-process when selected as a built-in gateway driver. Use the temporary server-boundary shim shared with Podman and Docker so traces retain the shape they will have when Kubernetes moves to a separate process. Move the common ComputeDriver RPC tracing layer into openshell-otel to keep all drivers aligned. Propagate the active W3C context through the controller-reserved Sandbox annotation and enable Agent Sandbox OTLP export in the local k3s workflow. This connects asynchronous controller reconciliation spans to the originating OpenShell create trace. Expose gateway OTLP configuration through Helm and add an Aspire collector to the local k3s workflow. Extend helm:k3s:forward with OTLP ingest and trace UI forwarding for Kubernetes and local container gateway development. Signed-off-by: Kris Hicks <khicks@nvidia.com> |
||
|
|
4e992093f8 |
fix(docker): trace standalone driver over OTLP (#2923)
Signed-off-by: Evan Lezar <elezar@nvidia.com> |
||
|
|
0e79653a7f |
feat(tui): show persisted guidance on rejected policy chunks (#2908)
* feat(tui): show persisted guidance on rejected policy chunks The reviewer's rejection note is stored in PolicyChunk.rejection_reason and already reaches the TUI in GetDraftPolicyResponse.chunks, but openshell-tui never read the field. The note was dropped at the last step, so a reviewer had no way to recall why a chunk had been rejected. Render it in two places, following the truncate-in-list / full-in-popup convention in the TUI development guide: a truncated, dimmed suffix on the list row, and a "Guidance:" line in the detail popup. Gate the accessor on status == "rejected" rather than on the field alone. Approving a chunk passes None for the reason and the gateway writes the field only when Some, so a chunk that was rejected and later approved still carries the old note. Reading the field unconditionally would surface a stale rejection on an approved rule. Part of #1098, Definition of Done item "Rejected chunks show persisted guidance". The other items on that issue are untouched. Signed-off-by: Vyncint Ng <115854244+vyncint@users.noreply.github.com> * fix(tui): scroll the draft detail popup so long guidance stays reachable A rejection reason has no server-side length cap, so a paragraph-length one overflowed the fixed 22-row detail popup: the tail was clipped and the later fields and both action hints were pushed off screen. The same overflow already affected the unbounded rationale and security_notes fields, so fix the popup rather than special-case the guidance line. Split the popup's inner area into a scrolling body and a pinned hint row, following the pattern in ui/create_provider.rs, and drive it with j/k, the arrow keys, PageUp, PageDown, g and G. The approve and close controls now stay on screen at every scroll position, and the bottom border carries the scroll position. Wrap free-form values explicitly instead of relying on Paragraph's Wrap, so the rendered row count is exactly lines.len() and the scroll clamp cannot under-run the content. Paragraph::line_count would answer the same question but sits behind ratatui's unstable-rendered-line-info feature. Add deterministic TestBackend coverage at 80x24 with a 2,000-character reason, covering the head and tail, the pinned hints, over-scroll clamping, and the pre-existing long-rationale case. Document the reviewer-facing guidance in the policy advisor page. Part of #1098. Signed-off-by: Vyncint Ng <115854244+vyncint@users.noreply.github.com> * fix(tui): wrap the denial row and hint footer on a narrow popup Making the scroll clamp exact meant dropping Paragraph's Wrap, which also removed wrapping from the lines that do not go through push_wrapped. At 70 columns the denial row clipped mid-value and lost the last-seen timestamp, and at 60 a pending chunk with scrollable content pushed [Esc] Close past the right edge of the single hint row. The key still worked, but that is the same controls-not-visible failure one axis over. Keep the denial row on one line while it fits and wrap it onto the label indent when it does not, so the common 80-column layout is unchanged. Pack the footer hints into as many rows as they need without splitting a hint, and derive the footer height from that: adding a hint row shrinks the body and can itself change whether the content scrolls, so the two settle together. The earlier tests were all 80x24, which is why they missed this. Cover the denial timestamps and the approve, reject and close hints at 60, 70 and 80 columns, plus the packing helper directly. Part of #1098. Signed-off-by: Vyncint Ng <115854244+vyncint@users.noreply.github.com> * fix(tui): measure popup wrapping in display columns, not chars wrap_value packed rows by chars().count(), and a CJK glyph is one char but two terminal columns. A double-width rejection reason therefore produced rows about twice the popup's width: the right half of every row was clipped, and unlike the vertical case there is no horizontal scroll to recover it. Measure width with Span::width(), the same measurement ratatui applies when it lays cells out, so the wrap agrees with the renderer by construction and no new dependency is needed. This covers word packing, the hard-break path for a word wider than the line, the label indent, the footer hint packing, and the denial row's fits-on-one-line check. The list row's shortened copy had the same defect through truncate_str, which counts chars. Leave truncate_str alone for its two existing callers and add truncate_display for the guidance row. Cover it with an 80x24 render regression that interleaves markers through the CJK text: a trailing marker lands on its own short row in both the broken and fixed layouts, so it would not detect this. The two unit tests measure with an independent column oracle rather than the helper under test, which would make them tautological. Part of #1098. Signed-off-by: Vyncint Ng <115854244+vyncint@users.noreply.github.com> --------- Signed-off-by: Vyncint Ng <115854244+vyncint@users.noreply.github.com> |
||
|
|
18ce13b9b1 |
feat(providers): expose actionable OAuth refresh failures (#2887)
* fix(providers): classify OAuth refresh failures Signed-off-by: Mrunal Patel <mrunalp@gmail.com> * test(providers): add Keycloak refresh e2e lane Signed-off-by: Mrunal Patel <mrunalp@gmail.com> * fix(providers): harden OAuth refresh recovery Signed-off-by: Mrunal Patel <mrunalp@gmail.com> * fix(providers): classify post-mint refresh failures Signed-off-by: Mrunal Patel <mrunalp@gmail.com> * fix(providers): tolerate malformed OAuth subtypes Signed-off-by: Mrunal Patel <mrunalp@gmail.com> --------- Signed-off-by: Mrunal Patel <mrunalp@gmail.com> |
||
|
|
5206bc51b2 |
feat(sdk): add OAuth Client Credentials support to SDKs (#2907)
* feat(sdk): add renewable client credentials auth Implement lazy OAuth client-credentials acquisition and renewal for the Python, TypeScript, and Go SDK clients, with shared security conformance coverage and service-account documentation.\n\nCloses #2803 Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(sdk): address client credentials review Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(sdk): address client credentials review Signed-off-by: Seth Jennings <sjenning@redhat.com> --------- Signed-off-by: Seth Jennings <sjenning@redhat.com> |
||
|
|
455883905a |
fix(python): remove CLI from wheel (#2321)
The Maturin-based wheel packaging was a historical remnant from when the local gateway launch path and OpenShell CLI were coupled in one binary. The gateway and CLI now ship as standalone artifacts, so the Python distribution should contain only the SDK. Build a single platform-independent setuptools wheel, verify that it cannot contain native code or an openshell entry point, and simplify the release jobs and documentation for SDK-only PyPI installs. Signed-off-by: Simon Scatton <sscatton@nvidia.com> |
||
|
|
aa848f164d |
fix(core): fall back to podman CLI when no API socket is found (#1858)
* fix(core): fall back to podman CLI when no API socket responds Auto-detection only checked well-known Podman socket paths, so a Podman machine exposing its API socket at a non-standard location went undetected. The symlink at a well-known path is not always present; it varies by Podman version, machine provider, and platform. Extend detect_podman_socket() to fall back to podman CLI discovery when no well-known candidate responds: podman info --format json determines whether the service is local or remote, and podman machine inspect resolves the host-side forwarded socket for VM-backed machines. All existing callers (driver auto-detection, the Podman driver, and the VM driver's container-engine fallback) pick this up without change. Select the machine backing the active Podman connection instead of the first entry from podman machine inspect, honoring podman's connection precedence: CONTAINER_CONNECTION, then CONTAINER_HOST (mapped to a connection by URI), then the containers.conf default. An explicit endpoint that maps to no known machine is left unresolved rather than guessing an unrelated machine. When CONTAINER_HOST is an explicit unix:// socket, use that path directly since podman info connects through it. Update the gateway config reference and the Podman driver README, which described probe-only detection. Signed-off-by: Russell Bryant <rbryant@redhat.com> * fix(core): inspect podman machine by name and bound discovery probes Address two startup defects in Podman socket auto-detection: - Remote discovery ran `podman machine inspect` with no arguments, which inspects only `podman-machine-default`. On a host whose default connection is a different machine, selection fell back to the first (wrong) entry, so the gateway could operate against the wrong backend. Resolve the active machine first and inspect it by name, trying the `-root`-stripped machine for rootful connections and returning None when no machine can be mapped instead of substituting another. - `podman info`, `podman machine inspect`, and `podman system connection list` used unbounded `Command::output()`, so a stalled machine, SSH connection, or helper could hang gateway startup indefinitely. Route all three through a bounded runner that kills and reaps the child on a documented deadline and returns None so detection can continue. Signed-off-by: Russell Bryant <rbryant@redhat.com> * fix(core): bound podman probe stdout drainage and kill descendants The previous timeout only bounded waiting for the direct child to exit. Draining stdout still called read_to_end, which waits for pipe EOF — a Podman probe can leave a daemonized descendant (SSH multiplexer, gvproxy) that inherited the stdout pipe, so drainage could wait indefinitely even after the direct child exited, defeating the deadline. Run each probe as its own process-group leader, cap stdout drainage by the same deadline via a channel, and on expiry kill the whole process group so descendants holding the pipe are terminated and EOF is reached. Add a regression test where the direct child exits but a backgrounded descendant keeps stdout open. Signed-off-by: Russell Bryant <rbryant@redhat.com> * fix(core): make podman probe deadline absolute against escaped descendants The prior drainage timeout still ended in an untimed rx.recv() after killing the process group. A descendant that escaped the probe's process group (e.g. via setsid) while holding the stdout pipe would not be killed, so read_to_end never saw EOF, the reader never sent, and that recv() could block gateway startup indefinitely. On drainage timeout, best-effort kill the process group and return None immediately, abandoning the reader thread instead of waiting on it again. The deadline now bounds the whole call regardless of what descendants do. Add a regression test whose stdout-holding descendant escapes the process group (`set -m`) and assert the call still returns promptly. Signed-off-by: Russell Bryant <rbryant@redhat.com> --------- Signed-off-by: Russell Bryant <rbryant@redhat.com> |
||
|
|
0a1f246587 |
feat(sandbox,podman): trust corporate CA for https:// proxies and intercepted TLS (#2512)
* feat(sandbox,podman): trust corporate CA for https:// proxies and intercepted TLS The corporate proxy chaining only accepted plain http:// proxy URLs, so operators whose forward proxy terminates TLS with a private corporate CA had no way to reach it, and TLS-intercepting proxies (mitmproxy, squid ssl-bump) that re-sign tunneled server certificates broke every upstream handshake after CONNECT. The supervisor now accepts https:// proxy URLs: it wraps the connection to the proxy in TLS before the CONNECT handshake, verifying the proxy certificate against the built-in Mozilla roots, the system CA bundle, and an optional operator corporate CA bundle. The upstream dial returns a Plain/Tls stream enum consumed generically by the relay paths. The corporate CA is delivered as a driver-supplied command-line argument (--upstream-proxy-ca-bundle), never an environment variable, matching the hardened proxy-config model where a sandbox image cannot influence the operator's egress boundary. It is folded into the sandbox combined trust bundle (write_ca_files) and the L7 upstream verification store (build_upstream_client_config) at startup, so intercepted upstream handshakes succeed and sandbox workloads trust the re-signed certificates. Configuration is fail-closed: a CA bundle set without a proxy, or an unreadable or certificate-free file, is fatal rather than silently weakening the trust boundary. The shared parse_upstream_proxy_url validator accepts https:// (recording the scheme so the driver and supervisor agree), keeping the explicit-port requirement. The Podman driver gains a proxy_ca_bundle operator setting (TOML, --sandbox-proxy-ca-bundle, OPENSHELL_SANDBOX_PROXY_CA_BUNDLE) that bind-mounts the host PEM read-only into the sandbox (a CA certificate is not secret) and points --upstream-proxy-ca-bundle at it, with a create-time readability check. The standalone dev gateway task passes OPENSHELL_SANDBOX_PROXY_CA_BUNDLE through to the generated podman config, so a local gateway can be pointed at a TLS-intercepting proxy without hand-editing the regenerated TOML. Refs #1792 Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(sandbox): reject CA bundles with valid PEM framing but invalid X.509 DER The proxy CA bundle validation counted PEM blocks that base64-decoded successfully, but did not verify the decoded bytes were accepted as trust anchors by RootCertStore. A bundle with syntactically valid PEM framing but invalid DER would pass the startup check while contributing zero usable anchors, causing opaque TLS failures at runtime instead of a fail-closed startup error. Validate decoded certificates through RootCertStore::add_parsable_certificates and reject the bundle unless at least one is accepted. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(sandbox,podman): allow proxy auth without insecure acknowledgement for https:// proxies For an https:// proxy the Proxy-Authorization credential travels inside the verified TLS session, so the proxy_auth_allow_insecure acknowledgement is unnecessary. Previously both http:// and https:// proxies required it, producing a misleading cleartext-risk diagnostic for a path that is already encrypted. Skip the requirement when the proxy URL uses https://; the acknowledgement is still tolerated if set. Updated in both the supervisor and Podman driver validation paths, with docs and tests. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix: format Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(kubernetes): use truly unsupported scheme in proxy validation test https:// is now a supported proxy scheme after a13c4dce, so the unsupported-scheme test must use a genuinely unsupported scheme. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(e2e): sign the https proxy fixture listener cert with a CA The corporate-proxy E2E fixture served a single `openssl req -x509` certificate as its TLS listener identity. OpenSSL marks that certificate `basicConstraints: critical, CA:TRUE`, and rustls refuses a CA certificate presented as an end-entity certificate (CaUsedAsEndEntity). The supervisor's TLS handshake with the proxy therefore failed, the upstream dial errored, and the workload's CONNECT was dropped without a response, so podman_corporate_proxy_trusts_ca_bundle_for_https_proxy failed on the approved destination while policy denial still worked. Generate a corporate CA and a separate listener leaf signed by it, serve the leaf chain, and publish only the CA as the bundle the supervisor trusts. This is what an intercepting proxy actually presents, and it exercises the corporate-CA trust path rather than pinning the listener certificate itself. Refs #1792 Signed-off-by: Philippe Martin <phmartin@redhat.com> --------- Signed-off-by: Philippe Martin <phmartin@redhat.com> |
||
|
|
e457974a52 |
feat(docker): export driver traces over OTLP (#2851)
Mirror the VM and Podman driver tracing setup for Docker. Export Docker driver spans through OTLP/gRPC as the distinct openshell-driver-docker service, preserve gateway trace context, record lifecycle and asynchronous provisioning operations, and report gRPC failures. Docker currently runs in-process when selected as a built-in gateway driver. Add the same temporary server-boundary shim used by Podman so traces retain the shape they will have when Docker moves to a separate process. Generalize the gateway provider selection for both in-process drivers and share the OTLP collector fixture across Docker, Podman, and VM tracing tests. Signed-off-by: Kris Hicks <khicks@nvidia.com> |
||
|
|
40d1b48666 |
feat(provider): support for SPIFFE backed token exchange (#1970)
* feat(provider): add ability to request token exchange instead of client credentials as OAuth grant_type Signed-off-by: Gordon Sim <gsim@redhat.com> * test(proxy): add further tests for token exchange Signed-off-by: Gordon Sim <gsim@redhat.com> * test(provider): add runnable example for token exchange Signed-off-by: Gordon Sim <gsim@redhat.com> * test(e2e): cover Podman token exchange grants Signed-off-by: Gordon Sim <gsim@redhat.com> * refactor(oauth): extract duplicated functionality from server and supervisor Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(provider): evict nearest-to-expiry entry from intermediate token cache Signed-off-by: Gordon Sim <gsim@redhat.com> * doc(supervisor): add podman example for token exchange Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(provider): withhold token-exchange subject credentials Signed-off-by: Gordon Sim <gsim@redhat.com> --------- Signed-off-by: Gordon Sim <gsim@redhat.com> |
||
|
|
679fe4c334 |
fix(policy): validate the applicable advisor candidate (#2850)
* fix(policy): bind reviews to applicable candidates Build and validate the exact effective-policy candidate before approval, bind review to live policy/provider/credential inputs, and preserve inspected endpoint contracts during mechanistic expansion. Closes #2821 Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * fix(policy): canonicalize advisor review inputs Serialize nested protobuf maps in stable key order for proposal review tokens and effective-policy hashes. Narrow reused multi-port endpoint contracts to the denied port so advisor proposals cannot widen binary access. Add regressions for both cases. Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * test(e2e): keep advisor sandbox running Create the issue 2821 regression sandbox detached with a durable canonical main process so policy denial, approval, and hot-reload checks run before lifecycle exit. Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * fix(policy): apply reviewed draft batches atomically Signed-off-by: John Myers <johntmyers@users.noreply.github.com> --------- Signed-off-by: John Myers <johntmyers@users.noreply.github.com> Co-authored-by: John Myers <johntmyers@users.noreply.github.com> |
||
|
|
7adc05af7a |
feat(supervisor): expose sandbox name to middleware request context (#2771)
* feat(supervisor): expose sandbox name to middleware request context Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * feat(supervisor): add workspace to middleware request context Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> --------- Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> |
||
|
|
3be2cd8a29 |
fix(helm): preflight Agent Sandbox APIs (#2867)
* fix(helm): preflight Agent Sandbox APIs Signed-off-by: Evan Lezar <elezar@nvidia.com> * test(kubernetes): share Agent Sandbox setup Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(e2e): wait for Agent Sandbox CRD status Signed-off-by: Evan Lezar <elezar@nvidia.com> * ci(canary): sparse-checkout sandbox helper Signed-off-by: Evan Lezar <elezar@nvidia.com> --------- Signed-off-by: Evan Lezar <elezar@nvidia.com> |
||
|
|
ef296806f5 |
feat(sandbox): add canonical main process (#2726)
* feat(sandbox): add canonical main process Closes #2710 Persist and supervise one canonical workload per sandbox, attach sandbox connect to its retained session, and make every unexpected main-process exit terminal. Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(sandbox): simplify canonical main process contract Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): preserve legacy VM main compatibility Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): preserve main status across driver updates Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): satisfy macOS process lint Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): gate Linux exit acknowledgement publisher Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(sandbox): simplify main process plumbing Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(supervisor): make controlling tty ioctl portable Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(supervisor): initialize canonical process environment Signed-off-by: Drew Newberry <anewberry@nvidia.com> * perf(supervisor): optimize retained main session Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(sandbox): detach main session on ctrl-c Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): use explicit main detach keys Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(dev): atomically stage Docker supervisor Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(test): align Docker main environment assertion Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(sdk): expose canonical main process fields Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
b2ea81822b |
feat(network): enable Docker and Podman policy DNS and transparent TCP (#2723)
* feat(network): enable Docker transparent TCP egress Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(e2e): cover Docker transparent TCP egress Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * feat(network): correlate transparent TCP audit events Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * docs(examples): add transparent TCP Redis demo Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * docs(examples): demonstrate blocked TCP connections Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * docs(examples): focus Redis demo audit output Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(network): close transparent TCP policy bypasses Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(sandbox): reject unsupported TCP policy reloads Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(ci): satisfy Linux transparent TCP lints Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * feat(podman): enable transparent TCP egress Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(podman): permit policy DNS port binding Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(e2e): use qualified transparent TCP hostname Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(podman): preserve exact policy DNS names Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(podman): route policy DNS over TCP Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(dns): serve multiple TCP queries per connection Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * docs(network): explain native DNS and TCP egress Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(sandbox): reconcile runtime reload with upstream Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(network): harden transparent DNS capture Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(podman): preserve resolver behavior for native tcp Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * docs(network): clarify native tcp runtime constraints Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(network): remove unused transparent tcp pin Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(network): admit redirected transparent tcp Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(network): restore podman transparent networking Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(podman): permit alpine busybox binaries Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(podman): use portable alpine keepalive Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(podman): build musl networking fixture Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(podman): isolate musl DNS probe Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(podman): keep privileged port capability dropped Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(network): preserve transparent TCP port 53 Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(network): report synthetic pool pressure by family Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(podman): bind tcp fixtures before readiness Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(podman): grant fixture low-port bind Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> --------- Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> |
||
|
|
4d7f402ce2 |
feat(policy): establish direct TCP egress foundation (#2711)
* feat(policy): accept explicit tcp endpoint protocol Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * refactor(network): snapshot authoritative egress decisions Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * docs(policy): document explicit tcp protocol Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * docs(policy): defer transparent TCP release guidance Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * chore(go): regenerate sandbox protobuf bindings Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(network): complete tcp egress foundation Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * docs(policy): document explicit tcp contract Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(network): fail closed on authorization errors Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(podman): fence delayed exit events before restart Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(policy): validate network endpoint destinations Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(providers): opt in tcp credential fixture Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(policy): require dns host for transparent tcp Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> --------- Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> |
||
|
|
4c5fce6e59 |
refactor(compute): support external driver parity (#2744)
* refactor(compute): negotiate external driver behavior Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(compute): revert external driver documentation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(compute): negotiate gateway-managed lifecycle Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(compute): remove driver feature negotiation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(compute): let drivers declare gateway lifecycle Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(compute): clarify lifecycle ownership Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
7909fb5d0f |
refactor(compute): unify gateway restart reconciliation (#2743)
* refactor(compute): unify gateway restart reconciliation Remove the Docker-specific gateway shutdown cleanup and reconcile persisted running intent through ComputeDriver::StartSandbox for Docker, Podman, and VM drivers. Explicitly stopped sandboxes remain stopped. Refs #2417 Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(compute): stop local sandboxes on shutdown Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(e2e): match managed Podman containers Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(compute): synchronize lifecycle sweeps Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
6e90f3d5a0 |
feat(providers): store refresh credentials in credential drivers (#2801)
* feat(providers): store refresh credentials in credential drivers Signed-off-by: Mrunal Patel <mrunalp@gmail.com> * fix(providers): harden refresh credential lifecycle Signed-off-by: Mrunal Patel <mrunalp@gmail.com> * fix(providers): migrate legacy refresh secrets before skip Signed-off-by: Mrunal Patel <mrunalp@gmail.com> * refactor(providers): defer credential migration Signed-off-by: Mrunal Patel <mrunalp@gmail.com> * fix(providers): make refresh configuration atomic Signed-off-by: Mrunal Patel <mrunalp@gmail.com> --------- Signed-off-by: Mrunal Patel <mrunalp@gmail.com> |
||
|
|
998db04780 |
feat(policy): allow non-root sandbox identities (#2785)
Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
2eb0880a00 |
feat(cli): support OIDC device authorization grant for headless login (#2795)
* feat(cli): support OIDC device authorization grant for headless login Closes #2793 Add OAuth 2.0 Device Authorization Grant (RFC 8628) support to the OpenShell CLI's OIDC login flow. When running in a headless environment (OPENSHELL_NO_BROWSER=1) without a client secret configured, the CLI now uses the device code flow instead of the browser-based PKCE flow. The device code flow: - Requests a device code and user code from the IdP's device authorization endpoint - Displays a verification URL and user code to the user - Polls the token endpoint until the user completes authorization or the code expires - Supports slow_down responses per RFC 8628 by increasing the polling interval This implementation: - Extends OidcDiscovery to optionally capture device_authorization_endpoint - Adds oidc_device_code_flow function with proper error handling for all RFC 8628 error codes - Updates gateway add and gateway login to dispatch to device flow when browser is suppressed - Adds comprehensive unit tests for device flow structs and response parsing - Updates gateway authentication documentation to describe the device code fallback Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(cli): validate OIDC device token responses Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(cli): add PKCE to OIDC device flow Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * docs(cli): document PKCE device flow Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> --------- Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> |
||
|
|
3a16012dbe |
fix(cli): prompt for fresh OIDC login after logout (#2773)
Signed-off-by: Gordon Sim <gsim@redhat.com> |
||
|
|
0d708d6d51 |
fix(policy): gate uninspected credentialed endpoints (#2493)
* fix(policy): gate uninspected credentialed endpoints Signed-off-by: Adrien Langou <alangou@nvidia.com> * refactor(cli): extract allowed-ip option parsing Signed-off-by: Adrien Langou <alangou@nvidia.com> * fix(policy): gate endpointless credential bindings Signed-off-by: Adrien Langou <alangou@nvidia.com> --------- Signed-off-by: Adrien Langou <alangou@nvidia.com> |
||
|
|
8d67250a5d |
fix(providers): keep refresh credential handles stable (#2780)
* fix(providers): keep refresh credential handles stable Signed-off-by: Mrunal Patel <mrunalp@gmail.com> * fix(providers): protect refresh-owned credentials Signed-off-by: Mrunal Patel <mrunalp@gmail.com> --------- Signed-off-by: Mrunal Patel <mrunalp@gmail.com> |
||
|
|
600bbae845 |
feat(ocsf): emit AI inference events via ai_operation profile on ApiActivity (#2664)
Apply the official OCSF ai_operation profile (introduced in v1.8.0) to ApiActivity [6003] events when the inference proxy routes a model call through inference.local. Attaches an ai_model object (name, ai_provider) and puts token counts and latency in unmapped fields. ApiActivity [6003] is the schema-correct class for the ai_operation profile in v1.8.0 (HttpActivity only gets it in v1.9.0). In Splunk CIM, ApiActivity maps to the "Change" data model, naturally separating inference events from regular HTTP proxy traffic. Changes: - Add AiModel object and ai_model field on BaseEventData - Add ApiActivityEvent struct and ApiActivityBuilder - Add emit_ai_inference in proxy.rs using ApiActivity with ai_operation - Vendor OCSF v1.8.0 schemas including api_activity class, ai_model object, and ai_operation profile definitions - Bump OCSF_VERSION to 1.8.0 - Update schema validation to skip profile-gated required fields Shorthand: API:INFERENCE [INFO] claude-3-haiku via anthropic 701ms [POST /v1/messages] Splunk/SIEM backward compatibility (v1.1/v1.3 CIM mapping) is tracked separately in #2662 as a configurable serialization concern. Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
d51a653f9c |
feat(driver-podman): add userns config (#2562)
* refactor(driver): extract shared supervisor binary helpers Move supervisor binary extraction, caching, and validation helpers from the Docker driver into openshell-core::driver_utils so both Docker and Podman drivers can reuse them. Moved helpers: extract_first_tar_entry, write_cache_binary_atomic, supervisor_cache_path, temp_extract_container_name, and validate_linux_elf_binary. The shared extract_first_tar_entry gains entry-type and empty-payload checks that the Docker-local version lacked. supervisor_cache_path takes a driver_subdir parameter so each driver caches under its own namespace (docker-supervisor vs podman-supervisor). Signed-off-by: Giuseppe Scrivano <gscrivan@redhat.com> * feat(driver-podman): add userns config Add a `userns` option to the Podman compute driver that maps to Podman's user namespace modes. The mode string is split on the first colon into the API's `nsmode` and `value` fields so parameterized values like `auto:size=65536` and `keep-id:uid=1000,gid=1000` are forwarded correctly. When the mode is `auto`, the container spec also sets `idmappings.AutoUserNs = true` as required by the API. An allowlist validates the mode at startup: `auto` and `keep-id` accept optional parameters; `host`, `private`, and `nomap` reject them; everything else is an error. Podman image volumes use overlay mounts internally and the kernel does not support idmapped mounts on overlay (`mount_setattr` returns EINVAL). When userns is configured (any mode except `host`), the driver extracts the supervisor binary from the image to a host-side cache and bind-mounts it instead of using an image volume. Configurable via TOML `userns = "auto"`, CLI `--userns`, or environment variable `OPENSHELL_PODMAN_USERNS`. Signed-off-by: Giuseppe Scrivano <gscrivan@redhat.com> --------- Signed-off-by: Giuseppe Scrivano <gscrivan@redhat.com> |
||
|
|
44bf0df485 |
feat(middleware): inspect WebSocket text messages (#2477)
* feat(middleware): inspect websocket text messages Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(middleware): address websocket review feedback Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(network): bound websocket message assembly Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(network): harden websocket upgrade lifecycle Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * refactor(middleware): unify in-process and remote transports Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * feat(middleware): support regex websocket redaction Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(middleware): bound persistent streaming sessions Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(middleware): accept websocket sequence gaps Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * refactor(middleware): refine websocket introspection contract Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(middleware): clarify websocket preflight lifecycle Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(middleware): clarify websocket coverage semantics Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(network): type websocket frame failures Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(network): return 503 when middleware admission is exhausted Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(middleware): align streaming API contract Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(middleware): clarify WebSocket event result scope Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(rfc): simplify middleware revision history Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * feat(examples): add WebSocket content guard support Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(middleware): unify binding payload limits Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * refactor(middleware): align payload limit terminology Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(middleware): address websocket review feedback Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(network): address websocket review findings Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * test(network): allow Linux handler setup in preflight regression Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(network): harden websocket relay finalization Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(network): inspect compressed websocket messages Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * test(network): stabilize compressed websocket regressions Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(go-sdk): regenerate middleware protobuf binding Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(middleware): clarify websocket skip lifecycle Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(middleware): address WebSocket review feedback Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> --------- Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> |
||
|
|
59479f492a |
feat(k8s): add namespace-per-workspace support (RFC 0011 Phase 3) (#2656)
* feat(k8s): add namespace-per-workspace support (RFC 0011 Phase 3) Implement three workspace namespace modes for the Kubernetes compute driver: shared (default, preserves current single-namespace behavior), managed (auto-creates/deletes namespaces per workspace), and operator (pre-provisioned namespaces with dynamic discovery via label selector or drop-in allowlist file). Key changes: - WorkspaceMode enum and namespace resolution in driver config - Managed namespace lifecycle with ServiceAccount and OpenShift SCC annotation propagation - Cluster-wide sandbox CR watchers for managed/operator modes - NamespaceValidator (Exact/Prefix/Allowlist) for SA token auth - Workspace-aware credential secret storage - Helm ClusterRole for multi-namespace RBAC - Gateway config, architecture, and reference docs Signed-off-by: Derek Carr <decarr@redhat.com> * test(k8s): add e2e tests for workspace namespace modes Add end-to-end tests for managed and operator workspace modes introduced in RFC 0011 Phase 3. The managed mode tests verify namespace creation with correct labels, ServiceAccount provisioning, sandbox CR placement, and namespace survival with remaining sandboxes. The operator mode tests verify rejection of unlabeled and nonexistent namespaces. The positive operator path (sandbox in labeled namespace) is known to fail due to an RBAC gap and will be addressed separately. Also fixes Helm 4 compatibility: move SPDX license headers inside conditional guards in 8 chart templates to prevent empty comment-only documents, and fix a trailing whitespace trimmer in clusterrole.yaml that concatenated the license header with apiVersion. Adds cleanup sweep in with-kube-gateway.sh to remove managed and operator namespaces before Helm uninstall, and mise tasks for running each mode independently. Signed-off-by: Derek Carr <decarr@redhat.com> * feat(k8s): add operator namespace label watcher Spawn a background kube::runtime::watcher in the K8s driver that watches namespaces matching the configured label selector and populates the OperatorNamespaceAllowlist at runtime. The driver owns the allowlist and exposes its Arc so the server can share the same set with the SA token authenticator. create_sandbox now gates pod creation on the allowlist in operator mode — workspaces whose namespace is not yet labeled are rejected at resource render time rather than silently proceeding. Workspace lifecycle itself is unaffected; only sandbox (resource) creation is gated. Signed-off-by: Derek Carr <decarr@redhat.com> * fix(k8s): harden operator mode and address review findings Close the fail-open gap in operator mode when only operator_namespace_file is configured: the allowlist is now created unconditionally in operator mode (fail-closed from startup). Implement the namespace file watcher using the notify crate, following the TLS hot-reload pattern (parent-directory watch, 1s debounce, ConfigMap symlink-swap safe). The file format is a JSON array of namespace name strings. Additional fixes from the 10-reviewer audit: - Change allowlist rejection from InvalidArgument to FailedPrecondition so callers know the request may succeed later once the namespace is provisioned. - NamespaceValidator::Allowlist now holds the OperatorNamespaceAllowlist newtype instead of a raw Arc<RwLock<BTreeSet>>, eliminating silent denial on RwLock poison. - Verify LABEL_MANAGED_BY and LABEL_GATEWAY_ID ownership before deleting a managed namespace. - Replace fixed 5s sleep in operator e2e test with a 30s poll loop. - Add Helm validation for workspaceMode values. - Fix Helm README type column and description for operator fields. - Add insert/remove methods to OperatorNamespaceAllowlist; label watcher now uses them instead of reaching through shared(). - Reject configs with both operator_namespace_label and operator_namespace_file set. Signed-off-by: Derek Carr <decarr@redhat.com> * feat(k8s): add workspace-level compute driver RPCs and harden RBAC Decouple namespace lifecycle from sandbox lifecycle by adding EnsureWorkspace/DeleteWorkspace RPCs to the ComputeDriver service. Namespace creation now happens before credential storage and namespace deletion happens on workspace delete, fixing credential storage in managed workspace mode. - Add EnsureWorkspace and DeleteWorkspace proto RPCs with implementations across all compute drivers (K8s managed delegates to ensure_namespace/delete_namespace_if_empty; others no-op) - Wire ensure_workspace into provider create/update/refresh paths so the namespace exists before the credential driver writes secrets - Wire delete_workspace into workspace deletion for cleanup - Remove delete_namespace_if_empty from sandbox deletion path - Scope ClusterRole secrets access to non-shared workspace modes - Add TODO for TLS cert hot-reload in sandbox gRPC client - Harden e2e tests with control-plane sandbox resolution assertions - Fix docker image save --platform flag for OCI index manifests Signed-off-by: Derek Carr <decarr@redhat.com> * fix(k8s): address re-review findings and add test coverage - Use server-side apply for TLS secret sync (fixes second sandbox creation failure when TLS is enabled) - Scope gateway-ID label selector unconditionally across all workspace modes (fixes operator reads/watches/deletes seeing foreign sandboxes) - Validate operator allowlist in EnsureWorkspace and DeleteWorkspace RPCs (prevents credential writes to namespaces outside the allowlist) - Extend ClusterRole secrets patch+delete to all non-shared modes with credential driver enabled (fixes operator credential storage RBAC) - Validate namespace ownership on 409 conflict in ensure_namespace (prevents adopting unowned namespaces in managed mode) - Replace delete_namespace_if_empty with unconditional delete_namespace letting Kubernetes cascade cleanup (fixes stuck terminating CRs) - Strengthen NetworkPolicy TODO to cover both managed and operator modes - Extract selector and ownership logic into testable free functions - Add unit tests for gateway-ID selectors and namespace ownership - Add Helm ClusterRole RBAC tests for operator credential driver Signed-off-by: Derek Carr <decarr@redhat.com> * ci(k8s): add workspace managed and operator mode e2e to CI Wire the existing e2e:kubernetes:workspace-managed and e2e:kubernetes:workspace-operator mise tasks into the branch-e2e workflow so they run alongside the other core Kubernetes e2e suites. Both are gated by run_core_e2e and included in the Core E2E result gate. Signed-off-by: Derek Carr <decarr@redhat.com> * test(k8s): add e2e tests for workspace namespace modes Add 7 new e2e tests covering workspace namespace lifecycle, TLS secret copying, ownership conflict detection, DNS-1123 validation, operator namespace preservation, and dynamic label watcher behavior. Fix async sandbox deletion race condition in existing tests by polling sandbox list instead of asserting immediately after delete. Signed-off-by: Derek Carr <decarr@redhat.com> * fix(k8s): grant secrets/patch unconditionally and backfill gateway-id labels Address two review findings: 1. RBAC: server-side apply (PATCH) is used for TLS secret sync in multi-namespace modes, but the ClusterRole only granted patch when the kubernetes-secrets credential driver was enabled. Grant patch unconditionally for non-shared modes since TLS sync always needs it; keep delete gated on the credential driver. 2. Upgrade safety: the new gateway-id label selector would orphan legacy Sandbox CRs that predate its introduction. Add a startup backfill in shared mode that patches any managed Sandbox CR missing the gateway-id label before the driver begins serving requests. Signed-off-by: Derek Carr <decarr@redhat.com> * fix(k8s): address workspace namespace review findings Signed-off-by: Derek Carr <decarr@redhat.com> * fix(k8s): address follow-up review findings Signed-off-by: Derek Carr <decarr@redhat.com> * fix(k8s): preserve workspace lookup after rebase Signed-off-by: Derek Carr <decarr@redhat.com> * fix(helm): allow managed secret creation Signed-off-by: Derek Carr <decarr@redhat.com> * fix(k8s): stop pods in workspace namespace Signed-off-by: Derek Carr <decarr@redhat.com> * test(k8s): scope pod deletion check to v1alpha1 Signed-off-by: Derek Carr <decarr@redhat.com> * fix(k8s): address workspace namespace review findings Signed-off-by: Derek Carr <decarr@redhat.com> --------- Signed-off-by: Derek Carr <decarr@redhat.com> |
||
|
|
bdabb54cb3 |
fix(security): authenticate extension services (#2638)
* fix(security): authenticate extension services Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * feat(extension-core): verify gateway JWTs Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * refactor(extension-core): keep inbound verification external This should become an extension SDK package. Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(security): harden the extension authentication contract Follow-up hardening on the alpha extension authentication mechanism. Claim contract: - Extension tokens carry an explicit `typ` of `openshell-ext+jwt`. They share a signing key with sandbox-to-gateway admission tokens and were otherwise separated by audience alone, so a verifier that neglects to check `aud` could accept a gateway credential. The header is a second, independent discriminator. - Publish OIDC-shaped discovery at `/.well-known/openid-configuration` so a service configured with only the gateway URL can learn the exact expected issuer and the JWKS location. It is shaped, not compliant: `issuer` is the gateway identity, not the serving URL. Audience agreement: - `MiddlewareManifest` and `InterceptorManifest` gain `expected_audience`. The audience is otherwise configured independently on each side of the boundary, where a mismatch surfaces only as an opaque authentication failure on every call. OpenShell now compares the two and fails at startup. An empty field keeps the check off for existing services. Compatibility: - Add `allow_insecure_transport` per registration. Enabling gateway JWT signing previously made any plaintext endpoint a hard startup failure, including the endpoint form used in our own documentation. The opt-out attaches no credential, is refused by the gateway if a supervisor asks for one, and warns at every startup. - Make the transport requirement kind-aware. A middleware endpoint must be reachable from every sandbox supervisor, so only interceptors may use a gateway-local Unix socket. Credential lifecycle: - Replace the process-global slot map with a supervisor-owned `ExtensionCredentialStore` shared explicitly across the gateway connections the supervisor opens, removing test-order coupling. - Rotate only when a credential is missing or has passed four fifths of its lifetime. Configuration polling ran every ten seconds against fifteen-minute credentials, so each poll re-ran gateway effective-policy resolution and re-minted the gateway token. - Bound credential minting per sandbox, since each request resolves the caller's effective policy. Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs: record alpha extension authentication in RFC appendices Restore the RFC 0009 and 0010 bodies to their accepted text and move every extension-authentication update into appendices instead. An RFC records a decision at a point in time; superseding detail belongs alongside it rather than rewritten into it. RFC 0009's appendix carries the shared contract: claims, authorization, key distribution, the `allow_insecure_transport` replacement for the body's `allow_insecure`, and residual risks. RFC 0010's records only what differs for interceptors and links to it. The existing protocol-extensions appendix, which parked the phase 2 transport question, now points forward to what was built. Also document the audience handshake, the discovery endpoint, the `typ` requirement, and `jti` replay guidance in the extensibility and gateway configuration pages, and correct the middleware transport guidance: middleware endpoints must be reachable from sandbox supervisors, so Unix sockets are not an option there. Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * refactor(extension-core): abstract extension server trust Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(core): update middleware manifest example Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(extension-auth): preserve unsigned gateway compatibility Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(extension-auth): reject cross-domain token replay Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> --------- Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> |
||
|
|
f12f3ef8d5 |
fix(macos): restore Homebrew sandbox callbacks (#2739)
* fix(macos): restore Docker gateway callbacks Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(gateway): reuse reachable primary callback listener Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
d0c6dc3fd8 |
feat(kubernetes): support corporate upstream proxy (#2633)
* feat(kubernetes): support corporate upstream proxy Signed-off-by: loveRhythm1990 <qiuweimin@126.com> * fix(kubernetes): reject proxy_auth_secret_key values Kubernetes cannot create Gateway validation accepted proxy_auth_secret_key values that Kubernetes rejects when creating the Secret (keys longer than 253 bytes, or the reserved "."/".." names), turning an invalid deployment setting into repeated sandbox Pod-provisioning failures instead of a startup error. Reject them in validate_upstream_proxy_config so they fail closed at gateway startup. Signed-off-by: loveRhythm1990 <qiuweimin@126.com> * docs(skill): add corporate upstream proxy checks to debug-openshell-cluster Add a Kubernetes corporate upstream proxy troubleshooting section covering rendered [openshell.drivers.kubernetes] configuration, credential Secret volume events, supervisor arguments and mounts confined to the network- supervising container, and proxy reachability. Signed-off-by: loveRhythm1990 <qiuweimin@126.com> --------- Signed-off-by: loveRhythm1990 <qiuweimin@126.com> |
||
|
|
c4b500a7de |
feat(helm): cert-manager external issuer + OpenShift passthrough Route (#2468)
* fix(core): trust public root CAs alongside the sandbox mTLS CA The supervisor gRPC client only trusted the CA configured via OPENSHELL_TLS_CA, since tonic ClientTlsConfig starts with an empty root store unless with_native_roots()/with_webpki_roots() is also enabled. Deployments where the gateway server certificate is issued by a public CA (e.g. cert-manager against an ACME issuer) caused every supervisor connection to fail the TLS handshake with "UnknownCA", since the sandbox mTLS CA and the server cert issuer were no longer the same. Enable both native and webpki roots in addition to the configured CA. tonic root store is a union of all configured sources, so this does not weaken verification for existing self-signed deployments. webpki-roots (compiled in) is enabled alongside native-roots since the supervisor binary may run in minimal sandbox images without a populated system CA bundle. Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * feat(helm): support external cert-manager issuers and OpenShift Route passthrough Add certManager.serverIssuerRef/clientIssuerRef so the gateway and mTLS client certificates can be issued by a real Issuer/ClusterIssuer (e.g. ACME) instead of only the chart built-in self-signed CA. Add openshiftRoute template for exposing the gateway via a TLS passthrough Route so the gateway keeps terminating its own TLS/mTLS. The server Certificate excludes internal-only SANs (cluster-local, localhost, loopback) when an external issuer is configured, since ACME issuers reject those per CA/Browser Forum baseline requirements. A template-time fail guard catches the misconfiguration at helm install time rather than asynchronously at cert-manager issuance time. Includes Helm unittest coverage for both issuerRef overrides and Route rendering, plus a CI values overlay for lint coverage. Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * docs: document cert-manager external issuer and OpenShift Route Update managing-certificates.mdx with the serverIssuerRef workflow and install-time validation behavior. Add a production section to the OpenShift guide covering passthrough Route with a real certificate. Regenerate Helm README for new certManager and openshiftRoute values. Sync debug-openshell-cluster skill with new troubleshooting steps for ACME issuance failures and supervisor UnknownCA from mismatched CAs. Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(helm,core): address PR review feedback on cert-manager external issuer Addresses all five blocking review items from #2468: 1. Remove .with_native_roots() from supervisor gRPC client -- the supervisor runs inside the user-selected sandbox image, so the image CA bundle is not operator-controlled. Keep .with_webpki_roots() (compiled-in, not user-controlled) alongside the configured CA. 2. Fail at render time when serverIssuerRef.name is set but clientCaFromServerTlsSecret is still true. Add negative Helm test. 3. Remove clientIssuerRef -- changing only clientIssuerRef breaks both directions because trust bundles are not modeled separately. Change serverIssuerRef.kind default from ClusterIssuer to Issuer. 4. Add server.oidc.issuer and server.oidc.audience to the documented OpenShift production Helm command. Add Access Control prerequisite. 5. Fail at render time when openshiftRoute.enabled and disableTls are both true. Add negative Helm test. Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(drivers): strip GATEWAY_TLS_SERVER_NAME from Docker and Podman env Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(helm,drivers): guard default clientCaSecretName and add env-strip tests Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * feat(tls): SNI-based dual certificate for internal and external server TLS Split the gateway server certificate into two: an internal cert issued by the chart's own CA (for supervisor connections via cluster-local SANs) and an external cert issued by an operator-configured Issuer such as ACME/Let's Encrypt (for CLI and Route access via public SANs). The gateway uses SNI-based certificate selection: connections whose SNI hostname matches external_server_names receive the external cert; all others (including those with no SNI) receive the internal cert. Security improvement: remove .with_webpki_roots() from the supervisor gRPC client so supervisors trust only the chart CA, closing a MITM vector via publicly-trusted certificates in user-supplied container images. Key changes: - Add DualCertResolver with SNI-based cert selection and full test coverage - Add external_cert_path, external_key_path, external_server_names to TlsConfig - Validate partial external cert config (error on cert-without-key or vice versa) - Validate empty external_server_names when external cert is configured - Split cert-manager templates into internal + external Certificate resources - Add Helm guards for misconfigured external issuer (empty serverDnsNames, internal-only SANs with external issuer, conflicting clientCaFromServerTlsSecret) - Update gateway-config.mdx, managing-certificates.mdx, openshift.mdx docs - Update debug-openshell-cluster skill for dual-cert troubleshooting Signed-off-by: Pi Agent <agent@openshell.local> * fix(drivers): strip GATEWAY_TLS_SERVER_NAME in VM driver and correct comments Add the same GATEWAY_TLS_SERVER_NAME environment stripping to the VM compute driver that Docker, Podman, and Kubernetes drivers already perform. Without this, a sandbox user on the VM driver could override the TLS server name the supervisor verifies. Fix stale comments in Docker and Podman drivers that referenced 'with WebPKI roots trusted' — WebPKI roots are explicitly not trusted after the tls-webpki-roots removal. Use tls-ring instead of bare channel for tonic in openshell-core so the TLS API (ClientTlsConfig, Endpoint::tls_config) is available without pulling in any root certificate store. Signed-off-by: Pi Agent <agent@openshell.local> * fix(tls,helm): wildcard SNI matching and Route host validation Add RFC 6125 single-level wildcard matching to DualCertResolver so external_server_names entries like *.example.com correctly match SNI hostnames like gw.example.com. Previously only exact matches worked, silently falling back to the internal cert for wildcard configurations. Add a Helm fail guard in route.yaml that rejects openshiftRoute.host values not listed in certManager.serverDnsNames when an external issuer is configured — catches cert/route hostname mismatches at install time instead of at TLS connect time. Quote the host field in route.yaml for robustness. Signed-off-by: Pi Agent <agent@openshell.local> * fix(helm): address blocking review items — client-CA guard, wildcard Route, serverIssuerRef gate 1. Remove the obsolete guard rejecting serverIssuerRef + clientCaFromServerTlsSecret=true. The internal server certificate is always signed by the chart CA (the same CA that signs the client cert), so clientCaFromServerTlsSecret=true is correct — its filtered ca.crt is exactly the right trust anchor. The old workaround (mounting openshell-ca-tls directly) unnecessarily exposed the CA private key to the gateway container. Remove the client-CA overrides from docs, CI overlay, and production examples. 2. Route host validation now supports wildcard certificates per RFC 6125: single-level wildcards like *.example.com match gateway.example.com but not deep.sub.example.com. Require an explicit openshiftRoute.host when an external issuer is configured — without one, OpenShift generates a hostname absent from serverDnsNames. 3. Reject serverIssuerRef.name when certManager.enabled is false — the external certificate, its Secret mount, and the gateway TLS config all require cert-manager to be enabled. Validated on ROSA (dev.dyee.p3) with branch-built images: - Fresh install with letsencrypt-prod ClusterIssuer - SNI dual-cert: external hostname served Let's Encrypt cert - Supervisor mTLS via internal cert path: ConnectSupervisor accepted - Client CA volume: filtered ca.crt from internal server secret (no key) - CLI connected via Route + OIDC Helm tests: 81 pass across 7 suites. Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> --------- Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> Signed-off-by: Pi Agent <agent@openshell.local> |
||
|
|
35fb27ef14 |
feat(sdk): add TypeScript SDK (@nvidia/openshell-sdk) (#2122)
* feat(sdk): add TypeScript SDK (@nvidia/openshell-sdk) First native, per-language SDK for the OpenShell gateway: a thin, idiomatic TypeScript client over proto-generated gRPC stubs (connect-es), no FFI. Covers the v0.1 surface — sandbox lifecycle (create/get/list/delete + waitReady/ waitDeleted), health, and streamed exec. - sdk/typescript/: package, client/transport/errors, protoc + protoc-gen-es codegen (gen/ gitignored, absorbed into dist/ at build), committed lockfile. - tasks/typescript.toml: sdk:ts install/proto/typecheck/build/ci/publish; sdk:ts:typecheck wired into `check`; sdk-typescript job in branch-checks (typecheck, build, and a --dry-run publish that validates the release path). - Enforce SPDX headers on .ts/.tsx/.mts/.cts (skip node_modules and gen/); back-fill docs/_components/jsx.d.ts and fern/components/CustomFooter.tsx. - release.py gains an npm version format; release-tag.yml publishes to GitHub Packages on tag, stamping the version (0.0.0 placeholder in git); prerelease builds publish under the `next` dist-tag, not `latest`. Ships as @nvidia/openshell-sdk on GitHub Packages pre-GA; public npm (@openshell/sdk) follows at GA with an unchanged public API. Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com> * chore(sdk): adopt TypeScript 6, tidy @types/node range - typescript ^5.7.2 -> ^6.0.3 (6.0 is now `latest`; the old caret capped at 5.x) - @types/node ^24.0.0 -> ^24 (same range, tidier) No source changes; codegen, typecheck, and build pass on 6.0.3. Verified the emitted d.ts still type-check for downstream consumers on TypeScript 5.0.4 through 5.9.3, so this does not raise the SDK's consumer TS floor. Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com> * refactor(sdk): group operations under a composable SandboxClient Reshape the client from flat methods (createSandbox, listSandboxes, exec) to a scoped SandboxClient reached as `client.sandbox.create/get/list/delete/exec` (+ waitReady/waitDeleted), mirroring the CLI's noun-verb model and the Python SDK's SandboxClient. SandboxClient is also usable standalone via SandboxClient.connect(); OpenShellClient composes it over a single shared transport, so future service/provider clients reuse one connection. health() stays top-level as a gateway call. No behavior change; types are unchanged. Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com> * chore(sdk): generate TypeScript SDK stubs with buf Replace the protoc gen.sh with `buf generate` + buf.gen.yaml. `buf` (@bufbuild/buf) is a package devDependency and self-compiles the protos, so the TS SDK no longer depends on the mise-pinned protoc; it drives the same connect-es plugin. Generation stays limited to the client-surface closure (openshell/sandbox/datamodel) via the input paths. Output is byte-identical to the previous protoc + protoc-gen-es pipeline. Lays the groundwork for a shared buf.yaml (lint/breaking/LSP) as a follow-up. Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com> * build(proto): add repo-level buf module with lint Declare proto/ as a single buf v2 module in a root buf.yaml so buf generate, lint, breaking, and the editor LSP resolve imports the same way. Lint uses STANDARD with six documented exceptions for deviations the current protos intentionally make: the flat proto/ layout with nested packages (DIRECTORY_SAME_PACKAGE, PACKAGE_DIRECTORY_MATCH) and the established API shape with unsuffixed services and reused request/response messages (RPC_REQUEST_RESPONSE_UNIQUE, RPC_REQUEST_STANDARD_NAME, RPC_RESPONSE_STANDARD_NAME, SERVICE_SUFFIX). Every other STANDARD rule now enforces on future protos. Breaking uses FILE. Code generation stays package-scoped in sdk/typescript/buf.gen.yaml since it binds to that package's connect-es plugin and output dir; its inputs are unchanged and regeneration is byte-identical. Wire the check in via a proto:lint mise task that runs buf from the SDK devDependencies. It is a dependency of both sdk:ts:ci (so the TypeScript SDK CI job enforces it) and the top-level lint aggregate (so local pre-commit covers it). Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com> * chore(sdk): publish as unscoped openshell-sdk on public npm Rename the package from @nvidia/openshell-sdk to the unscoped openshell-sdk and target public npm (registry.npmjs.org) instead of GitHub Packages. GitHub Packages requires a scope matching the owning org, and the @openshell scope is blocked by an unrelated existing package, so an unscoped name on public npm is the lowest-friction distribution path and needs no org approval. Rework the release-tag publish job to auth against registry.npmjs.org with NPM_TOKEN (the job now only needs packages: read to pull the CI image). Update the README install instructions and usage imports. Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com> * chore(sdk): publish @nvidia/openshell-sdk to GitHub Packages Revert the unscoped-name switch. GitHub Packages only accepts scoped names matching the owning org, so shipping there first (which needs no external npm org or NPM_TOKEN, just the repo's GITHUB_TOKEN) requires the @nvidia scope. Keeping the @nvidia/openshell-sdk name also lets a later public-npm release use the same install specifier, so adding public npm becomes a second publish step rather than a rename. Restore the GitHub Packages publish auth in the release-tag job and the scoped install instructions in the README (keeping the buf codegen note). Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com> * feat(sdk-ts): add streaming exec, forward, ssh, provider, and config methods Grow SandboxClient to the surface the first two consumers need. execStream yields stdout/stderr chunks as they arrive and exec now drains it, keeping its buffered ExecResult and signature unchanged. execInteractive is the TTY + stdin transport primitive (start-first framing, output/write/resize/close/done, no terminal glue). forward binds a local TCP listener that tunnels each accepted connection into the sandbox for the process lifetime, minting and revoking a per-socket SSH session token around a forwardTcp bidi. Adds createSshSession / revokeSshSession, attach/detach/listProviders, and getConfig / setPolicy / setSetting (sandbox-scoped, network-policy-only, with an optional wait poll). Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com> * build(sdk-ts): add Biome and Vitest tooling The TypeScript SDK had no formatter or linter and no test runner. Add Biome (format + lint, generated src/gen excluded) enforcing 2-space indent, single quotes, semicolons, and a 120-column width, and reformat the existing hand-written sources accordingly. Add Vitest for unit tests. Wire sdk:ts:format, sdk:ts:lint, and sdk:ts:test mise tasks into the fmt/lint aggregates, the root test suite, and sdk:ts:ci so they run in CI. Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com> * test(sdk-ts): cover the sandbox surface with in-memory transport tests Exercise SandboxClient against an in-memory OpenShell service built with createRouterTransport: request assembly and id resolution, u64/int64 rendered as strings, enum lowercasing, fromConnect code mapping, the exec/execStream drain plus a backward-compat check on exec, execInteractive start-first ordering and done resolution, and a forward() byte relay against a loopback echo with close() teardown. Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com> * docs(sdk-ts): document the new surface and connect/upload/download boundaries Document execStream, execInteractive, forward, ssh sessions, providers, and config/policy in the SDK README, and record the intentional boundaries: interactive connect / PTY ownership, upload/download (no file-transfer RPC), and detached forwards stay out of scope. Note the Biome/Vitest dev commands. Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com> * feat(sdk-ts): support mTLS client authentication Add clientCert and clientKey to ConnectOptions so the SDK can authenticate to the default local gateway, which uses mTLS user authentication. Without a client certificate and key the SDK could verify the server but never authenticate the caller, so it could not connect to the standard Docker, VM, Homebrew, or Linux-package gateway. Validate the pair as both-or-neither and pass cert and key through to the Node TLS options for https gateways. The h2c path is unchanged. Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com> * chore(sdk-ts): drop the demo script and its tsx dependency Remove src/demo.ts, the demo npm script, the tsx devDependency, and the tsconfig build exclude for the demo. The demo was never part of the published package, and dropping it also removes the only place that logged part of an SSH session token. Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com> * fix(sdk-ts)!: harden exec streaming, waits, SSH, and forwarding Address review feedback on the sandbox surface. - Make the streamed command exit code observable from idiomatic for-await: the terminal exit is now an in-band ExecStreamEvent ({ type: 'exit', exitCode }) rather than the async generator return value, which for-await discards. A stream that ends without an exit event now throws instead of reporting success. - Bound waitReady, waitDeleted, and the setPolicy wait by their timeout: each poll RPC carries a per-iteration deadline and the waits accept an AbortSignal, so a stalled call can no longer leave a wait pending forever. Add waitTimeoutSecs to SetPolicyOptions. - Validate the CreateSshSession response against the proto charset and range contract before returning it or using its token, since the values feed an OpenSSH ProxyCommand. - Respect socket backpressure when relaying forwarded responses: pause reading the gRPC stream when the local socket buffer is full and resume on drain so memory stays bounded. - Expose create-time sandbox policy: add policy and an advanced rawSpec passthrough to SandboxSpec so the safety boundary is expressible at creation and new spec fields do not require an SDK change. BREAKING CHANGE: execStream and the interactive exec output now yield a terminal { type: 'exit', exitCode } event; consumers iterating the stream must handle that arm. The exit code is no longer the async generator return value. Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com> * feat(sdk-ts): export error contract, enum unions, and caller cancellation Address Tier-1 review feedback on the TypeScript SDK public surface (PR #2122). - Errors: export SdkError and SdkErrorCode so callers can use instanceof and exhaustively switch on .code. fromConnect preserves the originating ConnectError as .cause and its status as .connectCode, maps Aborted to a new 'aborted' code for optimistic-concurrency conflicts, and maps Canceled and DeadlineExceeded to 'canceled'. errorCode() behavior is unchanged. - Enums: replace the string-typed phase, status, scope, and policySource fields with lowercase literal unions (SandboxPhaseName, HealthStatus, SettingScopeName, PolicySourceName) backed by exhaustive Record maps. The unions are a hand-maintained mirror of the generated proto enums; a new drift test pins each literal to its generated member name. - Cancellation: accept an optional AbortSignal on exec, execInteractive, and forward, threaded into both sandbox resolution and the streaming RPC. forward tears down its local listener on abort. Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com> * feat(sdk-ts): add raw escape hatch for uncurated gateway RPCs The curated sub-clients reduce proto messages to ergonomic subsets (for example get() drops created_at_ms, the full spec, conditions, runtime endpoints, and current_policy_version), and not every gateway RPC has a typed helper yet. Rather than ship methods that exist but throw, expose a generated client for the full surface. OpenShellClient.raw and SandboxClient.raw are generated clients covering every gateway RPC, returning the verbatim wire messages so proto distinctions the curated types smooth over are preserved. .transport exposes the shared connection for building extra clients over one socket. Generated request/response types are published at the new @nvidia/openshell-sdk/raw subpath. Curated methods stay the default; raw is the always-available floor. Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com> * fix(sdk-ts): address review feedback on exec, forward, and auth transport Settle exec's `done` promise before yielding the exit event so a consumer that breaks on exit no longer leaves it pending forever, and give it a lone rejection handler plus a finally-settle so a stream error or early abandon can never surface as an unhandled rejection or a hang. Attach an 'error' listener to each accepted forward socket synchronously, before forwardConnection awaits CreateSshSession; a peer reset in that window previously emitted an unhandled 'error' and crashed the process. Reject ambiguous or unsafe transport configs at buildTransport: oidcToken and edgeToken together (silently OIDC-only), and any auth token sent over plaintext http:// to a non-loopback host unless allowInsecureAuth is set. Wrap versionPin so a non-u64 expectedResourceVersion raises SdkError('invalid_config') instead of a raw BigInt SyntaxError, and raise the Node engine floor to >=20.3 for AbortSignal.any(). Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com> * fix(sdk-ts): address review feedback Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(sdk-ts): defer published sdk guide Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com> Signed-off-by: Drew Newberry <anewberry@nvidia.com> Co-authored-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
0f8fad23c4 |
feat(sandbox): add stop and start operations (#2653)
* feat(sandbox): add suspend and resume operations Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(server): preserve lifecycle work after cancellation Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(server): reconcile ambiguous lifecycle outcomes Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(server): complete suspended session cleanup Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(vm): preserve suspension state on resume failure Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(server): retry retained lifecycle transitions Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(server): clean sessions after suspend reconciliation Signed-off-by: Seth Jennings <sjenning@redhat.com> * test(sandbox): cover deleting suspended sandbox Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(kubernetes): preserve progressing sandbox suspension Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(kubernetes): bound suspend status polling Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(kubernetes): detect legacy sandbox suspension Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(tui): render suspended sandbox phases Signed-off-by: Seth Jennings <sjenning@redhat.com> * refactor(sandbox): rename suspend and resume lifecycle Signed-off-by: Seth Jennings <sjenning@redhat.com> * perf(server): clean stopped sessions on transition Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(kubernetes): fail fast on rejected stop Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(compute): fence stale restart lifecycle events Signed-off-by: Seth Jennings <sjenning@redhat.com> --------- Signed-off-by: Seth Jennings <sjenning@redhat.com> |
||
|
|
dd2b4e3bc0 |
feat(cli): warn when --env values look like credentials (#2655)
* feat(cli): add credential env match validation Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * feat(cli): warn when --env values look like credentials Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * docs(sandbox): add flag --no-credential-warnings details + polishing Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * fix(cli): match credential keywords on underscore segments Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> --------- Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> |
||
|
|
170961997f |
chore(ci): disable telemetry in internal test runs (#2648)
* chore(ci): disable telemetry in internal test runs Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> * test(ci): remove brittle telemetry wiring test Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> * docs: trim CI telemetry guidance Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> * test(e2e): share telemetry default with OpenShift Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> --------- Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> |
||
|
|
0120535efc |
feat(proxy): bind static credentials to provider endpoints (#2510)
* feat(proxy): bind static credentials to provider endpoints Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * test(e2e): verify static credential endpoint isolation Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * docs(provider): explain static credential endpoint binding Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * fix(e2e): use valid endpoint isolation fixtures Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * docs(provider): explain static credential endpoint binding Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * fix(credentials): preserve binding identity across rotations Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(proxy): enforce bindings across request lifecycle Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(proxy): close credential relay gaps Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * docs(credentials): clarify binding failure behavior Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(credentials): hash selected provider profile scope Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(proxy): resolve credentials after request admission Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * docs(credentials): clarify binding failure diagnostics Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(proxy): align single-route credential denials Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(credentials): harden endpoint-bound rotation Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(credentials): enforce identity and authority binding Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(credentials): snapshot provider environment atomically Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(e2e): include authority port in query proxy requests Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(credentials): close credential revocation gaps Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * docs(proxy): explain authority mismatch diagnostics Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(credentials): enforce binding lifecycle invariants Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(provider): reject credential config collisions Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(network): capture credential scope atomically Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(network): distinguish origin and absolute targets Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(provider): isolate endpointless profile credentials Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(network): normalize IPv6 request authorities Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * docs(credentials): clarify endpointless profile isolation Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * feat(policy): bind endpointless provider credentials Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(credentials): use current GCP placeholder revision Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * docs(providers): explain policy credential bindings Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(credentials): cover endpointless fail-closed invariant Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(policy): expect ambiguity rejection at creation Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(server): authenticate rebased policy requests Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * refactor(proxy): share credential mismatch finding builder Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(credentials): cover malformed binding metadata Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(credentials): verify multi-key endpoint isolation Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(e2e): cover same-host credential path denial Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * docs(credentials): document serialized refresh contract Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * refactor(proxy): consolidate L7 log formatting Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * perf(credentials): precompile endpoint binding patterns Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * perf(credentials): share identity epoch revisions Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(proxy): require explicit request default ports Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(policy): validate SigV4 credential sources Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(credentials): preserve endpoint bindings for credential handles Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * feat(go-sdk): expose network credential bindings Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> --------- Signed-off-by: John Myers <johntmyers@users.noreply.github.com> Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> Co-authored-by: John Myers <johntmyers@users.noreply.github.com> |
||
|
|
5e2f0d1b37 |
fix(policy): prevent implicit authorization inheritance (#2499)
* fix(policy): prevent implicit authorization inheritance A network rule authorizes every listed binary to reach every listed endpoint, so unioning an AddRule operation's binaries and endpoints independently grants binary-by-endpoint pairs the operation never declared. Require AddRule to declare the complete product before merging, and reject an operation that would give one host and port two different MCP inspection contracts. The rejection names the binaries the operation still has to declare. Fix proposal coverage on the same surface. An any-binary proposal was vacuously covered by a binary-restricted loaded rule, and a complete product split across several loaded rules was reported as uncovered. Coverage compares merge-widened endpoint fields by containment and exact-matches only the fields the merge never widens, so a policy the gateway just merged always reads back as covered and the sandbox policy.local /wait long-poll cannot spin to its deadline. Signed-off-by: Shiju <shiju@nvidia.com> * fix(policy): close authorization-inheritance gaps in merge and coverage Coverage treated an unset proposal value for a field the merge retains as a request for the default, so a proposal that merged cleanly into an endpoint carrying enforcement, protocol, or tls read back as uncovered and left the policy.local /wait long poll spinning. Unset now means unspecified. Ports are a set on the wire but each port is an independent authorization. Coverage and the inheritance check both resolve one binary, host, path, and port at a time, so ports spread across loaded rules resolve and a complete declaration split across incoming endpoints is accepted. An incoming empty binary list means any binary. It now has to declare every merged endpoint like a new concrete path does, and once declared the promotion is applied instead of appending an empty list and leaving the restricted scope in place. The endpoint-overlap fallback folded a new rule name into an existing rule where inheritance validation then rejected it, leaving no way to grant a binary part of a rule. Folding now keeps the requested rule name when it would widen, and reports that it did. An MCP contract conflict still propagates because one host and port carry a single inspection contract. AddAllowRules and AddDenyRules select an endpoint by host and port alone and now reject a target that resolves more than once, including two paths on one rule. RemoveBinary rejects an any-binary rule rather than reporting a success that leaves the binary authorized. Signed-off-by: Shiju <shiju@nvidia.com> * fix(policy): gate port and MCP-contract widening independently of binary scope The changed-port check only rejected when the operation also left an existing binary undeclared. An operation listing every existing binary could therefore declare one port of a multi-port endpoint and have its widened fields land on the endpoint the merge shares across all of them, authorizing L7 rules on a port it never named. Each changed port must now be named by the operation on its own, independently of binary-scope coverage. MCP contract compatibility was checked inside the endpoint fold, which only compares endpoints agreeing on host, path, and a shared port. The sandbox resolves one extended configuration per host and port and never consults the path, so a second MCP endpoint under another path, or in another rule, left the effective strict-tool-name, method-profile, and body-limit contract decided by match order. Contract agreement is now enforced across the whole merged policy, including provider-composed rules. A conflict already present in the baseline is left alone so unrelated updates still apply. An empty binary list authorizes any binary. Appending an incoming named list made it non-empty and revoked every process the operation did not name, turning an additive update into a silent mass revocation. An already-empty scope is now kept and reported; only a restricted scope is replaced by an incoming any-binary scope. Warnings raised during a fold now name the rule that was actually modified rather than the rule name the operation requested, which differ when the endpoint-overlap fallback redirects the operation. Signed-off-by: Shiju <shiju@nvidia.com> * fix(policy): route undeclared-port conflicts through the separate-rule fallback The fold-only classifier decides which merge errors disappear when the incoming authorization stays on its own rule. UndeclaredPortWouldChange was added ahead of the existing-binary conflict but never classified, so a differently named narrow update against a multi-port endpoint failed outright instead of landing separately. A same-key update still returns the error, because there the operation chose the target. The classifier is now an exhaustive match rather than a matches! with an implicit false. A new variant defaulting to "not fold-only" is what withdrew the separate-rule remedy here, so adding one has to be an explicit decision. Inspection-contract agreement now covers protocol, not only MCP options. The sandbox resolves one extended configuration per host and port and never consults the path, so an MCP endpoint and a REST endpoint on the same host and port left the effective inspection protocol decided by match order. Endpoints with no protocol carry no contract and are skipped. Signed-off-by: Shiju <shiju@nvidia.com> * fix(policy): compare only MCP contracts when detecting endpoint conflicts The post-merge conflict scan was broadened to compare inspection protocol as well as MCP options on one host and port. The supervisor selects among matching endpoint configs by most-specific path, so a broad REST endpoint and a narrower GraphQL endpoint on the same host and port are unambiguous and supported. The broader comparison rejected those updates even though the equivalent full policy loads and serves correctly. MCP options are not selected that way, so the scan keeps comparing them: two MCP endpoints on one host and port still have to agree on strict tool names, method profile, and body limit, whatever paths or rules hold them. The policy page returns to describing the MCP-specific rule and the path-aware selection it sits alongside. Signed-off-by: Shiju <shiju@nvidia.com> * fix(policy): keep MCP off a host and port shared with other inspection Narrowing the conflict scan back to MCP options let an MCP endpoint sit under a path already covered by a broader REST endpoint. The supervisor picks the parser by most-specific path, so the MCP endpoint parses the request, but _policy_allows_l7 is existential over every endpoint matching it. A plain REST rule on the overlapping path can therefore make allow_request true for a JSON-RPC tool call the MCP endpoint never allowed, and the relay forwards it. The scan now records every inspected protocol on a host and port. MCP may not share one with a differently inspected endpoint, and two MCP endpoints there still have to agree on one contract. Endpoints that are not inspected carry no contract and never compete, and two non-MCP endpoints stay supported because they share one method-and-path rule vocabulary. Signed-off-by: Shiju <shiju@nvidia.com> --------- Signed-off-by: Shiju <shiju@nvidia.com> |
||
|
|
4cb77a900e |
fix(e2e): separate Podman Machine loopback listeners (#2622)
* fix(e2e): separate Podman Machine loopback listeners Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> * test(e2e): remove shallow harness checks Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> * refactor(e2e): trim Podman listener workaround Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> * fix(e2e): bypass proxies for Podman health probe Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> --------- Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> |
||
|
|
284da54de5 |
docs(readme): add theme-aware banner (#2619)
* docs(readme): add theme-aware banner Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(readme): exclude preview screenshot from tree Signed-off-by: Johnny Greco <jogreco@nvidia.com> --------- Signed-off-by: Johnny Greco <jogreco@nvidia.com> |
||
|
|
5548405fcb |
feat(credentials): add provider credential storage drivers (#2437)
* feat(credentials): add provider credential storage drivers Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(credentials): harden credential update handling Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(credentials): harden credential driver security, correctness, and performance Address review findings from the credential storage drivers PR: - Route additional_credentials through the driver on refresh to prevent silent data loss for multi-credential providers (e.g. AWS STS) - Clean up stored credential handles on CAS failure during refresh to prevent orphaned secrets in external backends - Enforce namespace validation in the Kubernetes Secrets driver to prevent cross-namespace credential access when allow_reference_namespace is not enabled - Cache Vault Kubernetes auth tokens with 80% TTL to avoid re-authenticating on every credential operation - Parallelize resolve_credentials in all three drivers using try_join_all for faster sandbox startup - Add existingSecret support for the KEK Secret to fix helm template/GitOps workflows where lookup returns empty and regenerates the key - Document RBAC blast radius for the Kubernetes Secrets credential driver and recommend a dedicated namespace Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com> Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> * fix(credentials): add optimistic concurrency, fix thundering herd, parallelize operations Use resourceVersion optimistic concurrency with retry loop for K8s Secret ownership checks to prevent TOCTOU races. Switch Vault token cache from RwLock to Mutex with double-check pattern to prevent thundering herd on cache miss. Parallelize credential store and delete operations across independent keys using try_join_all. Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com> Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> * fix(credentials): handle partial failures, add delete retry, consolidate cleanup Replace try_join_all with join_all in credential store/delete operations to handle partial failures — successfully-stored handles are cleaned up when another key fails. Add retry loop with conflict detection to db-credstore delete_credential, matching the K8s driver pattern. Consolidate 4 manual cleanup_pre_stored_provider_credentials call sites into a single error handler using an async block. Remove inconsistent .trim() from db-credstore validate_handle_owner. Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com> Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> * fix(credentials): fix retry loop guard and remove unprotected validation Remove attempt-count guard from 409/Aborted match arms in retry loops so the post-loop Status::aborted error is reachable after exhausting retries. Previously, last-attempt conflicts fell through to the catch-all error arm, producing misleading Status::unavailable errors. Remove duplicate validation calls that ran after prepare_provider_credential_update but outside the cleanup-protected async block, which would leak pre-stored handles on failure. Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com> Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> * fix(credentials): add workspace/provider UUID to credential backend paths Include workspace and provider ID in credential backend object paths to ensure cross-workspace uniqueness and prevent credential collision (GATOR-1806c9be-01). - Updated credential driver proto to include workspace and provider_id fields - Modified Vault driver to include workspace/provider_id in managed_secret_path - Modified Kubernetes Secrets driver to include workspace/provider_id in credential_owner_id and managed_secret_name - Updated all credential runtime calls to pass workspace/provider_id - Updated tests to use the new signatures This prevents two workspaces sharing the same external credential store from colliding on provider names, which was a critical security issue (CWE-639). * fix(credentials): preserve provider-level expiration for handle-backed credentials Compute effective expiration from both provider and driver values using the earliest non-zero timestamp and skip expired values before insertion (GATOR-1806c9be-02). - Modified resolve_provider_handles to check provider credential_expires_at_ms - Skip expired credentials during resolution instead of returning them - Use effective expiration (min of provider and driver) in resolution results - Fix inference.rs to preserve earliest expiration when merging This ensures handle-backed credentials respect the same expiration semantics as inline credentials. * fix(credentials): stage refresh changes under new handles before validation Stage credential replacements under new immutable handles instead of reusing existing handles to prevent overwriting committed values before validation/CAS (GATOR-1806c9be-03). - Stage credentials with empty existing_handles map to force new handle creation - Validate and CAS before the new values are committed to backend storage - Delete old handles only after successful CAS - On CAS failure, delete only the newly staged handles - This prevents CWE-362/CWE-367 race conditions where failed refreshes could still modify or delete the active credential The fix ensures that a rejected refresh cannot modify the backend object still referenced by the committed provider record. * fix(credentials): add timeouts to credential driver RPCs Apply configured timeouts to both startup capability negotiation and runtime RPCs to prevent indefinite hangs (GATOR-1806c9be-05). - Add DEFAULT_CREDENTIAL_DRIVER_RPC_TIMEOUT_SECS constant (30s) - Apply timeout to GetCapabilities during startup connection - Apply timeout to all runtime RPCs (store, delete, resolve) - Use tokio::time::timeout to bound the entire GetCapabilities operation during startup, not just the socket connection - Return contextual deadline errors on timeout This prevents a faulty or overloaded driver from hanging gateway operations indefinitely. * fix(credentials): fix test to use consistent workspace/provider identity The Kubernetes auth Vault resolve test was constructing a managed path with test-workspace/test-provider-id but sending default/prov-123 in the request, causing validation to reject the request (GATOR-18e32351-01). - Update test to use test-workspace and test-provider-id in the request to match the logical_path construction - This ensures the test exercises the intended code path and validates Kubernetes auth resolution properly The test now passes and correctly validates identity enforcement. * fix(credentials): use unique staging ID for refresh to avoid overwrites Stage refresh replacements under genuinely distinct immutable handles using a unique staging ID to prevent overwriting committed values (GATOR-1806c9be-03). - Generate a unique staging ID using UUID for each refresh operation - Use this staging ID when storing credentials instead of the real provider ID - Pass the same staging ID during cleanup on failure to delete only staged objects - This ensures deterministic paths (Vault) and object names (K8s) don't collide with the committed provider's credentials The fix prevents failed refreshes from silently replacing active credentials or breaking providers by deleting still-referenced backend objects. * fix(credentials): wrap credential driver RPCs in local timeouts Add local tokio::time::timeout wrappers around credential driver RPCs to bound non-compliant or stalled UDS peers (GATOR-1806c9be-05). - Wrap StoreCredential, DeleteCredential, and ResolveCredentials in local timeouts - Return contextual deadline_exceeded errors when timeouts occur - Keep existing gRPC timeout metadata for compliant implementations - GetCapabilities during startup was already wrapped in previous commit This ensures a faulty local driver cannot hang gateway operations indefinitely, even if it accepts the connection but never responds to the RPC. * fix(credentials): preserve ownership for staged refreshes Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(credentials): bound startup capability probe Signed-off-by: Seth Jennings <sjenning@redhat.com> * test(provider): authenticate credential handler requests Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(ci): grant actions read to credential driver e2e Signed-off-by: Seth Jennings <sjenning@redhat.com> --------- Signed-off-by: Taylor Mutch <taylormutch@gmail.com> Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com> Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> Signed-off-by: Seth Jennings <sjenning@redhat.com> Co-authored-by: Taylor Mutch <taylormutch@gmail.com> Co-authored-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> |
||
|
|
537805568d |
feat(sandbox): honor OCI image working directories (#2530)
* feat(sandbox): honor Docker OCI working directories Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> * fix(sandbox): honor effective workspace access Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> * test(sandbox): cover enforced workspace denial Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> * docs(docker): explain effective workdir checks Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> * fix(sandbox): validate effective workspace writes Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> * fix(sandbox): reserve supervisor control roots Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> * refactor(sandbox): centralize control paths Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> * fix(sandbox): reserve OCI runtime mount roots Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> --------- Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> |
||
|
|
8328412959 |
feat(vm): export driver traces over OTLP (#2564)
Continue distributed traces across the gateway-to-driver process boundary and export VM driver spans to the same OTLP/gRPC collector. The driver reports as the distinct openshell-driver-vm service. Updated the gateway architecture and configuration reference with a generic external-driver forwarding contract. Instrumented: - Every RemoteComputeDriver RPC injects the active W3C trace context into tonic metadata. Managed VM readiness and runtime initialization give startup capability probes stable parent operations rather than isolated root spans. - A tonic service layer creates fixed, low-cardinality server spans for every ComputeDriver RPC. New handlers inherit tracing automatically; failures record OpenTelemetry error status and the gRPC status code. - Background provisioning remains attached to CreateSandbox after the RPC returns without extending the RPC span lifetime. - Provisioning records image preparation, bootstrap image resolution, overlay preparation, lifecycle configuration, pre-launch hooks, guest preparation, and launcher spawn as child spans. - VM startup reconciliation roots one trace for the persisted-sandbox scan, with per-sandbox restore and provision operations beneath it. The root remains open until all spawned restore tasks finish. - Delete cleanup records its own child operation. Design notes: - The gateway forwards its configured OTLP endpoint to managed external drivers. SDK `OTEL_*` variables continue to own sampling, batching, limits, headers, and transport tuning. - The VM driver has its own tracer provider and service resource so trace backends preserve the service boundary. - RPC operation names come from an explicit method mapping, keeping cardinality bounded without parsing the protobuf descriptor set at runtime. - Propagation uses a remote SpanContext for spawned provisioning. This keeps one trace while allowing the CreateSandbox server span to finish when the RPC response is sent. - Startup restoration is independent of gateway requests. It begins at the VM driver reconciliation span rather than attaching to an unrelated RPC. - Existing tracing events remain on the logging path. The OpenTelemetry layer exports spans only and excludes the SDK exporter callsites to avoid recursive traces. - Export configuration failures do not prevent the driver from serving, and buffered spans are drained during graceful shutdown. - Trace fields identify drivers, sandboxes, images, lifecycle phases, and gRPC outcomes without recording credentials, sandbox tokens, or request query parameters. Refs #2507 Signed-off-by: Kris Hicks <khicks@nvidia.com> |
||
|
|
905b554c7c |
refactor(network): consolidate proxy egress pipeline (#2373)
* refactor(network): introduce shared egress pipeline Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(network): cover shared proxy egress paths Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * refactor(network): make destination authorization explicit Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * refactor(network): pin proxy relay policy context Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(network): lock relay generation contracts Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(network): establish phase zero compatibility baseline Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * feat(policy): detect ambiguous network endpoints Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * feat(sandbox): fail closed on invalid policy updates Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * refactor(network): invalidate relays on policy changes Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * docs(policy): document validation failure posture Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(network): cover validation and middleware egress Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(config): move policy failure mode to gateway toml Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(network): name proxy contracts by behavior Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(network): align overlap validation with endpoint selection Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(network): expect hard loopback denial Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(network): match declared endpoint denial Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(policy): preserve path-specific endpoint overrides Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(network): respect hard-blocked host gateways Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(network): reconcile proxy refactor with main Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(network): preserve CONNECT policy generation Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(policy): cover runtime endpoint glob semantics Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(server): reject ambiguous policies before persistence Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * docs(policy): explain ambiguity preflight behavior Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(policy): compare body limits within protocol Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(proxy): avoid global tracing capture race Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * chore(server): format rebased provider tests Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(server): authenticate rebased policy requests Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(sandbox): retain runtime on middleware outage Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(server): preflight provider composition activation Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(sandbox): distinguish runtime failure transitions Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> --------- Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> Signed-off-by: John Myers <johntmyers@users.noreply.github.com> Co-authored-by: John Myers <johntmyers@users.noreply.github.com> |
||
|
|
d220d89468 |
feat(compute): negotiate gateway callback listeners (#2492)
* feat(compute): query gateway listener requirements Signed-off-by: Evan Lezar <elezar@nvidia.com> * feat(compute): add Podman listener requirements Signed-off-by: Evan Lezar <elezar@nvidia.com> * test(docker): use default gateway bind address Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(gateway): avoid wildcard primary listener Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(podman): validate callback listener discovery Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(server): support split dual-stack listeners Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(podman): support legacy rootless listener discovery Signed-off-by: Evan Lezar <elezar@nvidia.com> * test(e2e): accept loopback plaintext rejection Signed-off-by: Evan Lezar <elezar@nvidia.com> * docs(agent): add callback listener diagnostics Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(server): restrict compute callback listeners Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(podman): validate local callback port Signed-off-by: Evan Lezar <elezar@nvidia.com> * test(server): clarify callback listener contract Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(podman): require pasta for local callbacks Signed-off-by: Evan Lezar <elezar@nvidia.com> * docs(gateway): document RPM listener default Signed-off-by: Evan Lezar <elezar@nvidia.com> * refactor(server): keep listener provenance diagnostic-only Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(compute): preserve callback listener isolation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(e2e): remove Podman callback relay Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(packaging): preserve Podman callback loopback Signed-off-by: Drew Newberry <anewberry@nvidia.com> * ci(e2e): run VM smoke on nested-virt runner Signed-off-by: Drew Newberry <anewberry@nvidia.com> * ci(e2e): gate VM smoke on usable KVM Signed-off-by: Drew Newberry <anewberry@nvidia.com> * ci(e2e): probe KVM through VM driver Signed-off-by: Drew Newberry <anewberry@nvidia.com> * ci(e2e): tolerate hosted KVM denial Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(server): close traced futures before assertions Signed-off-by: Drew Newberry <anewberry@nvidia.com> * revert: remove tracing test stabilization Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Evan Lezar <elezar@nvidia.com> Signed-off-by: Drew Newberry <anewberry@nvidia.com> Co-authored-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
fa24299092 |
feat(gateway): export traces over OTLP (#2534)
Add an opt-in OTLP/gRPC trace exporter to the gateway. Export is enabled
by the presence of an `[openshell.gateway.otlp]` table with an endpoint;
there is no separate toggle.
Instrumented:
- Inbound request server spans, named for the RPC (`$service/$method`) or
`{method} {path}` for plain HTTP. They continue valid W3C `traceparent`
context when present and start a new trace otherwise. gRPC spans also
carry `rpc.system`, `rpc.service`, `rpc.method`, and trailer-derived
`rpc.grpc.status_code`.
- Compute driver calls (create, delete, list, get, validate, watch) as
client spans anchored on the `ComputeDriver` contract.
- Store reads and writes as children of the current request or loop span.
- Work with no inbound request: compute driver initialization, the sandbox
reconcile sweep, provider credential refresh tick, and driver watch events.
Each roots one operation trace so its child work does not arrive as anonymous
single-span traces.
This is deliberately not exhaustive. Auth, policy evaluation, and
middleware remain uninstrumented, as do store lifecycle calls (`ping`,
`close`) that a readiness poll would turn into a span per tick. The aim is
a useful trace tree at a reviewable size; coverage can grow against real
traces.
Design notes:
- The TOML table owns whether and where to export. The SDK `OTEL_*`
variables own how; sampling, batching, and limits are not mirrored into
gateway config.
- The OpenTelemetry layer exports spans only. Existing `tracing` events
remain on the stdout and sandbox-log paths and are not copied into trace
payloads.
- Telemetry never blocks the gateway. A malformed endpoint logs an error
and disables export rather than failing startup, and buffered spans are
drained during graceful shutdown.
- Failed spans carry error status without a separate `error.type` attribute.
Request spans use HTTP status and gRPC response trailers; driver spans use
the returned gRPC status; autonomous loop spans record failed results
explicitly. Store spans exempt `UniqueViolation` and `Conflict`, because
those errors report expected contention such as a held lease or an
optimistic-concurrency retry.
- The compute driver is reachable only through `TracedDriver::call`, so a
call cannot skip its span. This is the client half of a client/server pair
and the single place to inject context if drivers move out of process.
- Tests share one process-wide subscriber and in-memory exporter because
`tracing` caches callsite interest globally.
Inbound W3C trace context is propagated into gateway request spans. Context
is not yet injected into outbound driver calls, so a future out-of-process
driver would still need propagation at the `TracedDriver` seam.
The Helm chart is intentionally unchanged, so OTLP export cannot yet be
enabled on a chart-deployed gateway.
Refs #2507
Signed-off-by: Kris Hicks <khicks@nvidia.com>
|