Commit Graph
1199 Commits
Author SHA1 Message Date
Simon Scatton 455883905a fix(python): remove CLI from wheel (#2321)
The Maturin-based wheel packaging was a historical remnant from when the local gateway launch path and OpenShell CLI were coupled in one binary. The gateway and CLI now ship as standalone artifacts, so the Python distribution should contain only the SDK.

Build a single platform-independent setuptools wheel, verify that it cannot contain native code or an openshell entry point, and simplify the release jobs and documentation for SDK-only PyPI installs.

Signed-off-by: Simon Scatton <sscatton@nvidia.com>
v0.0.113
2026-08-25 13:20:42 +00:00
Evan Lezar 4fe5b0f605 fix(ci): allow pasta to receive Podman stop signals (#2900)
* fix(ci): allow pasta to receive Podman stop signals

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(ci): fail on pasta SIGTERM denial

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(ci): retain AppArmor denial diagnostics

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-08-25 12:43:45 +00:00
Russell Bryant aa848f164d fix(core): fall back to podman CLI when no API socket is found (#1858)
* fix(core): fall back to podman CLI when no API socket responds

Auto-detection only checked well-known Podman socket paths, so a Podman
machine exposing its API socket at a non-standard location went
undetected. The symlink at a well-known path is not always present; it
varies by Podman version, machine provider, and platform.

Extend detect_podman_socket() to fall back to podman CLI discovery when
no well-known candidate responds: podman info --format json determines
whether the service is local or remote, and podman machine inspect
resolves the host-side forwarded socket for VM-backed machines. All
existing callers (driver auto-detection, the Podman driver, and the VM
driver's container-engine fallback) pick this up without change.

Select the machine backing the active Podman connection instead of the
first entry from podman machine inspect, honoring podman's connection
precedence: CONTAINER_CONNECTION, then CONTAINER_HOST (mapped to a
connection by URI), then the containers.conf default. An explicit
endpoint that maps to no known machine is left unresolved rather than
guessing an unrelated machine. When CONTAINER_HOST is an explicit unix://
socket, use that path directly since podman info connects through it.

Update the gateway config reference and the Podman driver README, which
described probe-only detection.

Signed-off-by: Russell Bryant <rbryant@redhat.com>

* fix(core): inspect podman machine by name and bound discovery probes

Address two startup defects in Podman socket auto-detection:

- Remote discovery ran `podman machine inspect` with no arguments, which
  inspects only `podman-machine-default`. On a host whose default
  connection is a different machine, selection fell back to the first
  (wrong) entry, so the gateway could operate against the wrong backend.
  Resolve the active machine first and inspect it by name, trying the
  `-root`-stripped machine for rootful connections and returning None
  when no machine can be mapped instead of substituting another.

- `podman info`, `podman machine inspect`, and `podman system connection
  list` used unbounded `Command::output()`, so a stalled machine, SSH
  connection, or helper could hang gateway startup indefinitely. Route
  all three through a bounded runner that kills and reaps the child on a
  documented deadline and returns None so detection can continue.

Signed-off-by: Russell Bryant <rbryant@redhat.com>

* fix(core): bound podman probe stdout drainage and kill descendants

The previous timeout only bounded waiting for the direct child to exit.
Draining stdout still called read_to_end, which waits for pipe EOF — a
Podman probe can leave a daemonized descendant (SSH multiplexer, gvproxy)
that inherited the stdout pipe, so drainage could wait indefinitely even
after the direct child exited, defeating the deadline.

Run each probe as its own process-group leader, cap stdout drainage by
the same deadline via a channel, and on expiry kill the whole process
group so descendants holding the pipe are terminated and EOF is reached.
Add a regression test where the direct child exits but a backgrounded
descendant keeps stdout open.

Signed-off-by: Russell Bryant <rbryant@redhat.com>

* fix(core): make podman probe deadline absolute against escaped descendants

The prior drainage timeout still ended in an untimed rx.recv() after
killing the process group. A descendant that escaped the probe's process
group (e.g. via setsid) while holding the stdout pipe would not be killed,
so read_to_end never saw EOF, the reader never sent, and that recv() could
block gateway startup indefinitely.

On drainage timeout, best-effort kill the process group and return None
immediately, abandoning the reader thread instead of waiting on it again.
The deadline now bounds the whole call regardless of what descendants do.
Add a regression test whose stdout-holding descendant escapes the process
group (`set -m`) and assert the call still returns promptly.

Signed-off-by: Russell Bryant <rbryant@redhat.com>

---------

Signed-off-by: Russell Bryant <rbryant@redhat.com>
2026-08-25 11:49:22 +00:00
Russell Bryant 72b9c4acef test(e2e): pin direct podman calls to harness socket on macOS (#2909)
The podman_gateway_start e2e test shells out to `podman` directly (unlike
every other Podman e2e test, which routes through the gateway). The harness
`e2e/with-podman-gateway.sh` clobbers `XDG_CONFIG_HOME` with an empty dir to
isolate CLI/SDK gateway metadata. On macOS the podman client resolves its VM
connection config from `XDG_CONFIG_HOME`, so a bare `podman ps` falls back to a
nonexistent native rootless socket and fails with "unable to connect to Podman
socket". On Linux the socket resolves via `XDG_RUNTIME_DIR`, so the test passes
there and this is a macOS-only false failure.

Target the same API socket the gateway uses by passing `--url unix://$SOCKET`
when the harness-exported `OPENSHELL_PODMAN_SOCKET` is set, mirroring the
shell's `podman_cmd` helper. When the var is unset (running the test outside the
harness), fall back to plain `podman`, so Linux behavior is unchanged.

Signed-off-by: Russell Bryant <rbryant@redhat.com>
2026-08-24 22:04:25 +00:00
Philippe Martin 0a1f246587 feat(sandbox,podman): trust corporate CA for https:// proxies and intercepted TLS (#2512)
* feat(sandbox,podman): trust corporate CA for https:// proxies and intercepted TLS

The corporate proxy chaining only accepted plain http:// proxy URLs, so
operators whose forward proxy terminates TLS with a private corporate CA
had no way to reach it, and TLS-intercepting proxies (mitmproxy, squid
ssl-bump) that re-sign tunneled server certificates broke every upstream
handshake after CONNECT.

The supervisor now accepts https:// proxy URLs: it wraps the connection to
the proxy in TLS before the CONNECT handshake, verifying the proxy
certificate against the built-in Mozilla roots, the system CA bundle, and an
optional operator corporate CA bundle. The upstream dial returns a
Plain/Tls stream enum consumed generically by the relay paths.

The corporate CA is delivered as a driver-supplied command-line argument
(--upstream-proxy-ca-bundle), never an environment variable, matching the
hardened proxy-config model where a sandbox image cannot influence the
operator's egress boundary. It is folded into the sandbox combined trust
bundle (write_ca_files) and the L7 upstream verification store
(build_upstream_client_config) at startup, so intercepted upstream
handshakes succeed and sandbox workloads trust the re-signed certificates.
Configuration is fail-closed: a CA bundle set without a proxy, or an
unreadable or certificate-free file, is fatal rather than silently
weakening the trust boundary.

The shared parse_upstream_proxy_url validator accepts https:// (recording
the scheme so the driver and supervisor agree), keeping the explicit-port
requirement. The Podman driver gains a proxy_ca_bundle operator setting
(TOML, --sandbox-proxy-ca-bundle, OPENSHELL_SANDBOX_PROXY_CA_BUNDLE) that
bind-mounts the host PEM read-only into the sandbox (a CA certificate is not
secret) and points --upstream-proxy-ca-bundle at it, with a create-time
readability check.

The standalone dev gateway task passes OPENSHELL_SANDBOX_PROXY_CA_BUNDLE
through to the generated podman config, so a local gateway can be pointed at
a TLS-intercepting proxy without hand-editing the regenerated TOML.

Refs #1792

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(sandbox): reject CA bundles with valid PEM framing but invalid X.509 DER

The proxy CA bundle validation counted PEM blocks that base64-decoded
successfully, but did not verify the decoded bytes were accepted as
trust anchors by RootCertStore. A bundle with syntactically valid PEM
framing but invalid DER would pass the startup check while contributing
zero usable anchors, causing opaque TLS failures at runtime instead of
a fail-closed startup error.

Validate decoded certificates through RootCertStore::add_parsable_certificates
and reject the bundle unless at least one is accepted.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(sandbox,podman): allow proxy auth without insecure acknowledgement for https:// proxies

For an https:// proxy the Proxy-Authorization credential travels inside
the verified TLS session, so the proxy_auth_allow_insecure
acknowledgement is unnecessary. Previously both http:// and https://
proxies required it, producing a misleading cleartext-risk diagnostic
for a path that is already encrypted.

Skip the requirement when the proxy URL uses https://; the
acknowledgement is still tolerated if set. Updated in both the
supervisor and Podman driver validation paths, with docs and tests.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix: format

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(kubernetes): use truly unsupported scheme in proxy validation test

https:// is now a supported proxy scheme after a13c4dce, so the
unsupported-scheme test must use a genuinely unsupported scheme.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(e2e): sign the https proxy fixture listener cert with a CA

The corporate-proxy E2E fixture served a single `openssl req -x509`
certificate as its TLS listener identity. OpenSSL marks that certificate
`basicConstraints: critical, CA:TRUE`, and rustls refuses a CA
certificate presented as an end-entity certificate (CaUsedAsEndEntity).
The supervisor's TLS handshake with the proxy therefore failed, the
upstream dial errored, and the workload's CONNECT was dropped without a
response, so podman_corporate_proxy_trusts_ca_bundle_for_https_proxy
failed on the approved destination while policy denial still worked.

Generate a corporate CA and a separate listener leaf signed by it, serve
the leaf chain, and publish only the CA as the bundle the supervisor
trusts. This is what an intercepting proxy actually presents, and it
exercises the corporate-CA trust path rather than pinning the listener
certificate itself.

Refs #1792

Signed-off-by: Philippe Martin <phmartin@redhat.com>

---------

Signed-off-by: Philippe Martin <phmartin@redhat.com>
2026-08-24 21:10:25 +00:00
Artem Lytvyn d7e137f12e fix(proxy): normalize trailing-dot CONNECT hosts before policy evaluation (#2248)
* fix(proxy): normalize trailing-dot CONNECT hosts before policy evaluation

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* fix(supervisor-network): keep trailing dot on upstream proxy CONNECT host and add regression test

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* fix(supervisor-network): drop redundant UpstreamProxyArgs import in tests

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

---------

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>
2026-08-24 20:25:48 +00:00
krishicks e457974a52 feat(docker): export driver traces over OTLP (#2851)
Mirror the VM and Podman driver tracing setup for Docker. Export Docker
driver spans through OTLP/gRPC as the distinct openshell-driver-docker
service, preserve gateway trace context, record lifecycle and asynchronous
provisioning operations, and report gRPC failures.

Docker currently runs in-process when selected as a built-in gateway driver.
Add the same temporary server-boundary shim used by Podman so traces retain
the shape they will have when Docker moves to a separate process. Generalize
the gateway provider selection for both in-process drivers and share the OTLP
collector fixture across Docker, Podman, and VM tracing tests.

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-08-24 18:39:21 +00:00
Simon Scatton 905e99aa2a feat(build): add Nix-native Linux toolchains (#2875)
* feat(nix): add glibc 2.28 development shell

* feat(nix): add musl development shell

* feat(build): use mold in musl development shell

* feat(flake): add nix remote cache

* feat(build): use mold in default development shell
v0.0.112
2026-08-24 12:16:00 +00:00
krishicks 7fc6138981 docs(agent): warn on missing workflow labels (#2815)
Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-08-24 08:10:10 +00:00
grs 40d1b48666 feat(provider): support for SPIFFE backed token exchange (#1970)
* feat(provider): add ability to request token exchange instead of client credentials as OAuth grant_type

Signed-off-by: Gordon Sim <gsim@redhat.com>

* test(proxy): add further tests for token exchange

Signed-off-by: Gordon Sim <gsim@redhat.com>

* test(provider): add runnable example for token exchange

Signed-off-by: Gordon Sim <gsim@redhat.com>

* test(e2e): cover Podman token exchange grants

Signed-off-by: Gordon Sim <gsim@redhat.com>

* refactor(oauth): extract duplicated functionality from server and supervisor

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(provider): evict nearest-to-expiry entry from intermediate token cache

Signed-off-by: Gordon Sim <gsim@redhat.com>

* doc(supervisor): add podman example for token exchange

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(provider): withhold token-exchange subject credentials

Signed-off-by: Gordon Sim <gsim@redhat.com>

---------

Signed-off-by: Gordon Sim <gsim@redhat.com>
2026-08-24 05:42:30 +00:00
Prekshi Vyas e3dc011208 fix(provider): isolate unbound static credentials (#2862)
Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>
2026-08-23 20:36:19 +00:00
Drew Newberry 2f7fb65591 chore(docs): relicense Fern stylesheet under Apache 2.0 (#2882)
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-08-21 19:50:00 +00:00
krishicks 6c38646c59 feat(dev): add dedicated gateway:podman task (#2880)
Previously, Podman could be selected through automatic driver detection or with
`mise run gateway -- --driver podman`, but it did not have a dedicated task
like the Docker and VM drivers.

This adds a gateway:podman task and moves the Podman-specific setup into its
own script. The generic gateway task now delegates Podman launches to that
script.

Additionally:

Unlike Docker, which rebuilds and bind-mounts the supervisor binary, Podman
uses a dev-tagged supervisor image that can become stale. The default Podman
supervisor image is therefore rebuilt on each launch.

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-08-21 19:11:19 +00:00
John T. MyersandJohn Myers 679fe4c334 fix(policy): validate the applicable advisor candidate (#2850)
* fix(policy): bind reviews to applicable candidates

Build and validate the exact effective-policy candidate before approval, bind review to live policy/provider/credential inputs, and preserve inspected endpoint contracts during mechanistic expansion.

Closes #2821

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* fix(policy): canonicalize advisor review inputs

Serialize nested protobuf maps in stable key order for proposal review tokens and effective-policy hashes. Narrow reused multi-port endpoint contracts to the denied port so advisor proposals cannot widen binary access. Add regressions for both cases.

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* test(e2e): keep advisor sandbox running

Create the issue 2821 regression sandbox detached with a durable canonical main process so policy denial, approval, and hot-reload checks run before lifecycle exit.

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* fix(policy): apply reviewed draft batches atomically

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <johntmyers@users.noreply.github.com>
2026-08-21 19:05:53 +00:00
Artem Lytvyn 7adc05af7a feat(supervisor): expose sandbox name to middleware request context (#2771)
* feat(supervisor): expose sandbox name to middleware request context

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* feat(supervisor): add workspace to middleware request context

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

---------

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>
2026-08-21 18:50:13 +00:00
Drew Newberry 56c45a997e fix(providers): honor configured profile sources in sandboxes (#2878)
Closes #2877

Use the gateway's configured provider profile catalog for sandbox creation and provider attachment validation.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-08-21 17:29:04 +00:00
Evan Lezar de4c1fecf5 fix(test-guest): support RPM installs with DNF5 (#2864)
* fix(test-guest): support RPM installs with DNF5

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(test-guest): clarify RPM install arguments

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-08-21 14:40:47 +00:00
Evan Lezar 3be2cd8a29 fix(helm): preflight Agent Sandbox APIs (#2867)
* fix(helm): preflight Agent Sandbox APIs

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(kubernetes): share Agent Sandbox setup

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(e2e): wait for Agent Sandbox CRD status

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* ci(canary): sparse-checkout sandbox helper

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-08-21 14:32:19 +00:00
Drew Newberry 20d2e867e0 fix(sandbox): reject stale exit during restart (#2857)
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
v0.0.111
2026-08-21 03:56:35 +00:00
Drew Newberry 82f62fa3ca docs(rfc): define stable release policy (#2695)
* docs(rfc): define stable release policy

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(rfc): link review pull request

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(rfc): summarize release proposal

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(rfc): replace nightlies with release candidates

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(rfc): add breaking change examples

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(rfc): simplify compatibility proposal

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(rfc): add SELinux Podman coverage

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(rfc): simplify capability release rules

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(rfc): streamline release stability proposal

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(rfc): simplify release qualification criteria

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(rfc): clarify alpha exit motivation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(rfc): refine release qualification policy

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(rfc): define API maturity and conformance opt-outs

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(rfc): adopt Preview and Stable API maturity

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(rfc): define Stable and Experimental APIs

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(rfc): allow feature-driven minor releases

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(rfc): clarify release build audiences

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(rfc): define pre-release train semantics

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-08-21 03:55:45 +00:00
Drew NewberryandEvan Lezar 40f822906c feat(compute): add standalone first-party drivers (#2822)
* feat(compute): add standalone first-party drivers

Build Docker, Podman, Kubernetes, and VM drivers as external binaries and
exercise each through the public compute-driver API. Keep the external E2E
setup complete at introduction, including VM image selection, Kubernetes
post-renderer isolation, supervisor reuse, and scoped Podman coverage.

External Kubernetes endpoints support shared and managed workspace modes.
Operator mode remains restricted to the in-process driver because gateway
authentication and the driver must share a dynamic namespace allowlist.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci(e2e): run managed and external drivers independently

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(compute): cover external driver socket contract

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* ci(e2e): install bundled Z3 build dependency

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(server): gate in-tree driver tracing

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>
Co-authored-by: Evan Lezar <elezar@nvidia.com>
2026-08-21 02:49:55 +00:00
Drew Newberry dfb06e6a39 test(e2e): align detached sandbox assertions (#2856)
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-08-20 17:45:01 -07:00
John T. Myers 0300e6dcb2 fix(sandbox): order sidecar provider updates by generation (#2849)
Closes #2847

Separate sidecar delivery ordering from opaque provider environment fingerprints and cover live bind/unbind convergence.

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-08-20 23:53:48 +00:00
Drew Newberry 6c34a3c645 fix(sandbox): stabilize canonical main process tests (#2854)
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-08-20 16:33:15 -07:00
Drew Newberry ef296806f5 feat(sandbox): add canonical main process (#2726)
* feat(sandbox): add canonical main process

Closes #2710

Persist and supervise one canonical workload per sandbox, attach sandbox connect to its retained session, and make every unexpected main-process exit terminal.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(sandbox): simplify canonical main process contract

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): preserve legacy VM main compatibility

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): preserve main status across driver updates

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): satisfy macOS process lint

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): gate Linux exit acknowledgement publisher

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(sandbox): simplify main process plumbing

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(supervisor): make controlling tty ioctl portable

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(supervisor): initialize canonical process environment

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* perf(supervisor): optimize retained main session

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(sandbox): detach main session on ctrl-c

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): use explicit main detach keys

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(dev): atomically stage Docker supervisor

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(test): align Docker main environment assertion

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(sdk): expose canonical main process fields

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-08-20 22:18:23 +00:00
Drew Newberry 9ae3760768 refactor(compute): register compiled drivers (#2786)
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-08-20 21:42:31 +00:00
John T. Myers b2ea81822b feat(network): enable Docker and Podman policy DNS and transparent TCP (#2723)
* feat(network): enable Docker transparent TCP egress

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(e2e): cover Docker transparent TCP egress

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* feat(network): correlate transparent TCP audit events

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(examples): add transparent TCP Redis demo

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(examples): demonstrate blocked TCP connections

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(examples): focus Redis demo audit output

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): close transparent TCP policy bypasses

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(sandbox): reject unsupported TCP policy reloads

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(ci): satisfy Linux transparent TCP lints

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* feat(podman): enable transparent TCP egress

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(podman): permit policy DNS port binding

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(e2e): use qualified transparent TCP hostname

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(podman): preserve exact policy DNS names

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(podman): route policy DNS over TCP

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(dns): serve multiple TCP queries per connection

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(network): explain native DNS and TCP egress

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(sandbox): reconcile runtime reload with upstream

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): harden transparent DNS capture

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(podman): preserve resolver behavior for native tcp

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(network): clarify native tcp runtime constraints

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): remove unused transparent tcp pin

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): admit redirected transparent tcp

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): restore podman transparent networking

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(podman): permit alpine busybox binaries

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(podman): use portable alpine keepalive

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(podman): build musl networking fixture

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(podman): isolate musl DNS probe

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(podman): keep privileged port capability dropped

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): preserve transparent TCP port 53

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): report synthetic pool pressure by family

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(podman): bind tcp fixtures before readiness

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(podman): grant fixture low-port bind

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-08-20 19:26:53 +00:00
krishicks 2c0adf4868 feat(podman): export driver traces over OTLP (#2782)
Mirror the VM driver tracing setup for Podman. Export standalone driver
spans through OTLP/gRPC as the distinct openshell-driver-podman service,
propagate W3C context across ComputeDriver RPCs, record bounded RPC names
and failures, and flush buffered spans during graceful shutdown.

Podman still runs in-process when selected as a built-in gateway driver.
Add a temporary tracing shim that partitions gateway and Podman spans by
target into separate tracer providers while preserving their shared trace
and parentage. The shim also emits the same ComputeDriver server boundary
that the tonic layer emits out of process, keeping the observable trace
shape stable when Podman is eventually extracted.

Trace container create preparation, image and storage setup, lifecycle
operations, and cleanup. Document the service boundary and cover it with
isolated and repeated tracing tests.

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-08-20 18:43:16 +00:00
John T. Myers 0c6a3443ec feat(network): add policy DNS correlation store (#2713)
* feat(network): add policy DNS correlation foundation

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(network): describe dormant policy DNS boundary

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): harden policy DNS publication

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): audit policy DNS failures

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): preserve DNS answer order

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-08-20 18:17:57 +00:00
John T. Myers 4d7f402ce2 feat(policy): establish direct TCP egress foundation (#2711)
* feat(policy): accept explicit tcp endpoint protocol

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* refactor(network): snapshot authoritative egress decisions

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(policy): document explicit tcp protocol

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(policy): defer transparent TCP release guidance

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* chore(go): regenerate sandbox protobuf bindings

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): complete tcp egress foundation

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(policy): document explicit tcp contract

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): fail closed on authorization errors

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(podman): fence delayed exit events before restart

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(policy): validate network endpoint destinations

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(providers): opt in tcp credential fixture

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(policy): require dns host for transparent tcp

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-08-20 16:43:09 +00:00
Drew Newberry 4c5fce6e59 refactor(compute): support external driver parity (#2744)
* refactor(compute): negotiate external driver behavior

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(compute): revert external driver documentation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(compute): negotiate gateway-managed lifecycle

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(compute): remove driver feature negotiation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(compute): let drivers declare gateway lifecycle

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(compute): clarify lifecycle ownership

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-08-20 15:56:29 +00:00
Simon Scatton 9505ca5ed1 chore: remove Bazel build support (#2840)
Signed-off-by: Simon Scatton <sscatton@nvidia.com>
2026-08-20 15:08:09 +00:00
Drew Newberry 7909fb5d0f refactor(compute): unify gateway restart reconciliation (#2743)
* refactor(compute): unify gateway restart reconciliation

Remove the Docker-specific gateway shutdown cleanup and reconcile persisted running intent through ComputeDriver::StartSandbox for Docker, Podman, and VM drivers. Explicitly stopped sandboxes remain stopped.

Refs #2417

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(compute): stop local sandboxes on shutdown

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(e2e): match managed Podman containers

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(compute): synchronize lifecycle sweeps

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
v0.0.110
2026-08-20 10:12:41 +00:00
Piotr Mlocek 701382d015 fix(podman): wait for container stop completion (#2820)
* fix(podman): wait for container stop completion

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(podman): clarify stop restart race

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

---------

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-08-20 01:00:11 +00:00
Seth Jennings c90fd648f3 fix(cli): reuse sandbox provisioning display (#2816)
Signed-off-by: Seth Jennings <sjenning@redhat.com>
2026-08-19 22:02:17 +00:00
krishicks b7078dc2ea chore(gitignore): add Pi agent state (#2813)
And collect the other agent-related entries together.

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-08-19 21:42:44 +00:00
Mrunal Patel 6e90f3d5a0 feat(providers): store refresh credentials in credential drivers (#2801)
* feat(providers): store refresh credentials in credential drivers

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(providers): harden refresh credential lifecycle

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(providers): migrate legacy refresh secrets before skip

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* refactor(providers): defer credential migration

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(providers): make refresh configuration atomic

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

---------

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
2026-08-19 21:37:24 +00:00
Drew Newberry 998db04780 feat(policy): allow non-root sandbox identities (#2785)
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-08-19 21:20:14 +00:00
Jesse Jaggars 2eb0880a00 feat(cli): support OIDC device authorization grant for headless login (#2795)
* feat(cli): support OIDC device authorization grant for headless login

Closes #2793

Add OAuth 2.0 Device Authorization Grant (RFC 8628) support to the OpenShell CLI's OIDC login flow. When running in a headless environment (OPENSHELL_NO_BROWSER=1) without a client secret configured, the CLI now uses the device code flow instead of the browser-based PKCE flow.

The device code flow:
- Requests a device code and user code from the IdP's device authorization endpoint
- Displays a verification URL and user code to the user
- Polls the token endpoint until the user completes authorization or the code expires
- Supports slow_down responses per RFC 8628 by increasing the polling interval

This implementation:
- Extends OidcDiscovery to optionally capture device_authorization_endpoint
- Adds oidc_device_code_flow function with proper error handling for all RFC 8628 error codes
- Updates gateway add and gateway login to dispatch to device flow when browser is suppressed
- Adds comprehensive unit tests for device flow structs and response parsing
- Updates gateway authentication documentation to describe the device code fallback

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(cli): validate OIDC device token responses

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(cli): add PKCE to OIDC device flow

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* docs(cli): document PKCE device flow

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

---------

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>
2026-08-19 17:44:46 +00:00
grs 3a16012dbe fix(cli): prompt for fresh OIDC login after logout (#2773)
Signed-off-by: Gordon Sim <gsim@redhat.com>
2026-08-19 15:46:27 +00:00
alangou 0d708d6d51 fix(policy): gate uninspected credentialed endpoints (#2493)
* fix(policy): gate uninspected credentialed endpoints

Signed-off-by: Adrien Langou <alangou@nvidia.com>

* refactor(cli): extract allowed-ip option parsing

Signed-off-by: Adrien Langou <alangou@nvidia.com>

* fix(policy): gate endpointless credential bindings

Signed-off-by: Adrien Langou <alangou@nvidia.com>

---------

Signed-off-by: Adrien Langou <alangou@nvidia.com>
2026-08-19 15:18:02 +00:00
Mrunal Patel 8d67250a5d fix(providers): keep refresh credential handles stable (#2780)
* fix(providers): keep refresh credential handles stable

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(providers): protect refresh-owned credentials

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

---------

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
v0.0.109
2026-08-18 20:45:36 +00:00
Adel ZaaloukandClaude Opus 4.6 600bbae845 feat(ocsf): emit AI inference events via ai_operation profile on ApiActivity (#2664)
Apply the official OCSF ai_operation profile (introduced in v1.8.0) to
ApiActivity [6003] events when the inference proxy routes a model call
through inference.local. Attaches an ai_model object (name, ai_provider)
and puts token counts and latency in unmapped fields.

ApiActivity [6003] is the schema-correct class for the ai_operation
profile in v1.8.0 (HttpActivity only gets it in v1.9.0). In Splunk CIM,
ApiActivity maps to the "Change" data model, naturally separating
inference events from regular HTTP proxy traffic.

Changes:
- Add AiModel object and ai_model field on BaseEventData
- Add ApiActivityEvent struct and ApiActivityBuilder
- Add emit_ai_inference in proxy.rs using ApiActivity with ai_operation
- Vendor OCSF v1.8.0 schemas including api_activity class, ai_model
  object, and ai_operation profile definitions
- Bump OCSF_VERSION to 1.8.0
- Update schema validation to skip profile-gated required fields

Shorthand: API:INFERENCE [INFO] claude-3-haiku via anthropic 701ms [POST /v1/messages]

Splunk/SIEM backward compatibility (v1.1/v1.3 CIM mapping) is tracked
separately in #2662 as a configurable serialization concern.

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-08-18 16:30:03 +00:00
Roland Huss dc374e8878 chore(sdk/go): remove coverage.out from tracking (#2774)
Test coverage artifact was accidentally committed. Already in .gitignore.

Signed-off-by: Roland Huß <rhuss@redhat.com>
v0.0.108
2026-08-18 11:56:48 +00:00
Evan Lezar 2115b0c42b fix(driver-podman): compile container spec on macOS (#2789)
* ci(drivers): lint portable drivers on macOS

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(driver-podman): compile container spec on macOS

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-08-18 11:54:07 +00:00
alangou 877ddbacb4 fix(supervisor-network): canonicalize dot-segments before policy evaluation (#2699)
* fix(supervisor-network): strip path parameters before dot-segment resolution

Signed-off-by: Adrien Langou <alangou@nvidia.com>

* fix(supervisor-network): scope allow_encoded_slash to the matched L7 endpoint

Signed-off-by: Adrien Langou <alangou@nvidia.com>

* fix(supervisor-network): check the canonical target for encoded slashes

Signed-off-by: Adrien Langou <alangou@nvidia.com>

---------

Signed-off-by: Adrien Langou <alangou@nvidia.com>
2026-08-17 23:04:23 +00:00
Kirit Thadaka 6340d18745 Update docs.yml to remove warning banner (#2687)
* Update docs.yml

Remove warning banner from docs page.

* Update README.md
2026-08-17 22:45:13 +00:00
Shailendra Singh 4dfeff59c3 docs(rfc): add RFC 0013 native Windows support via MXC (#2071)
* docs(rfc): add RFC 0013 native Windows support via MXC

Propose native Windows 11 support through a build-only MSVC lane and a new
in-process, supervisor-free MXC compute driver, with host-side governed egress
and an OpenShell to MXC policy-translation seam.

Refs: #2050
Signed-off-by: Shailendra Singh <shailendras@nvidia.com>

* docs(rfc): address native Windows MXC review feedback

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>

* docs(rfc): update governed egress proxy topology

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>

* docs(rfc): clarify Windows proxy and gateway topology

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>

---------

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
2026-08-17 19:51:13 +00:00
Polite_realism 5d9b0f0476 fix(inference): prepend publisher prefix for Vertex non-Anthropic models (#2735)
* fix(inference): prepend publisher prefix for Vertex non-Anthropic models

Vertex AI's OpenAI-compatible endpoint requires the request body's model
field to carry a publisher prefix (e.g. google/gemini-2.5-flash), but
validate_vertex_model_id rejects slash as a path-traversal guard. This
created a deadlock: bare model IDs pass validation but are rejected by
Vertex with HTTP 400 "Malformed publisher model"; prefixed IDs are
rejected at configuration time.

Fix: in resolve_vertex_ai_route, compute body_model_id for non-Anthropic
routes by prepending the publisher from infer_vertex_publisher() or the
explicit VERTEX_AI_PUBLISHER config value. The bare model_id still goes
through the path-traversal validator unchanged. Anthropic rawPredict routes
encode the model in the URL path, not the body, and are unaffected. Both
the project/region path and the base-URL-override path apply the prefix.
For unrecognised models with no explicit publisher the bare ID is forwarded
unchanged; Vertex's 400 is the correct observable signal in that case.

Add an integration test in openshell-router that spins up a mock Vertex
endpoint accepting only the publisher-prefixed form and rejecting the bare
model name, verifying the body rewrite produces the required format.

Closes #2351

Signed-off-by: politerealism <burdcat17@gmail.com>

* fix(inference): propagate publisher-prefixed model_id through inference bundle

resolve_route_by_name_with_credentials built the ResolvedRoute with
config.model_id (the bare stored value) rather than resolved.route.model
(the publisher-prefixed value computed by resolve_vertex_ai_route). As a
result, the bundle delivered to sandboxes carried e.g. "gemini-2.5-flash"
instead of "google/gemini-2.5-flash", so live sandbox requests still hit
Vertex AI with the bare model name and received HTTP 400 "Malformed
publisher model".

Fix: use resolved.route.model in the bundle construction so the
publisher prefix survives the bundle boundary and the router sends the
correct body to Vertex AI.

Update the existing gemini bundle test to assert the prefixed model_id
and add a dedicated regression test that verifies the bundle carries the
publisher prefix for non-Anthropic Vertex routes.

Signed-off-by: politerealism <burdcat17@gmail.com>

* style(inference): apply rustfmt to new regression test

Signed-off-by: politerealism <burdcat17@gmail.com>

* fix(inference): address clippy lints in build_vertex_route

- Invert if !is_anthropic to satisfy clippy::if_not_else
- Replace match on Option with map_or_else to satisfy clippy::option_if_let_else

Signed-off-by: politerealism <burdcat17@gmail.com>

---------

Signed-off-by: politerealism <burdcat17@gmail.com>
2026-08-17 19:40:46 +00:00
alangou 6ebf10e2ec fix(build): preserve version prefixes in mise lockfile (#2778)
Signed-off-by: Adrien Langou <alangou@nvidia.com>
2026-08-17 19:26:02 +00:00