Commit Graph
397 Commits
Author SHA1 Message Date
Gordon Sim 6588c7601a fix(auth): bound OIDC CA bundle reads
Signed-off-by: Gordon Sim <gsim@redhat.com>
2026-09-25 22:20:37 +01:00
Gordon Sim 2af0cb53d2 fix(helm): scope OIDC CA trust to the OIDC client
Signed-off-by: Gordon Sim <gsim@redhat.com>
2026-09-25 22:20:37 +01:00
Drew NewberryandPiotr Mlocek 73a181d32f docs: streamline README, add policy prover to architecture docs (#3718)
* docs(readme): streamline README and move reference detail to docs

Restructure the README as a short path from overview to quickstart to
further reading. Move prerelease install steps into the installation
guide and telemetry build flags into a new observability page. Fix
broken docs links and outdated runtime and credential descriptions.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(readme): describe 0.1.0 as adding new isolation primitives

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(architecture): add policy prover as a gateway component

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs: describe OpenShell as a runtime for fleets of autonomous AI agents

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(architecture): describe policy prover as formal verification

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(architecture): name OpenShell Sandbox in component table

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(architecture): fold isolation backend into supervisor row

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(architecture): mention formal verification in overview

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(run-agent): use the published OpenCode image in the first-agent guide

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(inference): correct provider examples and readiness

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(providers): correct Google binding and provider selection

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(readme): sharpen value prop, how it works, and explore further

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(run-agent): use example Anthropic profile and add policy advisor step

Import the example Anthropic profile, which now allows OpenCode, instead
of editing it with sed. Add a step that shows how to review and approve
mechanistic policy proposals as the agent needs more access.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(providers): add OpenRouter example for OpenCode

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(run-agent): run OpenCode against OpenRouter with a free model

Add an example OpenRouter provider profile scoped to OpenCode and switch
the first-agent guide to it, using a free Nemotron model so readers do
not need OpenRouter credits. Revert the OpenCode binary added to the
example Anthropic profile.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(readme): link first-agent guide and add agent skills section

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Co-authored-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-09-25 13:53:45 -07:00
Johnny Greco d7f921190b docs(policy): refresh policy documentation and references (#3563)
* docs(policy): correct schema and default policy guidance

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): add network recipes and update command reference

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): organize lifecycle guidance and troubleshooting

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): split policy overview into concepts and management tasks

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): reorganize network recipes as a cookbook

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): restructure schema reference by field group and protocol

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): align troubleshooting, advisor, and reference pages

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): fix first policy tutorial and security guidance

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): keep overview high level and move network rules to their own page

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): focus policy management on CLI workflows and remove command reference

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): clarify policy views and sandbox deletion in management guide

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): streamline network rule concepts and examples

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): correct request path wildcard semantics

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): rewrite policy advisor guide for clarity

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): clarify policy advisor scope, setup, and review

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): rewrite policy prover guide for clarity

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): explain the two uses of the policy prover

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): describe policy prover uses, boundaries, and coverage

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): place prover before advisor and troubleshooting last

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): remove unsupported CI guidance from prover page

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): tighten policy prover introduction

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): move policy change behavior into management guide

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): name prover check types and note expanding coverage

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): prefix prover and advisor sidebar labels

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): streamline policy schema reference

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): place default policy before schema reference

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): fold troubleshooting into policy management guide

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): correct tutorial log samples and GitHub push policy steps

The first policy tutorial said the 403 body begins with error, policy, and
rule, but the proxy serializes the body with sorted keys. Its log samples also
showed the wrong CONNECT deny reason for a sandbox without network rules, and
the L7 deny sample omitted the :443 authority, the `l7` engine, and the reason
tag that the shorthand formatter emits.

The GitHub tutorial filtered denials with `--level warn`, which hides the INFO
level OCSF policy events, and showed the retired key=value log format. Its
hand-written policy also omitted /bin from the restrictive default, so
`policy set` would reject the file for removing a filesystem path on a live
sandbox. Start from `policy get --base` and add only the network rules.

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): improve flow and terminology across policy pages

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): correct network rule matching and protocol details

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): align policy management steps with CLI behavior

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): correct policy advisor proposal and approval details

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): correct policy section, default, and schema details

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): correct prover installation and coverage limits

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): recommend tls skip for server-first protocols

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): fix stale baseline path and interpreter examples

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): move policy pages under how-it-works and fix links

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): align native TCP guidance in security best practices

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): restore policy.local and policy DNS details from main

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): state exact glob matching rules

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

---------

Signed-off-by: Johnny Greco <jogreco@nvidia.com>
2026-09-25 18:19:16 +00:00
Drew Newberry 9244868056 docs: refresh architecture and agent guides (#3705)
* docs: refresh architecture and agent guides

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs: describe updated security architecture neutrally

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs: highlight new isolation primitives

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(sandboxes): clarify how to disconnect

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs: align architecture and guides with current navigation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(extensibility): streamline extension authentication guidance

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-25 02:38:00 -07:00
Piotr Mlocek aead95b7ab fix(policy): propose rules for unknown DNS hosts (#3707)
* fix(policy): propose rules for unknown DNS hosts

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(policy): clarify synthetic DNS use across protocols

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(policy): harden unknown-host DNS observations

- Emit the policy_dns_ineligible denial for every unknown name and
  report observation staging failures as DNS failure events.
- Refuse unknown names during fail-closed quarantine and after the
  observation budget, now a quarter of each address family's pool.
- Pin transparent TCP to the mapping of the deciding policy generation
  so a reload between DNS and authorization fails closed.
- Stop Docker workloads from inheriting host DNS search domains, which
  let the first expanded short name claim an observation address.
- Share mechanistic draft polling in conformance, register
  new-hostname-proposal in the installed suite, and update docs.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* test(policy): build policy DNS proxy tests on every target

The proxy tests name PolicyEndpointId, which proxy.rs imported only on
Linux, so the macOS test build failed. Import it for test builds too.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(policy): name DNS queries and mapped hosts in OCSF denials

DNS denial and failure events attached port 53 to the queried name,
which read as a connection to that host. They now carry only the name.
Transparent TCP denials for a policy DNS address show the mapped
hostname and keep the synthetic address in dst_endpoint.ip.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

---------

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-09-25 08:28:41 +00:00
Drew Newberry c93fd94a5f feat(cli): import provider profiles from HTTP URLs (#3706)
* feat(cli): import provider profiles from HTTP URLs

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(cli): initialize crypto provider for remote profiles

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-25 06:12:18 +00:00
Drew Newberry 7a50c0899f fix(cli): stream piped exec stdin beyond gRPC request limit (#3687)
* fix(cli): stream piped exec stdin across gRPC messages

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(cli): preserve small exec requests and surface stdin errors

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(cli): preserve exec stdin limit across gRPC streaming

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-25 04:45:57 +00:00
Drew Newberry 4688061882 fix(sandbox): deliver complete exec output before success (#3688)
* fix(sandbox): preserve exec output through channel close

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(exec): propagate output delivery failures before exit

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(cli): clarify exec output delivery failures

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-25 04:03:49 +00:00
Piotr Mlocek c9257c8447 fix(policy): restore policy.local and proposal conformance (#3689)
* fix(policy): restore policy.local and proposal conformance

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* test(conformance): select policy scenarios by name

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* test(conformance): reduce policy scenario timing flakes

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* test(conformance): assert proposals target Bash

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

---------

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-09-24 21:03:26 -07:00
Piotr Mlocek 5448fb4f98 fix(auth): skip renewal for non-expiring sandbox JWTs (#3686)
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-09-25 03:40:09 +00:00
Drew Newberry 08548713c5 fix(install): honor pinned releases and speed up prerelease discovery (#3681)
* fix(install): use native packages for prereleases and speed up discovery

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(install): honor pinned versions and guard Snap migration

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(install): omit Snap migration guard

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-24 17:57:42 -07:00
Piotr Mlocek f213e9a49b docs(middleware): reorganize and expand middleware guides (#3636)
* docs(extensibility): reorganize extensibility and middleware guides

Add an extensibility overview that introduces the extension protocol and
links every extension point, and move extension caller authentication to
a shared page used by middleware and gateway interceptors.

Split the supervisor middleware guide into an overview, a Configure and
Operate guide, and a Middleware Operations reference for service authors.
The overview explains when to use middleware, shows where it runs, lists
current limitations, and defines the service contract. Configure and
Operate covers policy attachment, service registration, failure behavior,
and observability. Middleware Operations describes HTTP request, HTTP
response, and WebSocket message operations with shared inputs and results,
per-operation diagrams, and detail accordions.

Pin page slugs so links resolve, redirect the replaced dev middleware
URL, and update reference and architecture links.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(middleware): clarify navigation and service contracts

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(middleware): drop obsolete dev URL redirect

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

---------

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-09-24 23:43:31 +00:00
Drew Newberry 52cb8ecee7 fix(kubernetes): remove NetworkPolicy acknowledgement (#3677)
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-24 22:41:45 +00:00
Oliver Calder 6c864ec9ab fix(install): avoid installing incompatible docker snap (#3666)
* fix(install): avoid installing incompatible docker snap

The work to land RFC-0012 added new restrictions when interacting with
Docker by setting `NoNewPrivs`. This prevents the `docker` snap from
transitioning its AppArmor profile from `snap.docker.dockerd` to
`docker-default` when it tries to launch a container. Thus, the `docker`
snap is currently incompatible with OpenShell.

This commit prevents `install.sh` from installing the `docker` snap
before installing the `openshell` snap, and instead requires the user to
install a non-snap Docker daemon before proceeding with installing the
snap. Systems without the `snap` command are unaffected, since they
install native packages without checking for the presence of Docker or
other compute providers.

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* fix(snap): only snapd 2.76 for openshell snap since store installs work

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* fixup! fix(install): avoid installing incompatible docker snap

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* fixup! fix(snap): only snapd 2.76 for openshell snap since store installs work

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

---------

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>
2026-09-24 19:46:00 +00:00
Jason T. GreeneandMrunal Patel 7691f88e0e perf(server): enable WAL for the SQLite store; relax sync only for SSH session issuance (#3543)
* perf(server): enable WAL and NORMAL sync for the SQLite store

On-disk SQLite stores ran with sqlx defaults: rollback journal
(`journal_mode=delete`) and `synchronous=FULL`. Every autocommit write paid
several fsyncs and blocked readers while it held the lock, so gateway hot
paths made of many small writes serialized on disk latency. The clearest
case is `openshell forward service`, which mints and revokes an SSH session
token around every forwarded TCP connection: two commits per connection,
tens of milliseconds each on a virtual disk, wall clock linear in the
number of concurrent connections, and enough queueing that bursts hit the
per-sandbox connection cap and get refused.

Switch on-disk databases to WAL with `synchronous=NORMAL`. The mode change
runs once on a single connection before the pool opens: entering WAL needs
exclusive access to the file, so doing it up front means pool connections
only ever re-apply the pragma to a file already in WAL mode, and a failure
surfaces as one clear connect error. The first start after upgrading an
existing database therefore needs the file to be otherwise unopened.
`synchronous` is applied through the connect options on every pooled
connection. In-memory databases keep their defaults. A crash can now roll
back the most recent transactions without corrupting the database, which
is the standard WAL trade-off and fits the single-node scope of the SQLite
backend.

Tests cover a fresh store, an existing rollback-journal file that must be
switched on connect, sidecar permissions, and concurrent readers under a
burst of insert-then-update writes. Architecture, configuration and Helm
docs describe the durability trade-off, the sidecar files, and the local
filesystem requirement.

Signed-off-by: Jason T. Greene <jason.greene@redhat.com>

* fix(server): keep SQLite commits durable except SSH session issuance

WAL with synchronous=NORMAL can roll back acknowledged commits after a
power loss or kernel crash, including SSH session revocations and other
authorization-tightening writes. Run the main pool with synchronous=FULL
so every acknowledged write is durable; in WAL mode that is a single
fsync of the WAL per commit.

Add Store::create_relaxed for inserts that are safe to lose, and use it
only for SSH session issuance: a dropped token just fails validation.
On file-backed SQLite it runs on a dedicated single-connection pool with
synchronous=NORMAL. Both pools share one WAL, so the next FULL commit
also makes earlier relaxed commits durable. Postgres treats it as an
ordinary durable MustCreate insert.

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

---------

Signed-off-by: Jason T. Greene <jason.greene@redhat.com>
Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
Co-authored-by: Mrunal Patel <mrunalp@gmail.com>
2026-09-24 18:42:43 +00:00
grs 8369bc11a5 fix(podman): restore host gateway alias mediation (#3606)
* fix(podman): restore host gateway alias mediation

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(podman-e2e-tests): enable broader test podman e2e coverage

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(tests): make test more reliable

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(podman): fix macos linting error

Signed-off-by: Gordon Sim <gsim@redhat.com>

---------

Signed-off-by: Gordon Sim <gsim@redhat.com>
2026-09-24 18:41:59 +00:00
Evan Lezar 6af5520f37 fix(vm): confine OCI layer application to the rootfs (#3550)
* refactor(vm): isolate OCI layer application

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(vm): confine OCI layer application

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(vm): require trusted bootstrap image

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(vm): validate bootstrap image in gateway

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(vm): configure bootstrap image in tracing test

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-09-24 17:22:46 +00:00
Drew Newberry 0518bd4c83 fix(sandbox): preserve local sessions across host sleep (#3573)
* fix(cli): recover sandbox connect transport

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(auth): propagate non-expiring local sandbox sessions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(cli): distinguish main exit from transport loss

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(cli): bound sandbox connect recovery

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-24 17:22:30 +00:00
Oliver Calder e60098d748 fix(snap): install openshell snap via install.sh when snap available (#3656)
* fix(snap): update stale snap docs and tests

Previously, the `openshell` snap required the `docker` snap. Now, it
works with any Docker daemon running on the system. Furthermore, the
`snap-declaration` assertion on the `openshell` snap when installed from
the Snap Store causes the `openshell` snap to always connect to the
system `:docker` slot, rather than a slot provided by the `docker` snap.

This commit updates the documentation, including the `description` field in
`snapcraft.yaml`, to ensure that all information is correct and
up-to-date.

Additionally, some tests connected the `openshell:docker` plug to the
`docker` snap's `docker:docker-daemon` slot, which is inconsistent with
how the `openshell` snap operates when installed from the store. Update
those tests to connect to the system `:docker` slot as well.

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* fix(snap): require snapd 2.76 for openshell snap

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* fix(nix): align snap gateway reproducer timeout with release-canary

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* fix(snap): require snapd 2.77 for openshell snap

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* fixup! fix(snap): require snapd 2.77 for openshell snap

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* feat(snap): install openshell snap via install.sh when snap available

Change `install.sh` to install the `openshell` snap by default when
snapd is installed on the host. This installs the snap from the
`latest/stable` channel, which should match the most up-to-date release
tag on github.

If `OPENSHELL_VERSION=dev` is set for `install.sh`, then it will install
the `openshell` snap from the `latest/edge` channel, which matches the
latest dev release available on github.

The `openshell` snap currently requires Docker in order to function. If
a Docker daemon is already installed on the system, it will be used by
the `openshell` snap. Otherwise, `install.sh` will install the `docker`
snap first, wait for the Docker daemon to be ready, and then install the
`openshell` snap.

Also, update the `release-canary.yml` to split the `ubuntu-snap` job
into `ubuntu-snap-system-docker` and `ubuntu-snap-provisions-docker`,
which test the two aforementioned scenarios. Previously, `ubuntu-snap`
manually installed a given snap artifact as built from CI, but with
these new jobs, it instead uses `install.sh` to install the published
`openshell` snap from the `latest/edge` track, thus matching the
behavior of the other release canary jobs.

Make corresponding changes to the `nix` guest reproducer, documentation,
and tests.

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* fix(snap): configure local gateway authentication

The `openshell` snap runs the gateway as a systemd system service, which
runs as root. Thus, the mTLS certs are generated by root and stored in a
root-owned directory to which non-root users do not have access. For
this reason, the snap's `openshell-gateway-wrapper` script sets
`OPENSHELL_DISABLE_TLS=true`.

This commit ensures that the `openshell` snap's gateway allows
unauthenticated local access by writing a default `gateway.toml`
configuration file during the install hook, which runs after the snap is
first installed but before services are started. The config file
contains sets `allow_unauthenticated_users = true`.

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* fix(install): configure snap gateway authentication

Recently, a new install hook was added which writes a default config
file for the `openshell` snap to allow unauthenticated local access to
the gateway. This is because the gateway service runs as root and the
mTLS certificates are not accessible to non-root users.

However, the `install.sh` script installs the `openshell` snap from the
snap store, and the published version may not yet have that new install
hook. Or, the user may already have the snap installed, in which case
the install hook does not run. In either case, we need `install.sh` to
ensure that the config file is written to set the gateway auth to
`allow_unauthenticated_users = true`, and then restart the openshell
gateway service.

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* fixup! feat(snap): install openshell snap via install.sh when snap available

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

---------

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>
2026-09-24 15:10:08 +00:00
Piotr Mlocek a8f98ec09d fix(auth): remove legacy sandbox JWT admission (#3562)
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-09-23 23:57:21 +00:00
Mark CampbellandKris Hicks 679b190677 feat(testing): support independent gateway and supervisor image overrides (#3341)
* feat(testing): normalize configurable test images

Signed-off-by: Bobbins228 <mcampbel@redhat.com>
Signed-off-by: Kris Hicks <khicks@nvidia.com>

* feat(helm): add global image overrides

Signed-off-by: Bobbins228 <mcampbel@redhat.com>

* feat(helm): support image registry overrides

Signed-off-by: Bobbins228 <mcampbel@redhat.com>
Signed-off-by: Kris Hicks <khicks@nvidia.com>

* refactor(helm): simplify image configuration

Signed-off-by: Bobbins228 <mcampbel@redhat.com>

* fix(e2e): avoid reloading reused kind sandbox image

Signed-off-by: Bobbins228 <mcampbel@redhat.com>

* fix(helm): default sandbox image to nvcr.io/nvidia/base/ubuntu:24.04

Signed-off-by: Kris Hicks <khicks@nvidia.com>

---------

Signed-off-by: Bobbins228 <mcampbel@redhat.com>
Signed-off-by: Kris Hicks <khicks@nvidia.com>
Co-authored-by: Kris Hicks <khicks@nvidia.com>
2026-09-23 23:35:30 +00:00
krishicks 0b351c4a9b fix(helm)!: reduce gateway Secret privileges (#3616)
* fix(driver-kubernetes-secrets)!: store provider credentials in one namespace

The Kubernetes Secrets credential driver now stores every credential in its
configured namespace in all workspace modes and rejects handles that reference
any other namespace before contacting the Kubernetes API. The gateway reaches
credential Secrets through the Role in that namespace; this allows removing the
Secret rules from the ClusterRole.

- Remove the workspace_mode, gateway_id, and allow_reference_namespace driver
  settings and stop rendering them from Helm. Configurations that set them fail
  at startup. Existing credential state is not migrated.
- Add server.credentialDrivers.kubernetesSecrets.createNamespace to provision a
  dedicated credential namespace. The namespace is kept on uninstall, adopted
  by a reinstall of the same release, and left untouched when something else owns
  it.
- Update the gateway config reference, Kubernetes setup docs, 0.1.0
  upgrade guide, compute-runtime architecture, and cluster debugging skill.

Signed-off-by: Kris Hicks <khicks@nvidia.com>

* fix(helm): reduce gateway Secret permissions

Remove the gateway's Secret list permission in every workspace mode and
grant source Secret reads through a Role in the sandbox namespace.
Bootstrap Secret cleanup deletes Secrets by exact name instead of listing
them.

- Grant get on the copied client TLS and image-pull Secrets through a Role in
  the sandbox namespace. The ClusterRole keeps get and patch on those names
  for the ownership check and server-side apply into workspace namespaces.
- Delete sandbox and supervisor bootstrap Secrets by exact name, derived from
  the runtime generation recorded on the Sandbox and, on restart, the target
  generation, tolerating 404. The generation annotation is cleared only after
  cleanup succeeds, and each bootstrap Secret has a Pod owner reference, so
  garbage collection removes any generation the driver does not name.
- Drop Secret list from the ClusterRole and the shared-mode sandbox Role.
- Extend the managed e2e RBAC checks to Secret list.
- Update the Kubernetes setup and sandbox runtime docs, compute-runtime
  architecture, and cluster debugging skill.

Signed-off-by: Kris Hicks <khicks@nvidia.com>

* fix(driver-kubernetes): stage workspace Secrets per runtime generation

Every Secret the Kubernetes driver writes into a workspace namespace is
now scoped to one sandbox runtime generation, immutable, and created
with create only, so the gateway never reads, patches, or adopts an
existing Secret there. This removes the gateway's cluster-wide get and
patch on the copied client TLS and image-pull Secret names.

- In managed mode, create an immutable copy of each configured
  image-pull Secret per generation, named os-pull-<id>-<generation>-<n>
  and owned by the generation's workload and supervisor Pods. Pods and
  the restarted Sandbox template reference those names, and generation
  cleanup deletes them by name. A Secret already holding a generation
  name fails the create.
- Outside shared mode, stage the gateway client TLS material into the
  supervisor bootstrap Secret instead of copying the client TLS Secret
  into the workspace namespace.
- Remove the fixed-name TLS and image-pull copies, the target ownership
  read, and the ClusterRole get and patch rule on the copied names.
  Source reads stay in the sandbox-namespace Role.
- Update the managed e2e to expect generation image-pull Secrets and
  client TLS material in the supervisor bootstrap Secret, and to check
  that the gateway cannot read the copied names in workspace namespaces.
- Update the gateway config and compute driver references, Kubernetes
  setup and sandbox runtime docs, compute-runtime architecture, driver
  README, and cluster debugging skill.

Signed-off-by: Kris Hicks <khicks@nvidia.com>

* fix(helm)!: grant operator-mode Secret permissions through the workspace chart

The operator-mode gateway ClusterRole grants no Secret permissions.
The openshell-workspace chart Role, installed in each operator-managed
namespace, grants the gateway create and delete on Secrets for sandbox
runtime generations. Operator-managed namespaces require the workspace
chart.

- Fail the chart tests on any ClusterRole rule that includes Secrets in
  operator and shared modes.
- Install the workspace chart when the operator e2e provisions a
  namespace, and assert that the gateway has no Secret permissions in a
  namespace without it.
- Update the Kubernetes setup docs, 0.1.0 upgrade guide, compute-runtime
  architecture, and cluster debugging skill.

Signed-off-by: Kris Hicks <khicks@nvidia.com>

---------

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-09-23 21:35:22 +00:00
Jim Meyer 62df64625b fix(identity): assess leaf and ancestor executable identities (#3633)
* fix(security): pin supplied executable identity chains

Signed-off-by: Jim Meyer <jimeyer@nvidia.com>

* docs(security): document executable identity chain pinning

Signed-off-by: Jim Meyer <jimeyer@nvidia.com>

* fix(binary-identity): compile Linux ancestry hashing

Signed-off-by: Jim Meyer <jimeyer@nvidia.com>

* test(binary-identity): avoid cross-label ancestry fixture

Signed-off-by: Jim Meyer <jimeyer@nvidia.com>

* test(e2e): choose distinct denied TCP port

Fix flaky test due to sequential port assignment on MacOS

Signed-off-by: Jim Meyer <jimeyer@nvidia.com>

* fix(identity): bound executable evidence cache

Reject identity chains atomically when the supervisor cache reaches its hard limit, and preserve existing pins without eviction. Classify malformed or conflicting evidence separately from policy denials at the staged TCP boundary.

Signed-off-by: Jim Meyer <jimeyer@nvidia.com>

---------

Signed-off-by: Jim Meyer <jimeyer@nvidia.com>
2026-09-23 21:07:24 +00:00
Drew Newberry a649aa42fc docs: add 0.1.0 upgrade guide outline (#3540)
* docs: add 0.1.0 upgrade guide outline

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs: add upgrade change provenance

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs: update 0.1.0 upgrade guidance

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs: move upgrade navigation below security

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs: remove release notes page

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs: announce OpenShell 0.1.0

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs: reorganize guides and require user-owned images

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs: reorganize navigation around core concepts

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs: simplify navigation and tutorial catalog

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs: refine 0.1.0 release highlights

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs: align navigation and 0.1.0 guidance

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-23 20:26:25 +00:00
Drew Newberry 52aac37866 fix(network): honor HTTP response connection closure (#3581)
* fix(network): honor HTTP response connection closure

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(network): fix EOF fixture socket setup

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(network): keep pipeline probe response reusable

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-23 12:01:46 -07:00
Mrunal Patel a408f5dd08 chore(kubernetes): update Agent Sandbox to v1.0.3 (#3578)
Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
2026-09-23 18:30:24 +00:00
Piotr Mlocek bed9e5eafc fix(supervisor): use better error message when sandbox connect is not available (#3572)
* fix(supervisor): explain unavailable main terminal attachments

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(supervisor): reset cursor after terminal attachment errors

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(supervisor): format read-only warnings for terminal clients

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs: drop terminal attachment documentation additions

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(supervisor): exit read-only viewers on Ctrl-C

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs: simplify read-only viewer Ctrl-C guidance

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

---------

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-09-23 17:31:13 +00:00
Drew Newberry 069ae6bd96 fix(vm): scope GPU filesystem enrichment to assigned workloads (#3580)
* fix(vm): scope GPU filesystem enrichment to assigned workloads

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(supervisor): clarify proxy baseline enrichment

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-23 07:14:19 +00:00
Drew Newberry b671171091 fix(sandbox): qualify task memory against workload child (#3574)
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-23 03:36:56 +00:00
Drew Newberry 84960e70a3 fix(kubernetes): scope resource admission RBAC (#3571)
* fix(kubernetes): scope resource admission RBAC

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(helm): gate PVC admission reads

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-23 00:02:10 +00:00
Evan Lezar 49df4d7ca7 fix(server): drain supervisor ownership cleanup on shutdown (#3547)
Closes #3546

Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-09-22 22:29:54 +00:00
Mrunal Patel bdffa102c3 feat(api): add durable exec launch admission (#3324)
* feat(api): add durable exec launch admission

Fence duplicate exec launches with keyed durable admission and producer-owned terminal completion. Keep uncertain launches unresolved and never replay output or interactive input.

Part of #3051 (phase 4a).

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(api): fence exec identity across authorization lookups

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

---------

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
2026-09-22 21:52:19 +00:00
Drew Newberry 1e34e8c576 fix(drivers): require admission labels for external resources (#3538)
* fix(drivers): require admission labels for external resources

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(drivers): address resource admission review findings

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(core): reserve driver-owned admission labels

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(core): clarify workspace admission label

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(drivers): clarify resource admission failures

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(e2e): configure resource admission fixtures

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): retry forbidden admission lookups

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(e2e): preserve external driver admission defaults

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-22 21:22:30 +00:00
Akram Ben Aissi 293fab75d4 fix(sandbox): support kernels < 5.19 via seccomp WAIT_KILLABLE_RECV fallback (#3420)
* fix(sandbox): fall back to plain seccomp listener when WAIT_KILLABLE_RECV is unavailable

SECCOMP_FILTER_FLAG_WAIT_KILLABLE_RECV was added in Linux 5.19. On older
kernels (for example RHEL 9.x / 5.14 nodes such as RHCOS on OpenShift) the
flag is rejected with EINVAL, which made the capability-free sandbox fail to
start during the notification probe with "notification launcher disappeared".

Install the notification listener with WAIT_KILLABLE_RECV when the kernel
supports it and fall back to a plain NEW_LISTENER on EINVAL. The fallback
listener records wait_killable_recv = false: its notification receive is
uninterruptible, but the sandbox is otherwise fully functional.

Signed-off-by: Akram <akram.benaissi@gmail.com>

* fix(sandbox): address review — legacy read-only listener mode + cancellation invariant

Follow-up to the PR review (GATOR-de00bfcc-01 / mrunalp): make the < 5.19
fallback cancellation-safe instead of racing broker writes.

- Record an explicit ListenerMode (Killable vs LegacyReadOnly); add
  writes_disabled()/mode() and emit the selected mode in qualification output
  (seccomp_listener_mode).
- Centralize task-memory output writes behind NotificationListener::
  write_task_output; in LegacyReadOnly mode getpeername, accept/accept4 with a
  non-null address, and sendmmsg length write-backs fail closed with EOPNOTSUPP.
  accept with a null address, socket/connect/bind/listen/sendto/sendmsg keep
  working (copied inputs, scalar responses, atomic ADDFD_SEND).
- Enforce the launch invariant `cancellation || task_memory_writes_disabled` in
  SandboxConfirmEvidence::validate() rather than dropping cancellation
  unconditionally; add task_memory_writes_disabled to SeccompEvidence.
- Add per-path fail-closed tests (write_task_output, write_socket_addr, a real
  plain listener installed on a modern kernel) and confirmation-invariant tests.
- Correct the flag-semantics comments and document both modes plus the reduced
  legacy syscall compatibility in architecture/sandbox.md.

Signed-off-by: Akram <akram.benaissi@gmail.com>

* docs: document legacy read-only sandbox mode on kernels before 5.19

Kernels without SECCOMP_FILTER_FLAG_WAIT_KILLABLE_RECV (< 5.19, e.g. RHEL 9.x /
RHCOS 5.14) run the sandbox in a legacy read-only cancellation mode where the
broker fails closed with EOPNOTSUPP on the mediated operations that write results
back into workload memory (getpeername, accept/accept4 with a non-null address,
sendmmsg length write-backs). Document this observable behavior and its syscall
limitations in the public Fern docs: the support-matrix kernel requirements and
the OpenShift runtime guidance.

Signed-off-by: Akram <akram.benaissi@gmail.com>

* style(sandbox): satisfy rustfmt and clippy doc_markdown

Match the pinned rustfmt (Rust 1.95.0) line-wrapping for the write_task_output
call, and backtick `legacy_read_only` in the qualification-report doc comment
so clippy::doc_markdown (-D warnings) passes.

Signed-off-by: Akram <akram.benaissi@gmail.com>

* fix(sandbox): migrate task_memory_writes_disabled into backend protocol

Add the `task_memory_writes_disabled` field to `SeccompEvidence` in the
backend protocol contract and relax the validation from requiring
`cancellation` to accepting `cancellation || task_memory_writes_disabled`.

This completes the rebase migration missed by the isolation-interface
refactor: the sandbox reports this field but the backend struct lacked
it, and the validator rejected every pre-5.19 legacy listener.

Signed-off-by: Akram <akram.benaissi@gmail.com>

---------

Signed-off-by: Akram <akram.benaissi@gmail.com>
2026-09-22 20:37:19 +00:00
Artem LytvynandJohn Myers 470a34635d fix(api): make WatchSandbox loss-aware and resumable (#3209)
* fix(api): emit warning on WatchSandbox broadcast lag instead of terminating

Broadcast lag on the status, log, and platform receivers was converted to a RESOURCE_EXHAUSTED status that terminated the whole watch stream. Lag is recoverable: the receiver resumes at the oldest surviving message. Emit a SandboxStreamWarning and continue streaming instead; keep terminating on Closed. Add helpers and unit tests covering the warning payload and receiver recovery after lag.

Partially addresses #3055 (cursor/resume follow up separately).

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* refactor(server): group per-sandbox log bus state and stamp sequence numbers

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* feat(proto): add resume cursor fields to sandbox watch API

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* feat(server): stamp watch cursors from a shared per-sandbox sequence

Allocate cursors from a single SeqAllocator shared by the log and
platform event buses, so a sandbox's merged watch stream carries
unique, strictly increasing cursors. A single resume_after_cursor can
then unambiguously locate a client's position across both sources.

Rewrite both publish paths to allocate the sequence, stamp
event.cursor, send, and append to the tail under one lock. This
removes the previous get_mut().expect() TOCTOU race where a concurrent
remove() between the two lock sections could panic.

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* feat(server): serve WatchSandbox resume from cursor with gap detection

Add tail_after() to the log and platform event buses, returning every
buffered event newer than a client's resume cursor. Each PerSandbox now
tracks last_trimmed_seq (the highest seq it has evicted) so a resume is
reported as an unrecoverable ResumeGap only when this bus dropped an
event the client still needs.

Judging gaps by evictions, not by the tail's oldest seq, is required
under the shared cursor space: each bus's tail is non-contiguous in the
global sequence because the other bus owns the missing seqs, so
comparing against tail.front() would flag false gaps.

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* feat(server): resume WatchSandbox from cursor across log and platform buses

Wire resume_after_cursor into the watch producer. On a non-zero cursor,
replay events strictly after it from both the log and platform buses,
merge by shared cursor, and emit in order before entering the live loop.
A trimmed range on either bus is an unrecoverable gap and terminates the
stream with OUT_OF_RANGE carrying the requested and earliest-available
cursors, distinct from recoverable lag which warns and continues.

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* test(server): cover WatchSandbox cursor resume paths

Add handler-level tests for the resumable watch stream: replay strictly
after the client cursor, merge log and platform events in shared-cursor
order, suppress duplicates when resuming at the latest cursor, and
terminate with OUT_OF_RANGE when the requested cursor has been trimmed.

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* docs(api): document WatchSandbox loss-awareness and resume

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* fix(server): deliver watch events once and harden cursor teardown

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* feat(sdk): add loss-aware resumable watch_logs to Rust SDK client

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* fix(server): keep watch cursors monotonic across teardown and restart

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* fix(server): merge live watch sources by cursor before emission

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* fix(api): bind watch cursors to a cursor space and merge tail sources

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* fix(server): revalidate the watch cursor space after collecting replay

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* test(server): update the public RPC schema fingerprint for the string cursor

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* fix(server): hold watch events above the publication watermark and emit the watch lag warning before its batch

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* test(server): synchronize the watch live-order test with the end of initialization

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* fix(sdk): use canonical sandbox name in watch_logs

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* test(sdk): guard canonical-name addressing in watch_logs

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* fix(server): fix public rpc schema

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* fix(api): reconcile watch resume rebase

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(server): bound interactive relay cleanup

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-09-22 19:45:55 +00:00
Yuedong Wu 718dba3430 fix(policy)!: reject removed tls endpoint values (#3414)
Signed-off-by: Yuedong Wu <dwcn22@outlook.com>
2026-09-22 17:58:19 +00:00
Philippe Martin 35e0a68e4a feat(kubernetes): support corporate proxy CA bundle (#3447)
* feat(kubernetes): support corporate proxy CA bundle

The Kubernetes driver had no way to supply a CA bundle for the corporate
egress proxy, so an `https://` proxy with a private CA, or a TLS-intercepting
proxy, could not be used. Podman and VM already expose `proxy_ca_bundle`.

Add `proxy_ca_bundle` to `[openshell.drivers.kubernetes]` as a path the
gateway Pod reads. The gateway stages the PEM into the existing
per-generation supervisor bootstrap Secret and passes
`--upstream-proxy-ca-bundle` on the supervisor argv. That Secret is already
immutable, owner-referenced and garbage-collected, and its volume mounts
every key at /.openshell/supervisor with no items filter, so this needs no
new object kind, volume, mount, or RBAC verb, and works in shared, managed
and operator workspace modes.

The bundle is deliberately read from the gateway's filesystem rather than
referenced as an object in the sandbox namespace. It becomes a trust anchor
for every upstream the sandbox reaches, so it must stay in the gateway's
trust domain; the immutable staging Secret also keeps the anchor from
changing underneath a running sandbox.

Bound the staged bundle at 256 KiB. The shared reader's limit is exactly the
apiserver's own Secret limit and the bootstrap Secret carries four other
keys, so a bundle between the two would pass gateway startup and then fail
every sandbox create with an opaque `data: Too long`.

Delegate the URL, no_proxy, connect_by_hostname and ca_bundle rules to the
shared validate_upstream_proxy_settings, keeping the Secret-specific
credential block local: this driver accepts an explicit
`proxy_auth_allow_insecure = false` without credentials, which the shared
rules reject. This also fixes the acknowledgement being demanded for an
`https://` proxy, where the credential travels inside the verified TLS
session. Add auth_setting_label so the inline-credential diagnostic names
the Secret keys instead of proxy_auth_file, which this driver rejects as an
unknown key.

Document that the bundle should carry only the CA that signs the proxy's
certificate, or that an intercepting proxy re-signs upstream certificates
with. Public roots already reach the sandbox through the supervisor image and
its TLS stack, and the bundle is concatenated with that system store into a
single boundary control frame, so a full merged trust bundle spends the frame
budget on duplicated roots. The frame, not the apiserver Secret limit, is the
tighter of the two ceilings in practice; raising the staging bound requires
checking it.

Closes #3443

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(helm): quote proxy CA ConfigMap references

Signed-off-by: Philippe Martin <phmartin@redhat.com>

---------

Signed-off-by: Philippe Martin <phmartin@redhat.com>
2026-09-22 17:57:40 +00:00
Varsha 7139df8ca5 fix(exec): preserve output after stdin EOF and verify stream completion (#3359)
* fix(exec): preserve output after stdin EOF and verify stream completion

Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>

* test(exec): cover fair duplex progress and live SDK completion

Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>

* test(sdk): compare large exec buffers with native equality

Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>

* test(exec): use workspace-scoped sandbox name in EOF regression

Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>

---------

Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>
2026-09-22 14:11:06 +00:00
50230616d5 refactor(runtime): retire Community image dependencies (#3386)
* feat(sandbox): default to official Alpine sandbox image

default_sandbox_image() now returns docker.io/library/alpine:3.22, a generic
version-qualified official image, so a fresh install no longer depends on the
community sandbox image catalog. All compute drivers (docker, podman,
kubernetes, vm) inherit this fallback.

Part of #3116.

Signed-off-by: Akram
Signed-off-by: Akram <akram.benaissi@gmail.com>

* feat(deploy): default deployment configs to the official Alpine sandbox image

Update the shared gateway default_image, Helm chart values, the standalone
Kubernetes manifest, and the dev gateway task scripts to use
docker.io/library/alpine:3.22 instead of the community base image, consistent
with default_sandbox_image(). GPU e2e image-build base is left unchanged (CUDA
needs a glibc base).

Part of #3116.

Signed-off-by: Akram
Signed-off-by: Akram <akram.benaissi@gmail.com>

* feat(driver): default to numeric non-root identity for USER-less images

With the default sandbox image now Alpine, images that declare no OCI USER
must start instead of being rejected. When the image declares no USER and
the policy requests none, the Podman and Docker drivers now supply a numeric
non-root identity (DEFAULT_SANDBOX_UID/GID = 1000) instead of rejecting,
matching the numeric-identity behavior of the Kubernetes and VM drivers. The
supervisor's resolved-identity path runs the sandbox as a synthesized
non-root account without the account existing in the image. Images that
declare a USER keep the OCI resolution path unchanged.

Part of #3116.

Signed-off-by: Akram <akram.benaissi@gmail.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(conformance): use Alpine workload image

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* refactor(policy): drop community image /app path from default policy

The restrictive default policy granted read-only access to /app, a directory
that only existed in the community base image. A generic Alpine default has no
/app, so remove it. Landlock best-effort already ignores absent paths; this
just stops advertising a community-specific layout in the default.

Part of #3116.

Signed-off-by: Akram
Signed-off-by: Akram <akram.benaissi@gmail.com>

* docs(config): document Alpine default images

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(podman): report early sandbox termination

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(podman): initialize rootless workspace ownership

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(sandbox): qualify NVIDIA Ubuntu default

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): initialize rootful default workspace

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(sftp): add native sandbox adapter

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sftp): gate runtime helper support to Linux

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sftp): support standard OpenSSH file operations

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sftp): harden rename and special file handling

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(runtime): remove community image dependencies

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(e2e): build provider readiness tool fixture

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(e2e): use a dedicated Noble fixture for Docker tests

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Akram
Signed-off-by: Akram <akram.benaissi@gmail.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Evan Lezar <elezar@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
2026-09-22 14:43:51 +02:00
John T. MyersandDrew Newberry c8a4ff5f19 fix(kubernetes): bind bootstrap to runtime identity (#3531)
* fix(kubernetes): bind bootstrap to runtime identity

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(compute): compensate runtime binding failures

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(compute): clean up backend on store failure

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(compute): merge runtime binding after start

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(auth): bind restarted sandbox sessions

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* refactor(compute): fold runtime binding into authentication

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
2026-09-22 05:08:37 +00:00
0a770d9173 feat(kubernetes): support HA gateway rebalancing (#1868)
* feat(kubernetes): support HA gateway rebalancing

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* perf(server): cache peer connections, tokens, and owner lookups

Every forwarded relay rebuilt its setup from scratch: an owner lookup, a
blocking read of the peer token, a TLS connect to the owning replica, and
a TokenReview plus Pod GET on the receiving side. Sandbox service routing
does this per HTTP request, so the apiserver calls scaled with traffic.

Cache all of it on ServerState:

- peer channels pooled per endpoint, so relays multiplex over one
  connection instead of redialing
- peer tokens keyed by SHA-256, expiring at min(ttl, token exp) so a hit
  cannot accept an expired token
- owner records for 3s against a 45s ownership TTL, still freshness
  checked before use

Entries are evicted when a relay fails. Also raise HTTP/2
max_concurrent_streams to 1024, since pooling funnels every relay between
two replicas onto one connection and hyper's default of 200 sits below
the 256 pending-relay budget.

Signed-off-by: divesh <dgude@nvidia.com>

* perf(server): pool upstream connections for sandbox services

Each HTTP request to a sandbox service opened its own supervisor relay,
paying a new TCP connection and HTTP/1 handshake every time. Worse, it
counted against the 32 in-flight relay cap, so a service handling more
than 32 concurrent requests failed outright.

Pool idle upstreams per endpoint and port, up to 8 each for 15s. Reuse is
safe because the pool only returns a connection hyper reports as ready,
and HTTP/1 cannot start a request until the previous body has drained.
Upgrades are never pooled since they take the connection over, and a
failed send evicts that endpoint. Pruning is bounded per key, with the
full sweep limited to once per 30s.

Signed-off-by: divesh <dgude@nvidia.com>

* fix(server): address HA gateway review findings (#3449)

- Let a gateway own supervisor sessions without a peer endpoint. Requiring
  one whenever the store is PostgreSQL broke every single-instance
  PostgreSQL deployment, because no sandbox supervisor could connect.
  A cross-replica request to an owner that advertises no endpoint now fails
  immediately naming the cause, instead of retrying until the wait timeout.
- Close a supervisor session on heartbeat only when another replica owns it,
  or after renewals fail for the ownership TTL. A database error no longer
  drops every session heartbeating during an outage.
- Clamp owner record ages at zero so a skewed or corrupt stored timestamp
  cannot produce a negative age.
- Bound the cross-object advisory lock with a lock timeout, so a stuck holder
  fails instead of blocking every mutation in the fleet.
- Refuse to start when a peer endpoint is configured on a multi-replica
  backend but peer authentication is unavailable, and warn when a
  multi-replica backend has no peer endpoint at all.
- Reject a plaintext peer endpoint when the gateway serves TLS.
- Skip the sandbox watch poller on single-replica backends, where the local
  update bus already sees every write.
- Rate-limit the peer owner cache sweep so an insert no longer scans the
  whole map under the lock.
- Retry GET and HEAD on a pooled upstream the sandbox closed, instead of
  returning 502, and drop an emptied endpoint from the pool right away.
- Document the gateway peer environment variables and the post-rollout
  ownership skew operators should expect.

Signed-off-by: divesh <dgude@nvidia.com>

* fix(server): harden HA supervisor ownership

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Signed-off-by: divesh <dgude@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: divesh <dgude@nvidia.com>
Co-authored-by: Divesh Chowdary <47188680+FrostGod@users.noreply.github.com>
2026-09-21 21:22:12 +00:00
Seth Jennings 2493d415c2 feat(extensions)!: normalize protocol negotiation (#3352)
* feat(extensions)!: normalize protocol negotiation

Closes #3057

Introduce a shared extension handshake, enforce protocol and capability compatibility across extension families, and expose immutable negotiated snapshots through gateway info and the Go SDK.

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(credentials): fail fast on negotiation errors

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(extensions): validate gateway handshake metadata

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(go-sdk): re-export extension kind constants

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(extensions): fail fast on credential handshake rejection

Signed-off-by: Seth Jennings <sjenning@redhat.com>

---------

Signed-off-by: Seth Jennings <sjenning@redhat.com>
2026-09-21 13:19:39 -07:00
Shiju fa8f6d3949 feat(cli): promote profile commands to top level (#3258)
* feat(cli): promote profile commands to top level

Add profile discovery and management commands with shared handlers for
the existing provider entry points. List a flat catalog across scopes
and follow continuation tokens through full and short pages.

Describe metadata, credentials, endpoints, TLS inspection, and MCP access
settings while preserving complete JSON/YAML definitions. Cover parser
equivalence, scope forwarding, pagination, and inspection settings with
focused unit and compiled-CLI integration tests.

Update docs, public skills, examples, and E2E command invocations.

Refs #2588

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(cli): remove redundant workspace selector qualification

Use the imported WorkspaceSelector in the provider integration helper so
the target passes Clippy with warnings denied.

Signed-off-by: Shiju <shiju@nvidia.com>

---------

Signed-off-by: Shiju <shiju@nvidia.com>
2026-09-21 17:18:55 +00:00
krishicks cbf026366d fix(ocsf): require network activity endpoints (#3355)
Previously, Network Activity could be constructed without a source or
destination endpoint, allowing connection, accept, relay, and configuration
events to violate the OCSF 1.8 endpoint constraint.

Now, NetworkActivityBuilder requires a source or destination endpoint at
compile time. Connection failures identify the workload peer or genuine
transparent destination, listener failures identify the listening endpoint,
and mediation-lane failures use Application Lifecycle rather than fabricated
network endpoints. Malformed forward requests use HTTP Activity with a
method-only request, generated 400 response, and workload peer.

Additionally, Unix relay-channel events use Base Event, policy-validation
warnings use Config State Change, and the unused bypass monitor is removed
because the current isolation architecture no longer uses it.

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-09-21 17:13:41 +00:00
Drew Newberry dee4f98dda chore(vm): refresh runtime defaults and hardening (#3446)
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-21 10:19:30 +00:00
Drew Newberry 1905069948 feat(sandbox): expose services during creation (#3439)
* feat(sandbox): expose services during creation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(providers): refresh Codex credentials in gateway

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(cli): normalize create-time service URLs

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(server): roll back failed service exposure

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(example): simplify Codex provider setup

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(example): separate provider setup commands

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): harden create-time service exposure

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(example): bundle Codex provider profile

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(example): allow npm-installed Codex binary

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-20 23:19:25 -07:00
Johnny GrecoandKirit93 5bce19ab47 feat(prover): check process, Landlock, and destination IP containment (#3394)
* feat(prover): check process, landlock, and destination IP containment

Signed-off-by: Kirit93 <kthadaka@nvidia.com>
Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(prover): reject ambiguous implicit IP modes

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(prover): validate implicit IP modes policy-wide

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(prover): tighten containment edge handling

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* test(prover): cover boundary v2 evidence shapes

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(prover): document boundary v2 evidence limits

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* refactor(prover): simplify result version contract

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* refactor(prover): clarify coverage terminology

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

---------

Signed-off-by: Kirit93 <kthadaka@nvidia.com>
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
Co-authored-by: Kirit93 <kthadaka@nvidia.com>
2026-09-18 21:38:08 +00:00
Drew Newberry 17ce738bfb fix(ci)!: remove gateway callback listener dependency (#3365)
* fix(ci): repair post-merge release canary

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(packaging): bootstrap canary runtime prerequisites

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci(canary): collect macOS VM diagnostics

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci(canary): pin libkrun-compatible macOS runner

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci(canary): limit macOS smoke test to package startup

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(compute)!: remove gateway callback listeners

Run Docker supervisors on host networking so they use the operator-configured primary gateway endpoint. Remove the unused compute-driver callback listener negotiation and listener-scoped routing machinery.

BREAKING CHANGE: The ComputeDriver API no longer exposes GetGatewayListenerRequirements or GatewayListenerRequirement. External drivers must regenerate bindings and connect supervisors to the configured primary gateway endpoint.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(docker): use sandbox runtime image in launcher

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(podman): exercise production endpoint selection

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(e2e): route supervisors to reachable gateways

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(e2e): align Podman endpoint fixtures

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): preserve host aliases for supervisors

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): align sandbox host gateway pin

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(e2e): address Docker fixtures by bridge IP

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(e2e): serialize sandbox lifecycle cases

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(e2e): host Docker TCP fixture with gateway

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(e2e): use loopback for host-network supervisor

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-18 20:55:55 +00:00
Drew Newberry d91b1999a0 feat(api)!: use sandbox names as canonical RPC references (#3272)
* feat(api)!: use sandbox names as canonical references

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(cli): update forward color fixture for workspace scope

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(supervisor): use sandbox names for settings lookup

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(e2e): use canonical sandbox request fields

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(e2e): use canonical sandbox receipt field

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(server): harden sandbox mutation handling

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(ci): update rebased sandbox references

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(api)!: standardize canonical entity references

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(api): codify protobuf API conventions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(api): preserve workspace selector semantics

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(api): restore workspace selector parity

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(api): preserve descriptive name fields

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(api): update e2e request fixtures

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(core): omit workspace selector during bootstrap

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* chore(api): remove proto convention checker

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* chore(api): refresh schema fingerprints after rebase

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(cli): use canonical provider receipt field

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-18 11:39:55 -07:00