Commit Graph
58 Commits
Author SHA1 Message Date
Johnny Greco 87929ad130 docs: keep page URLs aligned with file paths (#3713)
* docs: align page file names and nav labels with published URLs

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs: redirect moved dev pages and fix agent guide redirects

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* ci(docs): check that page URLs match their file paths

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(docs): align navigation checks with repo conventions

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs: omit historical redirect aliases

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(docs): preserve published overview and TypeScript URLs

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(docs): redirect unversioned overview and TypeScript URLs

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

---------

Signed-off-by: Johnny Greco <jogreco@nvidia.com>
2026-09-28 02:08:54 +00:00
Drew Newberry a3ed8c79cf docs: refresh architecture and extensibility pages (#3741)
* docs(readme): show Rust SDK installation command

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(architecture): lead with policy enforcement

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(architecture): reuse README how it works text

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(architecture): refresh architecture page and diagrams

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(extensibility): simplify extension points diagram and overview

Redraw the extension points diagram as two left-to-right rows for the
data plane and control plane, describe which layers each extension point
extends, and move authentication after Building Extensions as
Authenticating Extensions with updated cross-page links.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-27 02:18:29 +00:00
Drew Newberry 4ce767fc0c docs: describe 0.1 series in release callouts (#3732)
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-26 03:06:48 +00:00
Drew Newberry a67567e583 fix(snap): require mTLS for the snap gateway (#3726)
* fix(snap): require mTLS for the snap gateway

Replace the installer opt-in with an authenticated snap gateway. The wrapper
no longer forces plaintext, so the gateway serves TLS from the bundle it
already generates in $SNAP_COMMON/tls. The install hook writes a config that
enables mTLS user auth instead of unauthenticated access, and a new
post-refresh hook migrates the exact legacy default on existing installs.

install.sh waits for the gateway, detects whether it serves TLS, copies the
client bundle into the target user's snap state directory, and registers the
gateway over HTTPS. Older plaintext snap revisions still register over HTTP
with a warning. The release canary asserts mTLS auth and HTTPS registration.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(snap): pass config preflight and detect the mTLS gateway reliably

An explicit [openshell.gateway.mtls_auth] table fails config preflight,
which validates mTLS auth before the local TLS bundle supplies the client
CA. Write a default that pins the Docker driver instead; with the wrapper's
TLS bundle the gateway requires client certificates and enables mTLS user
auth automatically, as the native packages do.

The mTLS gateway rejects TLS handshakes without a client certificate, and
it still answers plaintext loopback HTTP for sandbox service routing, so the
installer could misdetect it as a legacy plaintext gateway. Probe HTTPS with
the root-owned client bundle, and treat a gateway as legacy only when a
plaintext gRPC Health call succeeds.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* chore(snap): simplify install hook comment

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(snap): migrate insecure gateway configs on refresh

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(snap): simplify mTLS detection and config migration

Detect the mTLS snap from the installed revision's post-refresh hook instead
of probing plaintext gRPC, and drop the scheme global. Remove the installer's
pre-hook config fallback, which is dead now that every channel ships the
install hook and which wrote the insecure default. Give the install hook a
single write path with a simple backup name, and shorten the manual client
certificate steps.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(snap): stop keeping a copy of replaced insecure configs

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(install): make the snap an opt-in install method

Stop selecting the OpenShell snap just because the snap command exists. Linux
installs default to the Debian or RPM package; OPENSHELL_INSTALL_METHOD=snap
(or deb, rpm) selects the package explicitly. Hosts that already have the
OpenShell snap keep refreshing it rather than gaining a second gateway on the
same port. The release canary and snap repro script opt in explicitly.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(snap): let the gateway auto-detect its compute driver

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(snap): restart the gateway after refresh

Published revisions use refresh-mode: endure, and snapd honors the old
revision's setting during a refresh, so the plaintext gateway kept running
with the migrated config unused until a manual restart. Restart the gateway
from the post-refresh hook so the mTLS config takes effect immediately.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(snap): drop refresh notes from the snap description

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(snap): trim snap refresh notes from installation docs

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-26 01:50:02 +00:00
Johnny Greco f155899527 docs(tutorials): run Pi with OpenRouter (#3722)
* docs(tutorials): run Pi with OpenRouter

Switch the Pi tutorial from Anthropic to OpenRouter, add fd to the Pi image
so Pi does not try to download it from GitHub inside the sandbox, and rename
the page to Run Pi with OpenRouter with a dev redirect from the old URL.

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(tutorials): drop redirect for dev-only Pi page

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(tutorials): make the Pi tutorial easier to follow

Add prerequisites and steps, explain providers and the profile fields in
plain terms, inspect the sandbox while Pi is still running, and add clean-up
and troubleshooting sections.

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(tutorials): clarify Pi's providers and sandbox additions

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(tutorials): make the Pi tutorial more conversational

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

---------

Signed-off-by: Johnny Greco <jogreco@nvidia.com>
2026-09-26 01:38:59 +00:00
Drew NewberryandPiotr Mlocek 73a181d32f docs: streamline README, add policy prover to architecture docs (#3718)
* docs(readme): streamline README and move reference detail to docs

Restructure the README as a short path from overview to quickstart to
further reading. Move prerelease install steps into the installation
guide and telemetry build flags into a new observability page. Fix
broken docs links and outdated runtime and credential descriptions.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(readme): describe 0.1.0 as adding new isolation primitives

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(architecture): add policy prover as a gateway component

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs: describe OpenShell as a runtime for fleets of autonomous AI agents

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(architecture): describe policy prover as formal verification

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(architecture): name OpenShell Sandbox in component table

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(architecture): fold isolation backend into supervisor row

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(architecture): mention formal verification in overview

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(run-agent): use the published OpenCode image in the first-agent guide

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(inference): correct provider examples and readiness

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(providers): correct Google binding and provider selection

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(readme): sharpen value prop, how it works, and explore further

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(run-agent): use example Anthropic profile and add policy advisor step

Import the example Anthropic profile, which now allows OpenCode, instead
of editing it with sed. Add a step that shows how to review and approve
mechanistic policy proposals as the agent needs more access.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(providers): add OpenRouter example for OpenCode

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(run-agent): run OpenCode against OpenRouter with a free model

Add an example OpenRouter provider profile scoped to OpenCode and switch
the first-agent guide to it, using a free Nemotron model so readers do
not need OpenRouter credits. Revert the OpenCode binary added to the
example Anthropic profile.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(readme): link first-agent guide and add agent skills section

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Co-authored-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-09-25 13:53:45 -07:00
Johnny Greco d7f921190b docs(policy): refresh policy documentation and references (#3563)
* docs(policy): correct schema and default policy guidance

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): add network recipes and update command reference

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): organize lifecycle guidance and troubleshooting

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): split policy overview into concepts and management tasks

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): reorganize network recipes as a cookbook

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): restructure schema reference by field group and protocol

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): align troubleshooting, advisor, and reference pages

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): fix first policy tutorial and security guidance

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): keep overview high level and move network rules to their own page

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): focus policy management on CLI workflows and remove command reference

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): clarify policy views and sandbox deletion in management guide

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): streamline network rule concepts and examples

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): correct request path wildcard semantics

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): rewrite policy advisor guide for clarity

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): clarify policy advisor scope, setup, and review

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): rewrite policy prover guide for clarity

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): explain the two uses of the policy prover

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): describe policy prover uses, boundaries, and coverage

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): place prover before advisor and troubleshooting last

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): remove unsupported CI guidance from prover page

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): tighten policy prover introduction

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): move policy change behavior into management guide

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): name prover check types and note expanding coverage

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): prefix prover and advisor sidebar labels

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): streamline policy schema reference

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): place default policy before schema reference

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): fold troubleshooting into policy management guide

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): correct tutorial log samples and GitHub push policy steps

The first policy tutorial said the 403 body begins with error, policy, and
rule, but the proxy serializes the body with sorted keys. Its log samples also
showed the wrong CONNECT deny reason for a sandbox without network rules, and
the L7 deny sample omitted the :443 authority, the `l7` engine, and the reason
tag that the shorthand formatter emits.

The GitHub tutorial filtered denials with `--level warn`, which hides the INFO
level OCSF policy events, and showed the retired key=value log format. Its
hand-written policy also omitted /bin from the restrictive default, so
`policy set` would reject the file for removing a filesystem path on a live
sandbox. Start from `policy get --base` and add only the network rules.

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): improve flow and terminology across policy pages

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): correct network rule matching and protocol details

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): align policy management steps with CLI behavior

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): correct policy advisor proposal and approval details

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): correct policy section, default, and schema details

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): correct prover installation and coverage limits

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): recommend tls skip for server-first protocols

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): fix stale baseline path and interpreter examples

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): move policy pages under how-it-works and fix links

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): align native TCP guidance in security best practices

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): restore policy.local and policy DNS details from main

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): state exact glob matching rules

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

---------

Signed-off-by: Johnny Greco <jogreco@nvidia.com>
2026-09-25 18:19:16 +00:00
Drew Newberry 9244868056 docs: refresh architecture and agent guides (#3705)
* docs: refresh architecture and agent guides

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs: describe updated security architecture neutrally

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs: highlight new isolation primitives

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(sandboxes): clarify how to disconnect

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs: align architecture and guides with current navigation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(extensibility): streamline extension authentication guidance

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-25 02:38:00 -07:00
Drew Newberry 08548713c5 fix(install): honor pinned releases and speed up prerelease discovery (#3681)
* fix(install): use native packages for prereleases and speed up discovery

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(install): honor pinned versions and guard Snap migration

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(install): omit Snap migration guard

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-24 17:57:42 -07:00
Piotr Mlocek f213e9a49b docs(middleware): reorganize and expand middleware guides (#3636)
* docs(extensibility): reorganize extensibility and middleware guides

Add an extensibility overview that introduces the extension protocol and
links every extension point, and move extension caller authentication to
a shared page used by middleware and gateway interceptors.

Split the supervisor middleware guide into an overview, a Configure and
Operate guide, and a Middleware Operations reference for service authors.
The overview explains when to use middleware, shows where it runs, lists
current limitations, and defines the service contract. Configure and
Operate covers policy attachment, service registration, failure behavior,
and observability. Middleware Operations describes HTTP request, HTTP
response, and WebSocket message operations with shared inputs and results,
per-operation diagrams, and detail accordions.

Pin page slugs so links resolve, redirect the replaced dev middleware
URL, and update reference and architecture links.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(middleware): clarify navigation and service contracts

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(middleware): drop obsolete dev URL redirect

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

---------

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-09-24 23:43:31 +00:00
Oliver Calder 6c864ec9ab fix(install): avoid installing incompatible docker snap (#3666)
* fix(install): avoid installing incompatible docker snap

The work to land RFC-0012 added new restrictions when interacting with
Docker by setting `NoNewPrivs`. This prevents the `docker` snap from
transitioning its AppArmor profile from `snap.docker.dockerd` to
`docker-default` when it tries to launch a container. Thus, the `docker`
snap is currently incompatible with OpenShell.

This commit prevents `install.sh` from installing the `docker` snap
before installing the `openshell` snap, and instead requires the user to
install a non-snap Docker daemon before proceeding with installing the
snap. Systems without the `snap` command are unaffected, since they
install native packages without checking for the presence of Docker or
other compute providers.

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* fix(snap): only snapd 2.76 for openshell snap since store installs work

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* fixup! fix(install): avoid installing incompatible docker snap

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* fixup! fix(snap): only snapd 2.76 for openshell snap since store installs work

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

---------

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>
2026-09-24 19:46:00 +00:00
Oliver Calder e60098d748 fix(snap): install openshell snap via install.sh when snap available (#3656)
* fix(snap): update stale snap docs and tests

Previously, the `openshell` snap required the `docker` snap. Now, it
works with any Docker daemon running on the system. Furthermore, the
`snap-declaration` assertion on the `openshell` snap when installed from
the Snap Store causes the `openshell` snap to always connect to the
system `:docker` slot, rather than a slot provided by the `docker` snap.

This commit updates the documentation, including the `description` field in
`snapcraft.yaml`, to ensure that all information is correct and
up-to-date.

Additionally, some tests connected the `openshell:docker` plug to the
`docker` snap's `docker:docker-daemon` slot, which is inconsistent with
how the `openshell` snap operates when installed from the store. Update
those tests to connect to the system `:docker` slot as well.

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* fix(snap): require snapd 2.76 for openshell snap

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* fix(nix): align snap gateway reproducer timeout with release-canary

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* fix(snap): require snapd 2.77 for openshell snap

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* fixup! fix(snap): require snapd 2.77 for openshell snap

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* feat(snap): install openshell snap via install.sh when snap available

Change `install.sh` to install the `openshell` snap by default when
snapd is installed on the host. This installs the snap from the
`latest/stable` channel, which should match the most up-to-date release
tag on github.

If `OPENSHELL_VERSION=dev` is set for `install.sh`, then it will install
the `openshell` snap from the `latest/edge` channel, which matches the
latest dev release available on github.

The `openshell` snap currently requires Docker in order to function. If
a Docker daemon is already installed on the system, it will be used by
the `openshell` snap. Otherwise, `install.sh` will install the `docker`
snap first, wait for the Docker daemon to be ready, and then install the
`openshell` snap.

Also, update the `release-canary.yml` to split the `ubuntu-snap` job
into `ubuntu-snap-system-docker` and `ubuntu-snap-provisions-docker`,
which test the two aforementioned scenarios. Previously, `ubuntu-snap`
manually installed a given snap artifact as built from CI, but with
these new jobs, it instead uses `install.sh` to install the published
`openshell` snap from the `latest/edge` track, thus matching the
behavior of the other release canary jobs.

Make corresponding changes to the `nix` guest reproducer, documentation,
and tests.

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* fix(snap): configure local gateway authentication

The `openshell` snap runs the gateway as a systemd system service, which
runs as root. Thus, the mTLS certs are generated by root and stored in a
root-owned directory to which non-root users do not have access. For
this reason, the snap's `openshell-gateway-wrapper` script sets
`OPENSHELL_DISABLE_TLS=true`.

This commit ensures that the `openshell` snap's gateway allows
unauthenticated local access by writing a default `gateway.toml`
configuration file during the install hook, which runs after the snap is
first installed but before services are started. The config file
contains sets `allow_unauthenticated_users = true`.

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* fix(install): configure snap gateway authentication

Recently, a new install hook was added which writes a default config
file for the `openshell` snap to allow unauthenticated local access to
the gateway. This is because the gateway service runs as root and the
mTLS certificates are not accessible to non-root users.

However, the `install.sh` script installs the `openshell` snap from the
snap store, and the published version may not yet have that new install
hook. Or, the user may already have the snap installed, in which case
the install hook does not run. In either case, we need `install.sh` to
ensure that the config file is written to set the gateway auth to
`allow_unauthenticated_users = true`, and then restart the openshell
gateway service.

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* fixup! feat(snap): install openshell snap via install.sh when snap available

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

---------

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>
2026-09-24 15:10:08 +00:00
Drew Newberry a649aa42fc docs: add 0.1.0 upgrade guide outline (#3540)
* docs: add 0.1.0 upgrade guide outline

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs: add upgrade change provenance

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs: update 0.1.0 upgrade guidance

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs: move upgrade navigation below security

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs: remove release notes page

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs: announce OpenShell 0.1.0

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs: reorganize guides and require user-owned images

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs: reorganize navigation around core concepts

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs: simplify navigation and tutorial catalog

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs: refine 0.1.0 release highlights

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs: align navigation and 0.1.0 guidance

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-23 20:26:25 +00:00
Yuedong Wu 718dba3430 fix(policy)!: reject removed tls endpoint values (#3414)
Signed-off-by: Yuedong Wu <dwcn22@outlook.com>
2026-09-22 17:58:19 +00:00
50230616d5 refactor(runtime): retire Community image dependencies (#3386)
* feat(sandbox): default to official Alpine sandbox image

default_sandbox_image() now returns docker.io/library/alpine:3.22, a generic
version-qualified official image, so a fresh install no longer depends on the
community sandbox image catalog. All compute drivers (docker, podman,
kubernetes, vm) inherit this fallback.

Part of #3116.

Signed-off-by: Akram
Signed-off-by: Akram <akram.benaissi@gmail.com>

* feat(deploy): default deployment configs to the official Alpine sandbox image

Update the shared gateway default_image, Helm chart values, the standalone
Kubernetes manifest, and the dev gateway task scripts to use
docker.io/library/alpine:3.22 instead of the community base image, consistent
with default_sandbox_image(). GPU e2e image-build base is left unchanged (CUDA
needs a glibc base).

Part of #3116.

Signed-off-by: Akram
Signed-off-by: Akram <akram.benaissi@gmail.com>

* feat(driver): default to numeric non-root identity for USER-less images

With the default sandbox image now Alpine, images that declare no OCI USER
must start instead of being rejected. When the image declares no USER and
the policy requests none, the Podman and Docker drivers now supply a numeric
non-root identity (DEFAULT_SANDBOX_UID/GID = 1000) instead of rejecting,
matching the numeric-identity behavior of the Kubernetes and VM drivers. The
supervisor's resolved-identity path runs the sandbox as a synthesized
non-root account without the account existing in the image. Images that
declare a USER keep the OCI resolution path unchanged.

Part of #3116.

Signed-off-by: Akram <akram.benaissi@gmail.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(conformance): use Alpine workload image

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* refactor(policy): drop community image /app path from default policy

The restrictive default policy granted read-only access to /app, a directory
that only existed in the community base image. A generic Alpine default has no
/app, so remove it. Landlock best-effort already ignores absent paths; this
just stops advertising a community-specific layout in the default.

Part of #3116.

Signed-off-by: Akram
Signed-off-by: Akram <akram.benaissi@gmail.com>

* docs(config): document Alpine default images

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(podman): report early sandbox termination

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(podman): initialize rootless workspace ownership

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(sandbox): qualify NVIDIA Ubuntu default

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): initialize rootful default workspace

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(sftp): add native sandbox adapter

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sftp): gate runtime helper support to Linux

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sftp): support standard OpenSSH file operations

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sftp): harden rename and special file handling

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(runtime): remove community image dependencies

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(e2e): build provider readiness tool fixture

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(e2e): use a dedicated Noble fixture for Docker tests

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Akram
Signed-off-by: Akram <akram.benaissi@gmail.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Evan Lezar <elezar@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
2026-09-22 14:43:51 +02:00
Drew Newberry 17ce738bfb fix(ci)!: remove gateway callback listener dependency (#3365)
* fix(ci): repair post-merge release canary

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(packaging): bootstrap canary runtime prerequisites

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci(canary): collect macOS VM diagnostics

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci(canary): pin libkrun-compatible macOS runner

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci(canary): limit macOS smoke test to package startup

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(compute)!: remove gateway callback listeners

Run Docker supervisors on host networking so they use the operator-configured primary gateway endpoint. Remove the unused compute-driver callback listener negotiation and listener-scoped routing machinery.

BREAKING CHANGE: The ComputeDriver API no longer exposes GetGatewayListenerRequirements or GatewayListenerRequirement. External drivers must regenerate bindings and connect supervisors to the configured primary gateway endpoint.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(docker): use sandbox runtime image in launcher

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(podman): exercise production endpoint selection

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(e2e): route supervisors to reachable gateways

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(e2e): align Podman endpoint fixtures

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): preserve host aliases for supervisors

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): align sandbox host gateway pin

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(e2e): address Docker fixtures by bridge IP

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(e2e): serialize sandbox lifecycle cases

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(e2e): host Docker TCP fixture with gateway

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(e2e): use loopback for host-network supervisor

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-18 20:55:55 +00:00
Drew NewberryandPiotr Mlocek 8b77925eb7 feat(installer): support prerelease installations (#3364)
* feat(installer): support prerelease installations

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(readme): clarify 0.1.0 production readiness

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(readme): refine production readiness message

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(release): publish rolling prerelease channels

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(release): keep prereleases out of GitHub releases

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(readme): refine 0.1.0 notice

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(readme): expand 0.1.0 notice

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(readme): simplify 0.1.0 notice

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(installer): simplify prerelease alias

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(installer): require successful prerelease runs

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(release): reuse exact prerelease artifacts

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(installer): include prover in prerelease bundle

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci: remove prerelease post-publish validation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(fern): sync announcement configuration

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs: remove stale prerelease canary claim

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(fern): defer announcement synchronization

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(fern): simplify release announcement

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Co-authored-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-09-18 00:53:36 +00:00
Philippe MartinandJohn Myers 04146692d9 refactor(providers)!: make provider profiles import-only (#3383)
* chore(providers): remove dead provider plugin modules

Twelve modules under crates/openshell-providers/src/providers/ were never
declared in providers/mod.rs, so they have not been compiled since the plugin
registry was narrowed to the two adapters it still registers. Four of them
(generic, gitlab, opencode, outlook) key off provider type identifiers that
normalize_provider_type already retired.

Keep google_cloud and vertex, which are the only plugins ProviderRegistry::new
registers.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* test(providers): load example profiles from providers/ at test time

Add an example_profiles module that reads the YAML under providers/ from the
source checkout at run time and parses it with the existing profile loader. It
is gated behind a new non-default example-profiles feature so it is available to
this crate's own tests and, once wired into dev-dependencies, to gateway and CLI
tests, while never reaching a release binary.

Nothing consumes it yet; later changes move the test fixtures off the compiled
catalog and onto these files.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* test: source profile fixtures from providers/ instead of the compiled catalog

Unit and integration fixtures reached into builtin_profiles() to get a profile
to work with, which ties the tests to the compiled catalog rather than to the
files an operator would import. Point them at the example_profiles loader
instead.

The profiles.rs tests keep their coverage unchanged and become explicit golden
tests over providers/*.yaml: they are what keeps those files valid once nothing
compiles them. builtin_profiles_are_sorted_by_id widens into a test that the
whole example set parses, sorts, and lints clean as one catalog.

The CLI fake gateways serve the example profiles the way a real gateway serves
what an operator imported, through a shared helper.

No behavior change: the gateway still loads the built-in source by default and
serves the same profiles.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* docs(providers): document the example provider profiles

Each file under providers/ now opens with a header naming its expected client
binary identities, the image layout those paths assume, the credential scope,
the endpoint access it grants, and a smoke test. Add a README covering the
import commands and why a profile should be copied and edited rather than
imported unchanged.

Several of these profiles bind network access to paths that only exist in the
OpenShell Community image — /sandbox/.venv, /app/.venv, /sandbox/.cursor-server,
/usr/lib/node_modules. Imported unchanged into another image the profile matches
nothing: the catalog still advertises it, but the credential is never injected
and the traffic is denied. The headers say so where it applies.

Comments only; the profile schema has no documentation fields and all fifteen
files still lint clean.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* test(server): seed unit-test state with the example provider profiles

Server unit tests inherited the builtin + user source default from ServerState,
so around a hundred and forty assertions about github, openai and the rest
resolved against the compiled catalog. Point the test state at the user source
alone and import the example profiles from providers/ into its store first, the
way an operator would. Every one of those assertions keeps passing unchanged,
which is the point: it proves the gateway behaves identically with an imported
catalog before the default moves.

Four tests asserted builtin-source semantics specifically. Under import-only the
only profiles a gateway cannot edit are the ones a non-user source vends, so
they now exercise a source-managed profile composed with the user source; the
read-only guard they cover is the one that still applies to interceptor
catalogs. The list test asserts the imported profile's user/platform identity
instead of builtin with an empty scope.

test_server_state_with_user_only_github_profile is gone: the default test state
now is a user-only gateway with github imported.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* refactor(cli): resolve credential suggestions from the gateway catalog

The --env credential warning scanned a profile table compiled into the CLI, so
its suggestions described the binary rather than the gateway the user is talking
to: a profile the gateway does not serve was suggested anyway, and an imported
custom profile never was.

Fetch the catalog from the connected gateway instead. The warning moves out of
argument parsing and into sandbox create and sandbox template create, where a
client already exists. A catalog fetch failure is not fatal — the warning
degrades to its generic form rather than blocking sandbox creation.

Extract the ListProviderProfiles paging loop from provider list-profiles into a
shared fetch_provider_profile_catalog, which the profile-driven paths now share.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* refactor(cli)!: infer providers from the gateway catalog

Command-to-provider inference went through a hardcoded alias table in
openshell-providers, a second copy of the built-in catalog's identifiers. A
custom profile could never be inferred no matter what binaries it declared, and
the table drifted from the profiles it mirrored.

Infer from the connected gateway's catalog instead: match the command's basename
against each profile's ID and against the basenames of the binaries the profile
authorizes. A profile that names /usr/bin/claude is the profile for running
claude. The match must be unique — where several profiles claim a command, the
user names one with --provider — and an empty catalog infers nothing.

No command is special-cased. `binaries` is the operator's authorization
statement, so a profile that declares a binary claims the command that runs it,
whatever that binary is; narrowing that belongs in the profile rather than in a
list compiled into the CLI, which could never cover an unbounded catalog anyway.

Breaking: the retired aliases stop resolving, and commands the old table never
listed can now infer. git, pip and uv are declared by the github and pypi
example profiles, so they infer where they previously did not.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* refactor(providers,server)!: resolve provider profiles by exact ID

normalize_provider_type was a hardcoded alias map — gh to github, claude to
claude-code, vertex to google-vertex-ai — and a second copy of the built-in
catalog's identifiers, independent of the YAML it mirrored. It let a provider
type resolve to a profile the operator never named, and it made the built-in IDs
behave as a reserved namespace.

Remove it, along with detect_provider_from_command and the alias-normalizing
ProviderRegistry::inject_env. A provider type now names a profile exactly:

- the effective catalog resolves an ID or reports it absent, with no alias retry
- plugins activate only for the ID of a profile the gateway resolved; a provider
  with no resolvable profile gets no plugin projection, instead of falling back
  to an alias guess
- the CLI surfaces the gateway's not-found instead of retrying under an alias
- telemetry buckets by profile ID, and the gitlab, opencode and outlook buckets
  go with the aliases that were their only source

Breaking: `--type gh`, `--type claude` and the other aliases no longer resolve.
Use the profile's own ID.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* feat(server): fail closed when a provider's profile is absent

A sandbox composed from a provider whose profile the gateway cannot resolve
started anyway: the credential and policy builders warned and skipped, so the
sandbox came up carrying none of that provider's credentials or network policy.
The operator learned about it later, as a denied connection or a missing
environment variable, rather than as the configuration error it is.

Check at the two composition boundaries — CreateSandbox and
AttachSandboxProvider — right beside the catalog snapshot already taken there,
and reject with a bounded diagnostic naming the provider, the profile it refers
to, and the import command that supplies it. Scope-aware: a platform-scoped
provider is told to import with --global.

Read paths are untouched. ListProviders, GetProvider and profile export keep
working so an operator can see and recover an affected provider, and the shared
policy builders keep their warn-and-skip for the diagnostic paths that also
reach them.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* refactor(server): scope the vendor base-URL pin to declared endpoints

provider_profile_endpoints_are_active withheld a credential and the provider's
policy layer when an openai or anthropic provider pointed its client somewhere
other than the public vendor endpoint. The guard was keyed on
profile.source == "builtin" and on those two profile IDs, so it protected only
profiles OpenShell shipped. Once profiles are import-only no profile is ever
builtin, and the guard would silently stop applying — including to an operator
who imported providers/openai.yaml verbatim.

Key it on the profile instead of on where the profile came from. A profile's
endpoints are the boundary its credential is bound to, so if the provider
configures a *_BASE_URL pointing at a host the profile does not declare, the
profile no longer describes where that credential goes and is treated as
endpointless. Host matching reuses the DNS-label-aware matcher in
openshell-core, so wildcard endpoints such as Vertex's
*-aiplatform.googleapis.com resolve correctly.

A profile with no declared endpoints has no boundary to contradict, and a config
value that names no host is not a redirect this can reason about. Both keep the
profile active. The control now covers every endpoint-bearing profile, including
an operator's own, and the last "builtin" string leaves the gateway.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* test(e2e): import example provider profiles during gateway bring-up

The e2e suites create providers from github, openai, nvidia, claude-code and
google-cloud, and the GitHub lane runs a real git clone through the profile's
binary attribution. A gateway serves only the profiles an operator imported, so
the lanes have to import them.

Add e2e_import_example_provider_profiles to the shared bring-up helpers and call
it from the Docker, Podman, Kubernetes and VM wrappers once the gateway is
healthy and registered. It runs the same command the upgrade notes give
operators, against the repository's own providers/ directory, so the lanes
exercise the documented path rather than a test-only shortcut.

Lands before the default changes: a same-ID user profile already shadows a
built-in, so importing works today.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* refactor(providers)!: make provider profiles import-only

OpenShell compiled fifteen provider profile YAML files into every release binary
and selected them by default, so a fresh gateway published a catalog it was
never configured with. Those profiles are not image-neutral: their binary
selectors name paths that exist in the OpenShell Community image, so changing
the sandbox image could make a profile inert while the catalog still advertised
it. Every endpoint, credential name and binary path in providers/ was also
effectively part of the 0.1.0 public contract.

A gateway's catalog is now exactly what an operator imported:

- providers/*.yaml is no longer include_str!'d, and builtin_profiles() is gone
  from the openshell-providers API
- the default provider_profile_sources is [{ type = "user" }], and a gateway
  with nothing imported reaches ready and serves an empty catalog — an empty
  catalog is a valid state, not a startup failure
- the builtin source type is removed from the configuration schema, and a
  gateway.toml that still names it is rejected at parse time with the import
  command rather than an unknown-variant error
- nothing reserves the canonical identifiers any more, so github, pypi,
  anthropic and the rest import at their own IDs and the imported profile is the
  only definition for that ID; the static_fallback that kept a shipped
  definition resident behind an imported one is gone with the source that
  produced it

A collision between an interceptor-vended profile and an imported one still
fails closed, which is the behavior that shadowing quietly bypassed for
built-ins.

Operators upgrading should export the profiles their deployment relies on
before upgrading, or copy them from the providers/ directory of the matching
release tag, then import them at the scope their providers use.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* docs(providers): document import-only provider profiles

Rewrite the provider profile documentation around a catalog the operator builds
rather than one the gateway ships.

- gateway-config: the default is [{ type = "user" }], the builtin source type is
  gone and rejected at startup, and an empty catalog is a valid ready state
- profiles: replace the built-in profile table and its shadowing semantics with
  the import workflow, and say plainly that a profile whose binary paths do not
  match the image is inert
- manage-providers: replace the two fixed provider-type tables with
  `provider list-profiles`, and describe catalog-driven command inference
- inference-routing: drop the "still loads built-in profiles for compatibility"
  transition text; its migration walkthrough is now the normal path
- quickstart and the Docker Compose, GitHub, AWS, Google Cloud and Vertex
  tutorials: import the profile before creating the provider, since nothing
  resolves without it
- release notes: a 0.1.0 migration section covering export-before-upgrade,
  importing from the release tag's providers/ directory, what happens to a
  provider whose profile is missing, and the removal of the legacy type aliases
- README, architecture, the openshell-cli and debug-openshell-cluster skills,
  and the governance interceptor example follow the same change; the example's
  smoke assertion now imports a profile to prove an authoritative interceptor
  hides it

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(cli,e2e): distinguish catalog failures and authenticate OIDC seeding

Addresses two findings from review (GATOR-c1bd9867-02 and -03).

The provider profile catalog lookup collapsed its error into an empty
catalog, so an authorization, availability or transport failure on
ListProviderProfiles was indistinguishable from a gateway that genuinely has
no profiles. Command inference then resolved nothing and the sandbox was
created without the provider it needed, deferring the failure to the workload.

Keep the lookup's outcome instead of discarding it. The credential warning is
advisory and still degrades to its generic form, but inference now consults the
catalog only when there is a command to resolve and surfaces the lookup failure
when there is, naming the gateway and pointing at explicit --provider selection.
Two regression tests cover it through a fake gateway whose ListProviderProfiles
returns UNAVAILABLE while every other RPC succeeds: sandbox creation fails
without sending a provider-less create request, and a sandbox with no trailing
command still succeeds because it needs no catalog.

The e2e profile seeding also ran unauthenticated in the OIDC lanes. Those lanes
deliberately skip gateway registration and start the gateway without a TLS
client CA, so no mTLS identity exists and no token has been acquired when the
import runs; the wrapper exited during setup. Skipping the import is not
sufficient because the provider tests now require the claude-code profile.

Add e2e_register_oidc_admin_session, which mints an administrator token with
Keycloak's password grant — the same grant the OIDC test helpers use — and
writes the gateway metadata and token bundle that an interactive login would
have stored, so the import runs as an authenticated administrator. The mTLS
lanes keep the direct import unchanged. Both affected wrappers are covered:
with-podman-gateway.sh had the same defect as with-docker-gateway.sh.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* refactor(cli)!: remove command-derived provider attachment

Addresses GATOR-c1bd9867-01.

A profile's `binaries` list authorizes a binary to reach that profile's
endpoints. It is not a statement that running the binary asks for the provider,
and reading it as attachment intent let a command silently gain provider
authority: the aws-s3 example declares /bin/bash, so `sandbox create -- bash`
resolved an existing aws-s3 provider and attached its credential-backed
capability and network policy to the shell and its descendants. Attachment of
an already-created provider never prompted, so the confirmation flow did not
guard it.

The distinction that would make inference sound — whether a declared binary is
a profile's client or merely a permitted runtime — cannot be expressed:
NetworkBinary carries only a path, and the field that encoded it was removed in
0.1.0. Any substitute is a guess. Restricting the guess to a unique claimant
does not help, because uniqueness measures how sparse the catalog is rather
than what the user intended, and a compiled list of "generic" commands could
never cover an unbounded operator catalog.

Remove trailing-command inference rather than approximate it. A provider is
attached only when named with --provider, which still creates a missing
provider from local discovery when the name matches an imported profile ID.
Sandboxes with no providers remain a normal, fully supported state.

Removing inference also settles GATOR-c1bd9867-02: with no consumer deriving
authority from the catalog, its only remaining use is the advisory credential
warning, so a failed lookup degrades that warning instead of blocking creation.
The regression test now asserts that an unreachable catalog still creates the
sandbox and attaches nothing.

While repurposing the deduplication test, a pre-existing defect surfaced:
repeating a name in --provider auto-created it twice and the second attempt
failed with "provider already exists", because the explicit pass never
consulted the set of names it had already handled. Guard it.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(e2e): load the OIDC token and trust the gateway when seeding profiles

Addresses the carried finding GATOR-c1bd9867-03.

e2e_register_oidc_admin_session established a session the CLI never used. The
CLI decides whether to load a stored bearer token by matching on the auth_mode
field of the gateway metadata alone; the helper omitted that field, so the
metadata fell through to the default arm and oidc_token.json stayed on disk
unread. The profile import went out unauthenticated exactly as it had before
the helper existed. The helper also installed no trust anchor, leaving the CLI
unable to verify the gateway's self-signed serving certificate.

Write auth_mode = "oidc" so the stored token is loaded, and install the CA at
<gateway>/mtls/ca.crt. Only the CA is installed: with no client certificate or
key on disk the CLI falls back to CA-only server verification and authenticates
with the bearer token, which is what these lanes need because they start the
gateway without --tls-client-ca. Certificate verification stays on; no insecure
transport override is introduced.

Assert the session before anything depends on it. ListProviderProfiles is
annotated auth_mode: "bearer", so it cannot succeed unless the token was loaded
and accepted. The helper now fails at that point, naming the gateway config
directory and echoing the CLI output, rather than letting the defect surface
later as an opaque profile import error.

Both affected wrappers pass the PKI directory and CLI binary the helper needs;
the mTLS lanes are untouched.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(e2e): request the OpenShell scopes when minting the admin token

The OIDC lanes failed at the session assertion added in b25dd0298 with
PERMISSION_DENIED and "scope 'provider:read' required". Authentication was
working -- the gateway logged ListProviderProfiles at gRPC status 7, which it
can only reach once the bearer token has been loaded and accepted. The token
simply carried no OpenShell scope.

sandbox:*, provider:*, config:*, workspace:* and openshell:all are optional
client scopes on the openshell-cli client in scripts/keycloak-realm.json, so
Keycloak mints them only when the request asks for them. The password grant
here asked for nothing, leaving the realm defaults (openid, profile, email,
roles, web-origins, acr) and an access token that authorizes no RPC.

Request "openid openshell:all", as e2e/python/oidc/oidc_auth_test.py already
does for its administrator tokens. openshell:all is SCOPE_ALL in
crates/openshell-server/src/auth/authz.rs, so one scope covers the setup calls
without enumerating them. Verified against the realm: the token goes from
"email profile openid" to "email profile openshell:all openid".

Record the same scopes in metadata.json. The helper writes the bundle
openshell gateway login would have stored, and oidc_scopes is the field that
login path reads back, so leaving it out would misdescribe the stored token.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(e2e): seed provider profiles once per VM lane gateway

The VM lanes failed on the second test target with "custom provider profile
'anthropic' already exists" for all fifteen example profiles, and the import
exited non-zero.

e2e_import_example_provider_profiles was called from run_e2e_test, so it ran
once per target -- four times against one long-lived gateway. That was
harmless while import overwrote silently, but profiles are now import-only:
ImportProviderProfiles is create-only and reports an existing id as an
error-severity diagnostic, with no overwrite flag on the request. The first
import therefore succeeds and every later one fails.

Hoist the call to just after the conformance run, which is where the docker,
podman and kube lanes already seed their catalogs. The profiles persist for
the gateway's lifetime, so every target still finds them, and the
E2E_TEST_OVERRIDE path is covered by the same single call.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* test(e2e): name the provider explicitly in the OIDC workspace-user test

user_can_create_sandbox_with_inferred_provider_command reached the gateway
once the admin session was fixed, and then failed: it asserted a
missing-provider error, but the sandbox was created and died provisioning
with "failed to spawn sandbox entrypoint process 'claude-code'".

The test drove provider resolution by passing claude-code as the trailing
command and relying on the CLI to infer the provider type from it. That
inference is what this branch removed, so the trailing word is now nothing
but an entrypoint, and the image has no such binary.

The regression the test guards is not inference itself: it is that a
workspace user resolving a provider is not gated behind Platform Admin.
Name the provider with --provider, the only remaining way to attach one.
The CLI still has to fetch the claude-code profile before it can auto-create
the provider, so the lookup a workspace user must be allowed to make still
happens, and auto-creation still stops at the non-interactive branch with
"missing required provider". Both assertions therefore keep their meaning.

Rename the test and rework its comments to describe what it now exercises;
the old name would otherwise outlive the behavior it was named for.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(providers): reconcile import-only profiles with main

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: Philippe Martin <phmartin@redhat.com>
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-09-17 22:59:27 +00:00
Johnny Greco 58b5f8f976 feat(prover): add standalone policy boundary checker (#3289)
* feat(prover): add standalone policy maximum checker

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* refactor(prover): simplify check scope schema

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(prover): align containment and cancellation with runtime

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(prover): stabilize containment checks in CI

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* test(prover): avoid solver in fast-path guard test

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(prover): align string containment with runtime

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(prover): reject ambiguous z3 string escapes

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(prover): align containment with runtime boundaries

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(prover-cli): harden cancellation and invalid input

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* feat(packaging): install policy prover with OpenShell

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(prover): clarify installation and check results

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(build): describe prover distribution directly

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* refactor(prover): rename maximum policy to boundary

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* refactor(prover): localize fail-closed validation

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* test(prover): cover fail-closed CLI surfaces

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* test(prover): allow CI load for REST solver proof

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(prover): use canonical policy schema for containment

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(prover): preserve uncertainty for runtime binary globs

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(prover): bound policy validation work

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(prover): document validation resource limits

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(prover): make containment API extensible

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(prover): define containment API contract

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(ci): integrate prover with consolidated builds

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(ci): declare release packaging dependency

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

---------

Signed-off-by: Johnny Greco <jogreco@nvidia.com>
2026-09-17 17:30:03 +00:00
VarshaandJohn Myers 0803c4aa4c refactor(policy)!: remove NetworkBinary harness field (#3222)
* refactor(policy)!: remove NetworkBinary harness field

Closes #3054

Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>

* test(policy): preserve unknown fields during migration

Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>

* fix(policy): preserve advisor provenance during merge

Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>

* fix(policy): remove redundant network binary default

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(policy): record harness schema migration

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-09-11 16:49:14 +00:00
Jesse JaggarsandDrew Newberry 02b664bb0d refactor(config): normalize and enforce gateway schema v2 (#2814)
* refactor(config): normalize compute driver field names

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* refactor(config): introduce canonical gateway fields

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* refactor(config): enforce gateway schema version 2

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): preserve compute driver runtime guarantees

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): address schema v2 review regressions

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): complete schema v2 migration safeguards

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(config): expand schema v2 regression coverage

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(config): add schema v2 parity manifest

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): correct parity manifest inventory

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* docs(config): record schema v2 intentional changes

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* docs(config): disposition schema v2 parity gaps

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): add dual schema parity harness

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): establish compute lifecycle parity baseline

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): preserve gateway option compatibility

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): record gateway option parity

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* docs(config): close gateway-wide parity gaps

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(podman): apply configured pids limit

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): validate Podman option parity

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): add Kubernetes option parity harness

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): record Kubernetes option parity

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): disposition VM parity lanes

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): add external driver parity lane

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(e2e): preserve external driver pull policy

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): attest parity artifacts and launches

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): require clean parity build sources

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): bind parity runtime artifacts

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(e2e): use isolated supervisor tags

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(e2e): qualify parity image tags

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(e2e): serve parity supervisor locally

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): isolate parity podman services

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): harden parity evidence provenance

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): pin parity sandbox artifacts

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): attest parity runtime inputs

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): bind parity runtime evidence

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): record compute boundary parity

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): disposition cross-cutting parity lanes

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(packaging): preflight gateway config upgrades

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): preserve rebase integration guarantees

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(ci): isolate temporary git signing config

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): update remaining schema v2 consumers

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(ci): provide e2fs tools to VM tests

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): align preflight with gateway startup

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(vm): preserve rootfs tar configuration

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* chore(config): adopt duration unit constructors

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(packaging): preflight RPM gateway config

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): address driver review findings

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(e2e): require fresh semantic parity evidence

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(docker): update tests for renamed sandbox label

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(gateway): preserve selective driver coverage after rebase

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
2026-09-11 05:00:24 +00:00
John T. Myers f4dc6be4b2 refactor(inference): remove managed inference routes (#3195)
* refactor(inference): remove managed inference routes

Closes #3172

Remove the inference route control plane, inference.local data path, built-in router crate, and SDK surface. Move inference workloads to explicitly imported provider profiles and native endpoints, with migration cleanup and updated tests and documentation.

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(policy): preserve alternate upstream isolation

Restore the provider policy activation guard so legacy OpenAI and Anthropic providers configured for alternate base URLs do not grant egress to the built-in public vendor endpoints.

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-09-09 18:47:22 +00:00
Simon Scatton 455883905a fix(python): remove CLI from wheel (#2321)
The Maturin-based wheel packaging was a historical remnant from when the local gateway launch path and OpenShell CLI were coupled in one binary. The gateway and CLI now ship as standalone artifacts, so the Python distribution should contain only the SDK.

Build a single platform-independent setuptools wheel, verify that it cannot contain native code or an openshell entry point, and simplify the release jobs and documentation for SDK-only PyPI installs.

Signed-off-by: Simon Scatton <sscatton@nvidia.com>
2026-08-25 13:20:42 +00:00
Drew Newberry f12f3ef8d5 fix(macos): restore Homebrew sandbox callbacks (#2739)
* fix(macos): restore Docker gateway callbacks

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(gateway): reuse reachable primary callback listener

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-08-14 18:38:20 +00:00
Evan LezarandDrew Newberry d220d89468 feat(compute): negotiate gateway callback listeners (#2492)
* feat(compute): query gateway listener requirements

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* feat(compute): add Podman listener requirements

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(docker): use default gateway bind address

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(gateway): avoid wildcard primary listener

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(podman): validate callback listener discovery

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(server): support split dual-stack listeners

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(podman): support legacy rootless listener discovery

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(e2e): accept loopback plaintext rejection

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* docs(agent): add callback listener diagnostics

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(server): restrict compute callback listeners

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(podman): validate local callback port

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(server): clarify callback listener contract

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(podman): require pasta for local callbacks

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* docs(gateway): document RPM listener default

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* refactor(server): keep listener provenance diagnostic-only

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(compute): preserve callback listener isolation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(e2e): remove Podman callback relay

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(packaging): preserve Podman callback loopback

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci(e2e): run VM smoke on nested-virt runner

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci(e2e): gate VM smoke on usable KVM

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci(e2e): probe KVM through VM driver

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci(e2e): tolerate hosted KVM denial

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(server): close traced futures before assertions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* revert: remove tracing test stabilization

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Evan Lezar <elezar@nvidia.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
2026-07-31 16:41:06 +00:00
Oliver Calder 994750e3a1 feat(snap): vendor ssh in openshell snap and remove ssh-keys interface (#2280)
Previously, the openshell snap used the ssh-keys interface to get access
to the host's ssh binary, which is used for sandbox connect/exec/forward.
However, ssh-keys is a privileged interface which also grants access to
the public and private ssh keys on the host. As such, it required manual
connection in order to be used.

This weakened the security sandbox of the snap, and hurt the UX of
installing it.

This commit changes this by removing the `ssh-keys` interface and
instead vendoring the `ssh` binary within the snap.

This is safe because OpenShell always invokes the `ssh` binary with
`StrictHostKeyChecking=no`, `UserKnownHostsFile=/dev/null`, and
`GlobalKnownHostsFile=/dev/null`, and never uses any host credentials or
ssh configuration. Openshell only ever access to `~/.ssh/config` to
write OpenShell-managed aliases, and this can safely live within the
snap sandbox, rather than leaking into the host environment.

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>
2026-07-15 03:05:33 +00:00
Mesut Oezdil ed8ce8208f docs: fix Docker version format from 28.04 to 28.0 (#2136) 2026-07-08 16:57:07 -07:00
Mesut Oezdil f852d07b6b docs: add Hermes Agent to supported agents table (#2131) 2026-07-03 17:55:26 +00:00
Piotr Mlocek e4d7d41656 fix(snap): use snap-owned XDG directories (#1972) 2026-06-25 11:46:34 -07:00
Piotr Mlocek 82d03f1c98 fix(linux): lower host glibc floor to 2.28 to support RHEL/Rocky 8 (#1934) 2026-06-22 14:09:26 -07:00
Zygmunt Krynicki 4025894a94 chore(snap): remove early snap packaging (#1648) 2026-06-08 22:09:09 -07:00
Eric Curtin 1c8417c4dd docs(container-gateway): fix Docker driver setup for containerized gateway (#1419)
The existing docs omitted or misstated several requirements when running
the gateway as a container with the Docker compute driver:

- OPENSHELL_GRPC_ENDPOINT is required; the Docker driver uses only the
  scheme (http/https) — host and port are substituted automatically with
  host.openshell.internal and the gateway's own bind port
- Supervisor binary must be extracted to a host path before starting the
  gateway; bind-mount sources are resolved by the host Docker daemon so
  the path must be identical inside and outside the gateway container
- Docker socket access requires adding the docker group (UID 1000 default)
- Port binding should remain 127.0.0.1; Docker driver adds a bridge
  listener automatically
- add --server-san host.openshell.internal to generate-certs for mTLS
- Complete the mTLS docker run with all Docker driver requirements
- Add deploy/docker/gateway.toml — TOML config for the Docker driver
- Add deploy/docker/docker-compose.yml referencing the TOML
- Add docs/get-started/tutorials/docker-compose.mdx tutorial page
- Remote gateway registration instructions (--remote flag)

Address reviewer feedback:
- Move Docker Compose tutorials card to the bottom of the list
- Replace inline YAML snippet in Docker Compose section with a reference
  to deploy/docker/ to avoid drift
- Clarify OPENSHELL_DB_URL is safe in compose.yml (plain SQLite path,
  no credentials); the TOML block targets credential-bearing DSNs
- Note that ./ in source: resolves relative to the compose file directory
- Clarify that only the scheme from OPENSHELL_GRPC_ENDPOINT matters
- Add note that the tilde volume mount resolves to the same absolute
  path on both host and container
2026-06-03 09:41:41 -07:00
Adam Miller f061b1d923 feat(providers): add Google Vertex AI inference provider (#1568)
* feat(providers): add Google Vertex AI provider

Adds Vertex AI provider profiles, routing, credential refresh plumbing, CLI support, docs, and regression coverage. Keeps the related NETLINK_ROUTE seccomp allowance needed by Vertex client tooling that calls getifaddrs.

* docs: add Vertex AI sandbox usage for Claude Code and OpenCode

Cover the full end-to-end setup for running Claude Code and OpenCode
inside an OpenShell sandbox via inference.local with a Vertex AI backend:

- google-vertex-ai.mdx: add 'Use from a Sandbox' section with tabbed
  examples for Claude Code (--bare flag, no /v1 suffix) and OpenCode
  (/v1 suffix required). Add providers_v2_enabled prerequisite and
  --no-verify note for global region. Document policy proposals table
  covering metadata.google.internal (always blocked), downloads.claude.ai,
  and storage.googleapis.com.

- inference-routing.mdx: expand 'Use the Local Endpoint' section with
  tabbed examples for Claude Code, OpenCode, Python OpenAI SDK, and
  Python Anthropic SDK. Add notes explaining the /v1 path suffix
  difference between clients.

- supported-agents.mdx: update Claude Code and OpenCode rows to mention
  inference.local support and correct base URL requirements.

* fix: address vertex review findings

* test(sandbox): retry on spurious Ok in fork-exec ambiguity test

On arm64 under heavy CI load, the /proc fd scan in
find_socket_inode_owners can transiently miss the parent process's
socket fd entry, returning only the child as an owner. This causes
resolve_process_identity to return Ok (single owner, no ambiguity
check fires) instead of the expected ambiguous-ownership Err.

Extend the retry loop to also handle unexpected Ok results, mirroring
the existing retry for transient Err results. 10 retries at 50ms gives
a 500ms settling window, which is sufficient for procfs to stabilize
on loaded arm64 runners.

* fix: address vertex review regressions

* docs(router): clarify stream_response semantics for Vertex rawPredict routing

Document the three call sites of prepare_backend_request and their
stream_response values in a caller table:

- send_backend_request: false → :rawPredict (unary endpoint)
- send_backend_request_streaming: true → :streamRawPredict
- verify_backend_endpoint: explicitly false to probe the unary endpoint

Cross-reference the table from build_provider_url and
is_vertex_anthropic_rawpredict_route so the stream_response=true guard
in the suffix upgrade branch is understood in full context.

Also note that is_vertex_anthropic_rawpredict_route is a structural
predicate (model_in_path + anthropic_messages + :rawPredict suffix),
not a named-provider check, so any future provider with the same route
shape inherits the transforms automatically.
2026-06-02 10:45:50 -05:00
Vegard Stikbakke 9e5aee4a56 docs: add Pi as supported sandbox (#1572) 2026-05-26 13:53:19 -07:00
Piotr Mlocek 0dc08a185e fix(release): build host Linux binaries with glibc floor (#1490) 2026-05-22 12:12:03 -07:00
Drew Newberry 603b3e27fa docs: update NemoClaw/OpenClaw references (#1529) 2026-05-22 10:31:59 -07:00
Taylor Mutch a3b16c18ab feat(auth): per-sandbox authentication to gateway (#1404) 2026-05-21 17:58:55 -07:00
Drew Newberry f257ed0193 refactor(packaging): rely on gateway runtime defaults (#1415)
* fix(packaging): use gateway TOML config in packages

* refactor(packaging): rely on gateway runtime defaults
2026-05-18 14:13:19 -07:00
Eric Curtin 9f8edb5a45 docs(installation): add container gateway page with docker run and compose examples (#1321)
Add a dedicated page documenting how to run the OpenShell gateway as a
container using docker run, docker-compose, or podman, without the
system package manager installer.

This is useful for users on immutable OS distributions (Fedora CoreOS,
bootc-based images, Silverblue) where the standard install.sh path is
not appropriate, and for container-first environments.

Covers a quick-start with TLS disabled (localhost-bound), a full mTLS
setup using the gateway's generate-certs subcommand, a docker-compose
example, and a Podman variant.

Closes the gap raised in #1285.

Signed-off-by: Eric Curtin <eric.curtin@docker.com>
2026-05-15 09:22:19 -07:00
Drew Newberry daa2a362d5 fix(packaging): enable mTLS for local packages (#1271) 2026-05-08 13:02:12 -07:00
Taylor Mutch d8f614fa31 docs(kubernetes): add initial reference docs (#1243)
* docs(kubernetes): add initial reference docs

Add a Kubernetes reference section under docs/reference/kubernetes/ covering
gateway setup, OIDC and reverse-proxy authentication, cert-manager PKI, and
Gateway API external access. Update installation.mdx to link to the new setup
page and nest the new folder under the Reference section in index.yml.

* docs(kubernetes): promote Kubernetes to top-level nav and retitle pages

Move docs/reference/kubernetes/ to docs/kubernetes/ so the section appears
as a top-level nav entry rather than nested under Reference. Retitle the
pages to broader, action-oriented names (Managing Certificates, Ingress,
Access Control) and rename the files to match. Update cross-references in
setup.mdx and the installation page.

* docs(nav): revert Reference to folder mode

The Reference section was switched to explicit section/contents to host
the nested Kubernetes subsection. With Kubernetes promoted to a top-level
folder, restore folder-mode auto-detection. Side effect: sandbox-compute-drivers
is picked up again, which was silently absent from the explicit list.

* docs(kubernetes): add OpenShift install page

Mirror the OpenShift install instructions from PR #1240 into the
Kubernetes docs section. Documents the SCC binding for sandbox pods
and the chart overrides required by OpenShift's Security Context
Constraints (disabled PKI init Job, plaintext TLS, cleared fsGroup
and runAsUser).
2026-05-07 12:55:56 -07:00
Drew Newberry 728165a12c docs: consolidate documentation structure (#1231) 2026-05-07 09:14:08 -07:00
Drew Newberry f56c09c7df docs: update gateway deployment architecture (#1108) 2026-05-04 22:52:29 -07:00
Miyoung Choi b7c763204d docs: refresh user-facing docs for recent sandbox and inference changes (#868)
* docs: refresh user-facing docs for recent sandbox and inference changes

- architecture: document system CA loading for upstream TLS, `tls: skip`
  as the opt-out, gateway state persistence across restarts, and OCSF
  structured logging surface.
- inference: document per-provider header allowlist, Authorization
  stripping, 120s streaming idle tolerance, and extended-thinking
  timeout guidance.
- manage-sandboxes: add "Execute a Command in a Sandbox" section for
  `openshell sandbox exec` with flag reference.
- security best practices: expand seccomp denylist (unconditional and
  conditional blocks), document two-phase Landlock probe, High-severity
  `landlock-unavailable` finding, and inference keep-alive closure.
- observability logging: document port in HTTP log URLs, `[reason:...]`
  denial suffixes, proxy 403/502 JSON error bodies, and Landlock
  CONFIG:ENABLED/CONFIG:OTHER events.

Signed-off-by: Miyoung Choi <miyoungc@nvidia.com>

* docs(architecture): tighten high-level architecture page

Trim implementation detail (CA bundle paths, deprecated TLS keys, SSH
handshake secret) from the high-level architecture page, fix accuracy
issues surfaced during deep audit, and expand uncommon acronyms on first
mention.

- Drop unsupported "cost-based routing" claim from Privacy Router row.
- Replace "brokers requests across the platform" with auth-boundary
  description.
- Add "inference" to Policy Engine constraint list per AGENTS.md.
- Expand Deny rule to include SSRF, blocked control-plane port, and L7
  deny paths in addition to deny-by-default.
- Switch Allow/Deny labels from hyphen to colon; remove em dashes and a
  double space.
- Expand LLM, SSRF, L7, TLS, CA, PEM, SSH, OCSF, and JSONL on first use.

Signed-off-by: Miyoung Choi <miyoungc@nvidia.com>
Made-with: Cursor

* docs: address audit feedback on refresh PR

- observability/logging: rewrite the allowed_ips paragraph after the
  Denial Reasons table; the previous wording said authors "can use"
  invalid entries while also stating they were rejected, which was
  contradictory and conflated load-time validation with the runtime
  per-CONNECT denial phrases the section documents.
- about/architecture: split compound sentences in the new Gateway
  Lifecycle and Observability sections so each clause stands alone.
- inference/about: drop the streaming-tolerance sentence from the prose
  paragraph since the dedicated Streaming reliability table row already
  covers it.

Signed-off-by: Miyoung Choi <miyoungc@nvidia.com>
Made-with: Cursor

* docs(architecture): address review feedback on architecture page

- Drop the Privacy Router row from the Components table; it is not
  yet a separately exposed customer-facing component.
- Update the page description and intro count to match the remaining
  three components (gateway, sandbox, policy engine).
- Split the policy decision into the three modes that the engine
  actually implements: Explicit Deny (deny rules and hardening rules,
  takes precedence), Allow, and Implicit Deny (no rule matched).

Signed-off-by: Miyoung Choi <miyoungc@nvidia.com>
Made-with: Cursor

* docs(architecture): convert policy decisions to a table

Promote the three policy decisions (Explicit Deny, Allow, Implicit
Deny) to a top-level table with Decision, When it applies, and
Outcome columns instead of a nested bulleted list under list item 5.
Top-level tables render reliably across markdown renderers, where
nested-in-list tables do not.

Signed-off-by: Miyoung Choi <miyoungc@nvidia.com>
Made-with: Cursor

* docs: repharse a bit

---------

Signed-off-by: Miyoung Choi <miyoungc@nvidia.com>
2026-04-17 13:15:59 -07:00
Piotr Mlocek 8b15ef772e docs(fern): move published docs into docs tree (#796)
Remove the legacy Sphinx pipeline and make docs/ the single source of truth so the published site matches the repository layout.
2026-04-09 15:02:27 -07:00
Parth Sareen ba19aade08 docs(ollama): update ollama tutorial and references to match latest (#511)
features
2026-03-20 14:42:54 -07:00
Hector Flores eff88b7011 feat(providers): add GitHub Copilot CLI agent provider (#476)
* feat(providers): add GitHub Copilot CLI agent provider

Source: https://docs.github.com/en/copilot/reference/copilot-allowlist-reference
2026-03-20 10:18:37 -07:00
Parth Sareen efb80e7382 docs(ollama): add ollama to community sandboxes catalog and supported agents (#383)
* docs(ollama): add ollama to community sandboxes catalog and supported agents

* update readme
2026-03-17 14:02:40 -07:00
Piotr Mlocek 85903b95e0 docs: add debug-inference skill, Ollama tutorial, and remove stale inference policy references (#353) 2026-03-16 07:26:25 -07:00
Miyoung ChoiandKirit93 c420109a6a docs: add dedicated gateway docs, network policy tutorial, and license page (#294)
* Added gateway and tutorial

* add new content and polish, add license terms

* small fix

* tiny fix

* polish bit more

* change warning to important

* fix support matrix

* small fix

* add a note about example script

* fix link

* fix link

* improve

* fix links

* unclutter sandbax index more

* improve titles

* tiny style fix

* style guide on inferences, release notes link to gh

* polish

* nit

* switch version to 0.0.3

* incorporate feedback

---------

Co-authored-by: Kirit93 <kthadaka@nvidia.com>
2026-03-14 09:53:01 -07:00