77 Commits
Author SHA1 Message Date
Oliver Calder 1ad4e428a6 fix(snap): simplify snap hooks (#3988)
* fix(snap): simplify snap hooks

The `post-refresh` hook runs after initial snap installation as well, so
there is no need to call the `install` hook from within the
`post-refresh` hook; instead, the logic can simply be moved into the
`post-refresh` hook directly, and the `install` hook removed.

Also, the existing `install` hook logic looked for an insecure
configuration, and if found, replaced the entire configuration file with
a minimal default in the current format. But OpenShell does that default
behavior without any config file, so we may as well simply remove the
configuration file entirely to keep up-to-date with the current default
behavior. Let OpenShell create a configuration file if it needs to,
rather than auto-create one via the packaging scripts.

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* fix(snap): remove the connect-plug-docker hook

The `openshell:docker` is auto-connected to the system `:docker` slot,
so there should not be a need to separately restart the gateway service
when the interface is connected.

For locally-built test snaps which were not published to the store, the
autoconnection is not made, but when the snap is installed, the gateway
will attempt to start anyway and fail to find any available compute
driver, so quickly restart until it hits the systemd start-limit, after
which systemd prevents the service from being started again. If a user
tries to manually connect their locally-built `openshell` snap to the
`:docker` slot, then the `connect-plug-docker` hook runs and triggers a
restart of the gateway, which will usually fail because the start limit
has already been hit. An error in the hook will thus cause the interface
connection to be undone, which is undesirable.

Thus, we can remove this hook entirely, and instead allow interface
connections to succeed as intended. The user still needs to manually
restart the gateway service after making a manual connection (as was the
case previously) and probably needs to `systemctl reset-failed` first,
but at least connection will succeed beforehand so they can proceed with
these steps.

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* fix(snap): set refresh-mode: endure again, with manual restart

Return to the previous behavior before commit a67567e58, where the
gateway is not stopped before refreshes. The `post-refresh` hook
now restarts the gateway if the TLS configuration was corrected, so we
don't have to enforce restarting the gateway on every refresh even when
not necessary. Thus, set `refresh-mode: endure`, and let the hook decide
when the gateway needs to be restarted.

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* fix(snap): update docs and tests to reflect snap hook changes

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* docs(snap): remove verbose explanation of snap gateway refresh behavior

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

---------

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>
2026-10-01 15:10:38 +00:00
Evan Lezar 0e8d9e55f9 test(e2e): remove schema parity campaign (#3864)
Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-10-01 13:17:20 +00:00
Drew Newberry 021400be8a refactor(auth): separate sandbox identity from TLS (#3110)
* refactor(auth): separate sandbox identity from TLS

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(auth): clarify gateway mTLS behavior

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(auth): include workspace scope in TLS authorization checks

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(e2e): bound service auth sandbox names for large PIDs

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-10-01 04:33:25 +00:00
Derek Carr 912a077bd6 feat(service): add bearer authorization passthrough (#3796)
* feat(service): add bearer authorization passthrough

Signed-off-by: Derek Carr <decarr@redhat.com>

* docs(sdk): add service authorization migration guide

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(server): remove stale version import

Signed-off-by: Derek Carr <decarr@redhat.com>

* docs(upgrade): remove service authorization SDK guide

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(e2e): relabel provider readiness TLS mount

Signed-off-by: Derek Carr <decarr@redhat.com>

* test(e2e): stabilize exposed service routing

Signed-off-by: Derek Carr <decarr@redhat.com>

* test(e2e): support HTTPS service routing

Signed-off-by: Derek Carr <decarr@redhat.com>

---------

Signed-off-by: Derek Carr <decarr@redhat.com>
2026-09-30 20:22:46 +00:00
Drew Newberry acbac9cb79 feat(sandbox): add main restart policy (#2798)
* feat(sandbox): add main restart policy

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): harden policy-driven restarts

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): address restart review feedback

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>

* fix(sandbox): port restart policy to current runtime

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* perf(sandbox): restart promptly after terminal delivery

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>
2026-09-29 02:18:30 +00:00
Drew Newberry d1a19c70ee fix(homebrew): overwrite generated gateway config during migration (#3752)
Closes #3746

Use Homebrew atomic_write for exact legacy config migrations and keep the first-install write distinct in the formula test.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-27 18:56:08 +00:00
Drew Newberry a67567e583 fix(snap): require mTLS for the snap gateway (#3726)
* fix(snap): require mTLS for the snap gateway

Replace the installer opt-in with an authenticated snap gateway. The wrapper
no longer forces plaintext, so the gateway serves TLS from the bundle it
already generates in $SNAP_COMMON/tls. The install hook writes a config that
enables mTLS user auth instead of unauthenticated access, and a new
post-refresh hook migrates the exact legacy default on existing installs.

install.sh waits for the gateway, detects whether it serves TLS, copies the
client bundle into the target user's snap state directory, and registers the
gateway over HTTPS. Older plaintext snap revisions still register over HTTP
with a warning. The release canary asserts mTLS auth and HTTPS registration.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(snap): pass config preflight and detect the mTLS gateway reliably

An explicit [openshell.gateway.mtls_auth] table fails config preflight,
which validates mTLS auth before the local TLS bundle supplies the client
CA. Write a default that pins the Docker driver instead; with the wrapper's
TLS bundle the gateway requires client certificates and enables mTLS user
auth automatically, as the native packages do.

The mTLS gateway rejects TLS handshakes without a client certificate, and
it still answers plaintext loopback HTTP for sandbox service routing, so the
installer could misdetect it as a legacy plaintext gateway. Probe HTTPS with
the root-owned client bundle, and treat a gateway as legacy only when a
plaintext gRPC Health call succeeds.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* chore(snap): simplify install hook comment

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(snap): migrate insecure gateway configs on refresh

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(snap): simplify mTLS detection and config migration

Detect the mTLS snap from the installed revision's post-refresh hook instead
of probing plaintext gRPC, and drop the scheme global. Remove the installer's
pre-hook config fallback, which is dead now that every channel ships the
install hook and which wrote the insecure default. Give the install hook a
single write path with a simple backup name, and shorten the manual client
certificate steps.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(snap): stop keeping a copy of replaced insecure configs

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(install): make the snap an opt-in install method

Stop selecting the OpenShell snap just because the snap command exists. Linux
installs default to the Debian or RPM package; OPENSHELL_INSTALL_METHOD=snap
(or deb, rpm) selects the package explicitly. Hosts that already have the
OpenShell snap keep refreshing it rather than gaining a second gateway on the
same port. The release canary and snap repro script opt in explicitly.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(snap): let the gateway auto-detect its compute driver

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(snap): restart the gateway after refresh

Published revisions use refresh-mode: endure, and snapd honors the old
revision's setting during a refresh, so the plaintext gateway kept running
with the migrated config unused until a manual restart. Restart the gateway
from the post-refresh hook so the mTLS config takes effect immediately.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(snap): drop refresh notes from the snap description

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(snap): trim snap refresh notes from installation docs

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-26 01:50:02 +00:00
Gaizka Menendez 123d95e2ed fix(pagination): document list contract and harden SDK pagers (#3279)
* fix(pagination): document list contract and harden SDK pagers

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* fix(pagination): bound pager token history

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* fix(pagination): preflight token history limits

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* fix(pagination): paginate sandbox providers

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* fix(pagination): harden TypeScript pager

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* fix(pagination): expose provider pagers in SDKs

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* docs(pagination): describe a uniform list contract

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* fix(pagination): use stable provider cursors

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* docs(pagination): clarify mutation semantics

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* fix(cli): expose sandbox provider pagination

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* fix(proto): refresh pagination schema inventory

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* fix(proto): refresh rebased schema inventory

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>
Signed-off-by: Kris Hicks <khicks@nvidia.com>

---------

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>
Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-09-23 18:51:59 +00:00
50230616d5 refactor(runtime): retire Community image dependencies (#3386)
* feat(sandbox): default to official Alpine sandbox image

default_sandbox_image() now returns docker.io/library/alpine:3.22, a generic
version-qualified official image, so a fresh install no longer depends on the
community sandbox image catalog. All compute drivers (docker, podman,
kubernetes, vm) inherit this fallback.

Part of #3116.

Signed-off-by: Akram
Signed-off-by: Akram <akram.benaissi@gmail.com>

* feat(deploy): default deployment configs to the official Alpine sandbox image

Update the shared gateway default_image, Helm chart values, the standalone
Kubernetes manifest, and the dev gateway task scripts to use
docker.io/library/alpine:3.22 instead of the community base image, consistent
with default_sandbox_image(). GPU e2e image-build base is left unchanged (CUDA
needs a glibc base).

Part of #3116.

Signed-off-by: Akram
Signed-off-by: Akram <akram.benaissi@gmail.com>

* feat(driver): default to numeric non-root identity for USER-less images

With the default sandbox image now Alpine, images that declare no OCI USER
must start instead of being rejected. When the image declares no USER and
the policy requests none, the Podman and Docker drivers now supply a numeric
non-root identity (DEFAULT_SANDBOX_UID/GID = 1000) instead of rejecting,
matching the numeric-identity behavior of the Kubernetes and VM drivers. The
supervisor's resolved-identity path runs the sandbox as a synthesized
non-root account without the account existing in the image. Images that
declare a USER keep the OCI resolution path unchanged.

Part of #3116.

Signed-off-by: Akram <akram.benaissi@gmail.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(conformance): use Alpine workload image

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* refactor(policy): drop community image /app path from default policy

The restrictive default policy granted read-only access to /app, a directory
that only existed in the community base image. A generic Alpine default has no
/app, so remove it. Landlock best-effort already ignores absent paths; this
just stops advertising a community-specific layout in the default.

Part of #3116.

Signed-off-by: Akram
Signed-off-by: Akram <akram.benaissi@gmail.com>

* docs(config): document Alpine default images

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(podman): report early sandbox termination

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(podman): initialize rootless workspace ownership

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(sandbox): qualify NVIDIA Ubuntu default

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): initialize rootful default workspace

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(sftp): add native sandbox adapter

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sftp): gate runtime helper support to Linux

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sftp): support standard OpenSSH file operations

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sftp): harden rename and special file handling

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(runtime): remove community image dependencies

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(e2e): build provider readiness tool fixture

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(e2e): use a dedicated Noble fixture for Docker tests

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Akram
Signed-off-by: Akram <akram.benaissi@gmail.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Evan Lezar <elezar@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
2026-09-22 14:43:51 +02:00
Drew Newberry 1905069948 feat(sandbox): expose services during creation (#3439)
* feat(sandbox): expose services during creation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(providers): refresh Codex credentials in gateway

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(cli): normalize create-time service URLs

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(server): roll back failed service exposure

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(example): simplify Codex provider setup

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(example): separate provider setup commands

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): harden create-time service exposure

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(example): bundle Codex provider profile

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(example): allow npm-installed Codex binary

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-20 23:19:25 -07:00
Drew Newberry 17ce738bfb fix(ci)!: remove gateway callback listener dependency (#3365)
* fix(ci): repair post-merge release canary

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(packaging): bootstrap canary runtime prerequisites

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci(canary): collect macOS VM diagnostics

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci(canary): pin libkrun-compatible macOS runner

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci(canary): limit macOS smoke test to package startup

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(compute)!: remove gateway callback listeners

Run Docker supervisors on host networking so they use the operator-configured primary gateway endpoint. Remove the unused compute-driver callback listener negotiation and listener-scoped routing machinery.

BREAKING CHANGE: The ComputeDriver API no longer exposes GetGatewayListenerRequirements or GatewayListenerRequirement. External drivers must regenerate bindings and connect supervisors to the configured primary gateway endpoint.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(docker): use sandbox runtime image in launcher

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(podman): exercise production endpoint selection

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(e2e): route supervisors to reachable gateways

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(e2e): align Podman endpoint fixtures

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): preserve host aliases for supervisors

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): align sandbox host gateway pin

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(e2e): address Docker fixtures by bridge IP

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(e2e): serialize sandbox lifecycle cases

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(e2e): host Docker TCP fixture with gateway

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(e2e): use loopback for host-network supervisor

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-18 20:55:55 +00:00
Drew Newberry d91b1999a0 feat(api)!: use sandbox names as canonical RPC references (#3272)
* feat(api)!: use sandbox names as canonical references

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(cli): update forward color fixture for workspace scope

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(supervisor): use sandbox names for settings lookup

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(e2e): use canonical sandbox request fields

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(e2e): use canonical sandbox receipt field

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(server): harden sandbox mutation handling

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(ci): update rebased sandbox references

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(api)!: standardize canonical entity references

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(api): codify protobuf API conventions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(api): preserve workspace selector semantics

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(api): restore workspace selector parity

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(api): preserve descriptive name fields

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(api): update e2e request fixtures

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(core): omit workspace selector during bootstrap

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* chore(api): remove proto convention checker

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* chore(api): refresh schema fingerprints after rebase

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(cli): use canonical provider receipt field

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-18 11:39:55 -07:00
Drew NewberryandPiotr Mlocek 8b77925eb7 feat(installer): support prerelease installations (#3364)
* feat(installer): support prerelease installations

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(readme): clarify 0.1.0 production readiness

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(readme): refine production readiness message

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(release): publish rolling prerelease channels

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(release): keep prereleases out of GitHub releases

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(readme): refine 0.1.0 notice

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(readme): expand 0.1.0 notice

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(readme): simplify 0.1.0 notice

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(installer): simplify prerelease alias

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(installer): require successful prerelease runs

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(release): reuse exact prerelease artifacts

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(installer): include prover in prerelease bundle

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci: remove prerelease post-publish validation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(fern): sync announcement configuration

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs: remove stale prerelease canary claim

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(fern): defer announcement synchronization

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(fern): simplify release announcement

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Co-authored-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-09-18 00:53:36 +00:00
Johnny Greco 58b5f8f976 feat(prover): add standalone policy boundary checker (#3289)
* feat(prover): add standalone policy maximum checker

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* refactor(prover): simplify check scope schema

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(prover): align containment and cancellation with runtime

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(prover): stabilize containment checks in CI

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* test(prover): avoid solver in fast-path guard test

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(prover): align string containment with runtime

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(prover): reject ambiguous z3 string escapes

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(prover): align containment with runtime boundaries

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(prover-cli): harden cancellation and invalid input

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* feat(packaging): install policy prover with OpenShell

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(prover): clarify installation and check results

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(build): describe prover distribution directly

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* refactor(prover): rename maximum policy to boundary

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* refactor(prover): localize fail-closed validation

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* test(prover): cover fail-closed CLI surfaces

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* test(prover): allow CI load for REST solver proof

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(prover): use canonical policy schema for containment

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(prover): preserve uncertainty for runtime binary globs

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(prover): bound policy validation work

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(prover): document validation resource limits

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(prover): make containment API extensible

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(prover): define containment API contract

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(ci): integrate prover with consolidated builds

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(ci): declare release packaging dependency

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

---------

Signed-off-by: Johnny Greco <jogreco@nvidia.com>
2026-09-17 17:30:03 +00:00
Mrunal Patel c502be9fd7 feat(api): return typed deletion outcomes with explicit missing-target semantics (#3317)
* feat(api): return typed deletion outcomes with explicit missing-target semantics

Implement phase 2 of #3051 across the public gateway API and first-party SDKs. Preserve asynchronous sandbox deletion and observed resource identities, reserve legacy wire fields, and document the coordinated migration.

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(sdk): bind deletion waits to sandbox identity

Track accepted deletions by original sandbox identity in Rust and TypeScript. Preserve name-only waits and cover replacement races and lookup failures.

Merge the phase-one cleanup fix and adapt its interceptor regressions to typed deletion outcomes.

Refs #3051.

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* test(e2e): accept asynchronous stopped sandbox deletion

Allow either completed or accepted deletion output, then continue polling for actual sandbox absence. This preserves the lifecycle assertion under the phase-two deletion contract.

Refs #3051.

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

---------

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
2026-09-16 22:11:19 +00:00
Derek Carr d68b7069c3 refactor(proto)!: use well-known time types (#3113)
* refactor(proto)!: use well-known time types

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(proto): preserve time migration behavior

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(proto): preserve timestamp boundary semantics

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(proto): convert sandbox token expiry to timestamp

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(proto): preserve time compatibility semantics

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(proto): preserve exact endpoint and profile times

Signed-off-by: Derek Carr <decarr@redhat.com>

* test(e2e): use duration for interactive exec timeout

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(sdk-go)!: remove legacy profile duration fields

BREAKING CHANGE: Go provider profile callers must use RefreshBefore, MaxLifetime, and CacheTTL with ProfileDuration instead of the whole-second fields.

Signed-off-by: Derek Carr <decarr@redhat.com>

---------

Signed-off-by: Derek Carr <decarr@redhat.com>
2026-09-16 19:51:45 +00:00
Drew Newberry c1f2e7189f feat(isolation): implement the RFC 0012 sandbox architecture (#2942)
* feat(isolation): add RFC 0012 backend contract

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>

* refactor(isolation): name the interface crate explicitly

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): expose trusted host gateway

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(agents): inventory the MXC driver

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): add mediated DNS transport

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): tighten interface error and digest contracts

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(isolation): remove unrelated driver inventory

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): define capability-free launch contract

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): seal confirmed boundary state

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): validate confirmation for external backend implementations

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): clarify mediated DNS identity

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(isolation): generalize loopback connector

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(isolation): unify typed network mediation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): bind launches to sandbox sessions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(mxc): initialize extended sandbox status

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): add boundary protocol and Linux primitives

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): harden signals and separate process status from transport

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): validate remote confirmation through public contract

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): validate wire state and propagate snapshot failures

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(isolation): import owned agent specification explicitly

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(isolation): describe mediated DNS channel

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): bound mediation attach without nested retries

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(isolation): generalize loopback protocol

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): add transport-neutral session authentication

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(isolation): separate sandbox backend protocol

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): harden runtime boundary controls

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): add terminal boundary operation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): split supervisor and sandbox runtimes

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): harden boundary isolation and lifecycle ownership

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): reject private root redirects and adopt typed errors

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): preserve accept thread ownership on musl

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(sandbox): isolate credential probes from filtered threads

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): return retained exec exit status to independent waiters

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): bound network mediation and preserve socket authorization

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): bound control admission and retire stale mediation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci(e2e): select migrated drivers per stack layer

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(sandbox): implement loopback connector

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): authenticate the Sandbox Protocol

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(supervisor): rotate launch-scoped authentication

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(sandbox): consume dedicated backend crate

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(sandbox): align topology session fixture

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): align projected bootstrap bundle

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(auth): validate refreshed credentials before rotation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): fail closed across supervisor disconnects

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): repair rebased sandbox CI

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* build(runtime): publish separate sandbox and supervisor images

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(config): configure the sandbox runtime image

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(ci): validate sandbox binary linkage

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(isolation): use backend and runtime terminology

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(sandbox): use a scratch runtime image

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(ci): refresh schema and dependency policy

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): bind reconnects to supervisor process

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs: align runtime split operational guidance

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* chore(security): document Kubernetes runtime RBAC

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): enforce runtime lifecycle invariants

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(compute): identify sandbox start generations

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(server): restore sandbox launch sessions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): support authenticated runtime replacement

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(auth): bind sandbox session successors

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(auth): retry pending sandbox successors

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(vm): run the supervisor outside the guest workload

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(vm): use sandbox backend protocol

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): repair rebase integration

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): use unified build toolchain

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(vm): use sandbox runtime terminology

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(vm): own guest network bootstrap

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): expose guest init version

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): select native supervisor artifacts

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): guard guest init Linux symbols

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): scope Linux test imports

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): avoid guest interface casts

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): reconcile admitted sandbox identity

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): share resolved sandbox identity

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): surface host supervisor failures

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): include guest logs on supervisor exit

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): rotate and clean runtime generations

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): make sandbox starts generation-aware

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): rotate restored sandbox sessions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(vm): keep shared paths in the base layer

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): bind sandbox session lineage

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(docker): isolate workloads behind the host supervisor

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(docker): rotate launch-scoped authentication

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(docker): use sandbox backend protocol

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(docker): use host networking for supervisor

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): preserve host gateway alias resolution

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(docker): use separate sandbox and supervisor images

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): restore startup validation after rebase

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(docker): name the sandbox runtime directly

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): narrow supervisor CA runtime storage

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): close companion isolation gaps

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(docker): align mediated network expectations

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(docker): exercise mediated network paths

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): attach supervisor to managed network

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): defer supervisor recovery until gateway is ready

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): make sandbox starts generation-aware

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): rotate restored sandbox sessions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): preserve workloads during session rotation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(docker): remove unrelated configuration RFC changes

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): bind sandbox session lineage

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(kubernetes): add proxy-pod isolation topology

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(kubernetes): use sandbox backend protocol

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): use stable sandbox service authority

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(kubernetes): split sandbox and supervisor images

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): adapt proxy pods to current runtime APIs

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(kubernetes): describe the single runtime placement

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(kubernetes): simplify sandbox orchestration

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): validate deployment prerequisites

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): update Trivy Helm profile inventory

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(kubernetes): update Trivy scan inventory count

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): reuse preloaded runtime images in e2e

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): type and clean runtime resources

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): make sandbox restarts recoverable

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): rotate restored sandbox sessions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): preserve supervisor egress

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): bind sandbox session lineage

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(podman): adopt isolated sandbox and supervisor containers

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): stage bootstrap archives at named volume destinations

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(podman): rotate launch-scoped authentication

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(podman): use sandbox backend protocol

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(podman): use host networking for supervisor

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(podman): split sandbox and supervisor images

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): repair rebase integration

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(podman): name the sandbox runtime directly

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): provision supervisor CA runtime storage

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): address isolation review findings

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): inspect Debian supervisor provenance

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): use libpod-compatible tmpfs options

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): bind verified sandbox runtime binary

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): provide external driver data directory

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): start sandbox before joining user namespace

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): separate supervisor user namespace

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): make sandbox starts generation-aware

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): rotate restored sandbox sessions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): bind sandbox session lineage

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* perf(isolation): add TCP and DNS benchmark harnesses

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(perf): align benchmark timing and supported protocols

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(perf): report TCP benchmark metrics accurately

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(perf): cancel failed worker startup

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): build matching local supervisor image

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): make local sandbox smoke test runnable

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): wire local sandbox runtime image

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): narrow sandbox service RBAC

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(ci): validate split runtime artifacts

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): harden runtime session handling

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(supervisor): add standalone network proxy role

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(rfc): remove implementation companion notes

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(vm): standardize runtime release name

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): pin renamed runtime artifacts

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(auth): persist sandbox runtime identity

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(runtime): restore branch validation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): reconcile main after rebase

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(network): close unframed HTTP 1.0 responses

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* chore(isolation): preserve upstream OCSF updates

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(security): close credential and TLS replay paths

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(auth): make sandbox refresh retries idempotent

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-16 00:49:14 +00:00
Mrunal Patel 39cf4823f7 feat(api): add structured gateway errors and SDK decoding (#3313)
* feat(api): expose structured gateway errors across SDKs

Refs #3051. Add standard validation, conflict, and retry details; preserve raw transport status in Rust, Go, TypeScript, and Python; document status and recovery guidance.

This is the structured-error foundation only. Mutation result shapes, allow_missing, durable request deduplication, and exec retry semantics remain follow-up work.

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(python): preserve wrapped RPC cleanup handling

Inspect the original gRPC call when handling missing sandboxes during deletion waits and managed cleanup. Add intercepted cleanup regressions and clarify the error-wrapper migration contract.

Addresses the cleanup review on #3313; part of #3051.

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

---------

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
2026-09-15 17:51:11 +00:00
Jesse JaggarsandDrew Newberry 02b664bb0d refactor(config): normalize and enforce gateway schema v2 (#2814)
* refactor(config): normalize compute driver field names

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* refactor(config): introduce canonical gateway fields

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* refactor(config): enforce gateway schema version 2

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): preserve compute driver runtime guarantees

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): address schema v2 review regressions

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): complete schema v2 migration safeguards

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(config): expand schema v2 regression coverage

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(config): add schema v2 parity manifest

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): correct parity manifest inventory

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* docs(config): record schema v2 intentional changes

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* docs(config): disposition schema v2 parity gaps

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): add dual schema parity harness

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): establish compute lifecycle parity baseline

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): preserve gateway option compatibility

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): record gateway option parity

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* docs(config): close gateway-wide parity gaps

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(podman): apply configured pids limit

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): validate Podman option parity

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): add Kubernetes option parity harness

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): record Kubernetes option parity

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): disposition VM parity lanes

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): add external driver parity lane

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(e2e): preserve external driver pull policy

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): attest parity artifacts and launches

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): require clean parity build sources

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): bind parity runtime artifacts

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(e2e): use isolated supervisor tags

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(e2e): qualify parity image tags

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(e2e): serve parity supervisor locally

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): isolate parity podman services

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): harden parity evidence provenance

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): pin parity sandbox artifacts

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): attest parity runtime inputs

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): bind parity runtime evidence

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): record compute boundary parity

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): disposition cross-cutting parity lanes

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(packaging): preflight gateway config upgrades

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): preserve rebase integration guarantees

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(ci): isolate temporary git signing config

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): update remaining schema v2 consumers

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(ci): provide e2fs tools to VM tests

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): align preflight with gateway startup

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(vm): preserve rootfs tar configuration

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* chore(config): adopt duration unit constructors

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(packaging): preflight RPM gateway config

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): address driver review findings

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(e2e): require fresh semantic parity evidence

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(docker): update tests for renamed sandbox label

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(gateway): preserve selective driver coverage after rebase

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
2026-09-11 05:00:24 +00:00
Drew Newberry 1860010850 feat(sdk): add lazy pagination pagers (#3256)
* feat(sdk): add lazy pagination pagers

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(sdk): harden pager edge cases

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(sdk): cover initial resume token

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(go): fix all-workspaces pager examples

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-11 00:02:08 +00:00
Drew Newberry 33bbda3d33 refactor(persistence): adopt continuation-token pagination (#3249)
* refactor(persistence): adopt continuation-token pagination

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(pagination): address continuation review findings

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(tui): recover completed list refreshes

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(pagination): address review scalability findings

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(pagination): repair branch validation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(go): use page size in template example

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-11 00:02:07 +00:00
Mrunal PatelandJohn Myers 90dbe5454b feat(api): add typed workspace selectors (#3245)
* feat(api)!: add typed workspace selectors

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(cli): preserve template workspace metadata

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* test(e2e): migrate workspace request selectors

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(api): update public schema inventory

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-09-10 20:38:50 +00:00
John T. Myers f4dc6be4b2 refactor(inference): remove managed inference routes (#3195)
* refactor(inference): remove managed inference routes

Closes #3172

Remove the inference route control plane, inference.local data path, built-in router crate, and SDK surface. Move inference workloads to explicitly imported provider profiles and native endpoints, with migration cleanup and updated tests and documentation.

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(policy): preserve alternate upstream isolation

Restore the provider policy activation guard so legacy OpenAI and Anthropic providers configured for alternate base URLs do not grant egress to the built-in public vendor endpoints.

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-09-09 18:47:22 +00:00
Evan Lezar 8d7db25403 fix(snap): recover gateway after Docker connection (#2866)
Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-09-03 08:21:43 +00:00
grs 74960ebfae feat(server): add sandbox templates (#2833)
* feat(server): add sandbox workload templates

Signed-off-by: Gordon Sim <gsim@redhat.com>

* feat(go-sdk): add sandbox workload template support

Signed-off-by: Gordon Sim <gsim@redhat.com>

* feat(rust-sdk): add sandbox workload template support

Signed-off-by: Gordon Sim <gsim@redhat.com>

* feat(python-sdk): add sandbox workload template support

Signed-off-by: Gordon Sim <gsim@redhat.com>

* feat(typescript-sdk): add sandbox workload template support

Signed-off-by: Gordon Sim <gsim@redhat.com>

* docs(agents): document sandbox workload templates

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(cli): support default GPU requests in sandbox templates

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(server): cap sandbox templates per workspace

Signed-off-by: Gordon Sim <gsim@redhat.com>

* docs(architecture): document sandbox workload template boundaries

Signed-off-by: Gordon Sim <gsim@redhat.com>

* feat(cli+sdk): expose sandbox workload template provenance

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(cli): include sandbox template annotations in output

Signed-off-by: Gordon Sim <gsim@redhat.com>

* feat(sandbox): add label selectors to template listing

Signed-off-by: Gordon Sim <gsim@redhat.com>

* test(e2e): cover sandbox template failure paths

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(server): preserve command and ttl when creating sandbox from template

Signed-off-by: Gordon Sim <gsim@redhat.com>

* test(server): add field coverage test for template merge

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(go-sdk): add pagination support to fake client

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(cli): warn on env vars that looks like secrets

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(sdk-ts): propagate sandbox workspace through lifecycle calls

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(docs): update workspace management docs

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(ts-sdk): support command and tty when creating from template

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(go-sdk): support command and tty when creating from template

Signed-off-by: Gordon Sim <gsim@redhat.com>

* test(python-sdk): verify command and tty handling when creating from template

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(go-sdk): update docs and ClientInterface

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(server): validate sandbox create specs before I/O

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(go-sdk): guard empty DNS-1123 label validation

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(python-sdk): allow empty template builder mappings

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(cli): align template GPU JSON default output

Signed-off-by: Gordon Sim <gsim@redhat.com>

---------

Signed-off-by: Gordon Sim <gsim@redhat.com>
2026-09-02 00:55:29 +00:00
Simon Scatton a4f9c762ce fix(release): handle prerelease tag builds (#3094)
Signed-off-by: Simon Scatton <sscatton@nvidia.com>
2026-09-01 16:12:01 +00:00
Simon Scatton c8f13205e3 ci(release): publish prerelease artifacts (#3093)
Signed-off-by: Simon Scatton <sscatton@nvidia.com>
2026-09-01 15:31:18 +00:00
Artem Lytvyn eb15e1a4c9 feat(sandbox): add --no-login-shell to skip shell startup files on exec (#2852)
* feat(sandbox): add --no-login-shell to skip shell startup files on exec

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* chore(sdk/go): regenerate proto bindings for no_login_shell

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* test(supervisor-process): cover login-shell flag selection

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* feat(sandbox): gate --no-login-shell on supervisor SSH banner

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

---------

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>
2026-08-31 16:17:58 +00:00
Drew Newberry 69a05ebb3b fix(sandbox): complete successful main processes (#2884) 2026-08-28 18:21:20 -07:00
Seth Jennings 5206bc51b2 feat(sdk): add OAuth Client Credentials support to SDKs (#2907)
* feat(sdk): add renewable client credentials auth

Implement lazy OAuth client-credentials acquisition and renewal for the Python, TypeScript, and Go SDK clients, with shared security conformance coverage and service-account documentation.\n\nCloses #2803

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(sdk): address client credentials review

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(sdk): address client credentials review

Signed-off-by: Seth Jennings <sjenning@redhat.com>

---------

Signed-off-by: Seth Jennings <sjenning@redhat.com>
2026-08-25 17:53:47 +00:00
Simon Scatton 455883905a fix(python): remove CLI from wheel (#2321)
The Maturin-based wheel packaging was a historical remnant from when the local gateway launch path and OpenShell CLI were coupled in one binary. The gateway and CLI now ship as standalone artifacts, so the Python distribution should contain only the SDK.

Build a single platform-independent setuptools wheel, verify that it cannot contain native code or an openshell entry point, and simplify the release jobs and documentation for SDK-only PyPI installs.

Signed-off-by: Simon Scatton <sscatton@nvidia.com>
2026-08-25 13:20:42 +00:00
Drew Newberry ef296806f5 feat(sandbox): add canonical main process (#2726)
* feat(sandbox): add canonical main process

Closes #2710

Persist and supervise one canonical workload per sandbox, attach sandbox connect to its retained session, and make every unexpected main-process exit terminal.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(sandbox): simplify canonical main process contract

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): preserve legacy VM main compatibility

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): preserve main status across driver updates

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): satisfy macOS process lint

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): gate Linux exit acknowledgement publisher

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(sandbox): simplify main process plumbing

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(supervisor): make controlling tty ioctl portable

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(supervisor): initialize canonical process environment

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* perf(supervisor): optimize retained main session

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(sandbox): detach main session on ctrl-c

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): use explicit main detach keys

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(dev): atomically stage Docker supervisor

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(test): align Docker main environment assertion

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(sdk): expose canonical main process fields

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-08-20 22:18:23 +00:00
Drew Newberry f12f3ef8d5 fix(macos): restore Homebrew sandbox callbacks (#2739)
* fix(macos): restore Docker gateway callbacks

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(gateway): reuse reachable primary callback listener

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-08-14 18:38:20 +00:00
Seth Jennings 0f8fad23c4 feat(sandbox): add stop and start operations (#2653)
* feat(sandbox): add suspend and resume operations

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(server): preserve lifecycle work after cancellation

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(server): reconcile ambiguous lifecycle outcomes

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(server): complete suspended session cleanup

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(vm): preserve suspension state on resume failure

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(server): retry retained lifecycle transitions

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(server): clean sessions after suspend reconciliation

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* test(sandbox): cover deleting suspended sandbox

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(kubernetes): preserve progressing sandbox suspension

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(kubernetes): bound suspend status polling

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(kubernetes): detect legacy sandbox suspension

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(tui): render suspended sandbox phases

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* refactor(sandbox): rename suspend and resume lifecycle

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* perf(server): clean stopped sessions on transition

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(kubernetes): fail fast on rejected stop

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(compute): fence stale restart lifecycle events

Signed-off-by: Seth Jennings <sjenning@redhat.com>

---------

Signed-off-by: Seth Jennings <sjenning@redhat.com>
2026-08-13 06:13:54 +00:00
2f96c53b8c feat(gateway,cli): windows compilation support (#2496)
* chore(windows): gate Unix-only workspace code for MSVC

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* feat(windows): stub unsupported compute drivers

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* ci(windows): add MSVC mise build lane

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* docs(windows): document MSVC build-only design

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* docs(agent): add Windows MSVC build skill

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* feat(windows): add Windows build support

Signed-off-by: Akber Raza <akberr@nvidia.com>

* refactor(windows): consolidate Windows-specific dependencies and improve build logic

Signed-off-by: Akber Raza <akberr@nvidia.com>

* feat(windows): add libclang path resolution and update cargo commands with bundled Z3 features

Signed-off-by: Akber Raza <akberr@nvidia.com>

* chore(tooling): lock Windows tool artifacts

Signed-off-by: Giedrius Burachas <gburachas@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* feat(windows): enhance libclang path resolution to support architecture-specific subdirectories

Signed-off-by: Akber Raza <akberr@nvidia.com>

* Fix Windows dependency gating after sync merge

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(z3): update Z3 header path requirements in Windows build documentation and scripts

Signed-off-by: Akber Raza <akberr@nvidia.com>

* docs(windows): relocate Windows MSVC build design to architecture/

Why: windows-msvc-build-design.mdx is a design document ("design decisions for
the native Windows MSVC build lane"), but it lived in the published, user-facing
docs/reference/ tree. Per AGENTS.md (Documentation) and architecture/README.md
("rfc/ vs architecture/"), design content belongs in architecture/ (or rfc/),
not in published reference. It also shared Fern sidebar "position: 6" with the
MXC compute-driver design page, colliding in the Reference nav ordering.

What:
- Move docs/reference/windows-msvc-build-design.mdx ->
  architecture/windows-msvc-build.md.
- Strip the Fern publish frontmatter and add a plain H1, matching the other
  architecture docs.
- Register it in the architecture doc index in architecture/README.md.
- Repoint the inbound references (build-openshell-mxc-windows skill + reference,
  implement-openshell-mxc-driver skill) to the new path.

With both design pages moved out of docs/reference/, the duplicate position-6
sidebar collision is resolved.

Signed-off-by: Akber Raza <akberr@nvidia.com>

* remove openshell-supervisor-network from unsupported driver package test exclusion list

Signed-off-by: Akber Raza <akberr@nvidia.com>

# Conflicts:
#	tasks/scripts/windows-msvc.ps1

* fix(interceptors): gate unix-only imports so the crate builds on Windows

openshell-gateway-interceptors failed to compile on Windows (E0432: no UnixStream in tokio::net), breaking any Windows build of openshell-server (which depends on it unconditionally). The connect_unix_endpoint fn was already #[cfg(unix)]-gated, but the imports it uses (UnixStream, TokioIo, Uri, service_fn) were left ungated. Gate those four imports with #[cfg(unix)] too. No behavior change on unix; Windows now compiles (no errors, no unused-import warnings).

Signed-off-by: Akber Raza <akberr@nvidia.com>

* feat(windows): add native ARM64 test support

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mise): skip Skaffold on Windows

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(windows): harden ARM64 toolchain discovery

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(windows): scope ARM64 toolchain preflight

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(windows): restore compatibility after GitHub sync

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(windows): avoid rate-limited Z3 source lookup

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mise): skip Helm checks on Windows

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(windows): support repository pre-commit checks

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(windows): stabilize native MSVC validation

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(windows): harden shared Z3 source cache

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>

* fix(windows): avoid leaking MSVC flags into clang-cl

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(windows): complete ARM64 migration audit

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(windows): restore ARM64 Ninja discovery

Signed-off-by: Akber Raza <akberr@nvidia.com>

* refactor(windows): separate platform crate roots

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(windows): restore proto include cfg gating

Signed-off-by: Akber Raza <akberr@nvidia.com>

* refactor: address lint errors

* fix(windows): add preflight check for proxy auth file path

* docs(windows): update GitHub checkout guidance

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(windows): restore CI after dependency updates

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mise): repair Windows sccache lock entry

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(windows): reconcile validation after rebase

Signed-off-by: Akber Raza <akberr@nvidia.com>

* refactor(server): exclude unsupported drivers on Windows

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* refactor(server): isolate platform driver config

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(windows): repair unsupported driver contract test

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(sandbox): remove stale dependencies

Signed-off-by: Akber Raza <akberr@nvidia.com>

* ci(windows): pin x64 workflow actions

Signed-off-by: Akber Raza <akberr@nvidia.com>

* ci(windows): align x64 Rust toolchain

Signed-off-by: Akber Raza <akberr@nvidia.com>

* ci(windows): align ARM64 workflow setup

Signed-off-by: Akber Raza <akberr@nvidia.com>

* refactor(windows): exclude unsupported runtime crates

Signed-off-by: Akber Raza <akberr@nvidia.com>

* refactor(windows): exclude unsupported crates at workspace boundary

Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com>

* refactor(server): gate builtin driver config by platform

Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com>

* fix(sandbox): restore crate documentation

Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com>

* ci(windows): make build workflow manual

Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com>

* ci(windows): temporarily enable pull request builds

Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com>

* ci(windows): cache Rust dependencies

Signed-off-by: Akber Raza <akberr@nvidia.com>

* refactor(windows): remove unnecessary platform changes

Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com>

* ci(windows): make build workflow manual

Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com>

* fix(ci): synchronize mise lockfile

Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com>

* fix(ci): normalize mise provenance metadata

Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com>

* refactor(python): isolate Windows atomic replace retry

Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com>

* fix(python): type Windows permission test errors

Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com>

---------

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>
Signed-off-by: Giedrius Burachas <gburachas@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com>
Co-authored-by: Shailendra Singh <shailendras@nvidia.com>
Co-authored-by: Giedrius Burachas <gburachas@nvidia.com>
Co-authored-by: Jamie King <jamiek@nvidia.com>
Co-authored-by: Piotr Mlocek <pmlocek@nvidia.com>
Co-authored-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com>
2026-08-11 21:00:36 +00:00
Evan LezarandDrew Newberry d220d89468 feat(compute): negotiate gateway callback listeners (#2492)
* feat(compute): query gateway listener requirements

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* feat(compute): add Podman listener requirements

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(docker): use default gateway bind address

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(gateway): avoid wildcard primary listener

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(podman): validate callback listener discovery

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(server): support split dual-stack listeners

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(podman): support legacy rootless listener discovery

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(e2e): accept loopback plaintext rejection

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* docs(agent): add callback listener diagnostics

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(server): restrict compute callback listeners

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(podman): validate local callback port

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(server): clarify callback listener contract

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(podman): require pasta for local callbacks

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* docs(gateway): document RPM listener default

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* refactor(server): keep listener provenance diagnostic-only

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(compute): preserve callback listener isolation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(e2e): remove Podman callback relay

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(packaging): preserve Podman callback loopback

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci(e2e): run VM smoke on nested-virt runner

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci(e2e): gate VM smoke on usable KVM

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci(e2e): probe KVM through VM driver

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci(e2e): tolerate hosted KVM denial

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(server): close traced futures before assertions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* revert: remove tracing test stabilization

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Evan Lezar <elezar@nvidia.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
2026-07-31 16:41:06 +00:00
Derek Carr 5952a5a23f feat(workspace): add workspace resource model with scoping, membershi… (#2243)
* feat(workspace): implement workspace model (Phase 1 of RFC 0011)

Implements workspace and membership model providing hard isolation
boundaries for multi-player OpenShell deployments.

Workspace CRUD with Kubernetes-style Terminating phase for graceful
deletion. All resources scoped by workspace via ObjectMeta. Membership
RPCs for workspace access control. Persistence migration shifts name
uniqueness to (object_type, workspace, name). Provider profiles support
platform and workspace scoping. Service routing uses workspace-prefixed
DNS labels. Inference routes renamed and workspace-scoped with
DeleteInferenceRoute RPC. Python SDK with WorkspaceClient, two-method
list pattern (workspace-scoped and for_all_workspaces), and workspace
parameter on all methods. CLI workspace flags, TUI workspace cycling.
K8s driver filters unmanaged CRs and uses delete preconditions. Podman
driver uses immutable container IDs. Label serialization fixed across
all put_if call sites.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(cli): delegate sandbox upload command to existing upload function

The standalone `sandbox upload` command reimplemented upload logic
inline with two bugs: it used `Path::exists()` which follows symlinks
(rejecting dangling symlinks), and it ran git-aware filtering on
symlink sources. The `run::sandbox_upload()` function already handles
both cases correctly via `sandbox_upload_plan()`. Replace the inline
logic with a call to the existing function.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(e2e): shorten sandbox names and fix test compatibility

Shorten the sandbox name in initial_sparse_policy_is_acknowledged_as_loaded
from 'e2e-2159-sparse-enrich' (22 chars) to 'e2e-sparse-enrich' (17 chars)
to comply with MAX_ROUTABLE_NAME_LEN (19 chars).

Also capture stderr in create_keep_with_args so future sandbox creation
failures include the actual CLI error instead of reporting empty output.

Signed-off-by: Derek Carr <decarr@redhat.com>

* test(workspace): add test coverage for workspace CRUD and persistence isolation

Add unit tests for workspace create happy path, get round-trip, get
not-found, get empty-name rejection, already-exists error, and
resolve_workspace not-found. Add persistence test proving cross-workspace
name uniqueness (same name in different workspaces produces separate
records). Add workspace name max-length boundary tests. Fix e2e harness
to include stderr in name-parse-failure error path. Align Python e2e
test_workspace_crud with try/finally pattern. Document provider profile
catalog workspace scoping gap in RFC 0011.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(examples): update examples for workspace model compatibility

Shorten sandbox names in demo scripts to fit the 19-character
MAX_ROUTABLE_NAME_LEN limit: policy-demo prefix to pd-, multi-agent
notepad derives a short SANDBOX_TAG from the run ID, governance
interceptor uses gs-PID-RANDOM. Update vscode-remote-sandbox.md SSH
host aliases from openshell-{name} to openshell-{name}.{workspace}
format.

Signed-off-by: Derek Carr <decarr@redhat.com>

* feat(sdk): add workspace-scoped client and workspace CRUD

Add WorkspaceScopedClient modeled after kube::Api::namespaced — captures
workspace once and injects it into every sandbox request. Add workspace
CRUD methods (create, get, list, delete) and list_sandboxes_all_workspaces
on OpenShellClient. Extend SandboxRef with workspace field and add
WorkspaceRef type. Include mock tests for all new operations.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(lint): resolve clippy warnings in workspace test assertions

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(docs): convert indented code blocks to fenced in RFC 0011

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(lint): resolve clippy warnings and apply cargo fmt across workspace

Auto-format with cargo fmt and fix clippy warnings exposed by the
reformat: unnecessary qualifications, map_unwrap_or, identical match
arms, unused variable prefix, dead code annotations, and let-unit-value
in e2e harness.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(workspace): address workspace scoping issues from review

- Add workspace field to settings JSON output (CLI)
- Skip Podman containers missing workspace label instead of defaulting
  to empty string, matching K8s driver behavior
- Add resource_version to list_by_scope SELECT in both SQLite and
  Postgres backends, with regression test
- Gate PolicyLocalContext proposal/lookup routes on workspace readiness,
  returning 503 when workspace is not yet discovered
- Block sandbox and provider creation in TUI all-workspaces mode
- Clear workspace vectors in TUI reset_sandbox_state

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(workspace): make provider profile catalog workspace-aware

Thread workspace through snapshot_catalog so the
EffectiveProviderProfileCatalog enforces workspace boundaries on both
read and write paths. UserProviderProfileSource now loads platform-scoped
profiles (workspace "") plus the target workspace's profiles, preventing
cross-workspace duplicate profile ID collisions that previously caused
global catalog failures.

Update RFC 0011 to reflect catalog scoping is implemented in Phase 1
rather than deferred to future work.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(persistence): include workspace column in atomic policy revision INSERT

put_policy_revision_atomic omitted the workspace column from the INSERT
into the objects table in both SQLite and Postgres backends, causing
atomically-written policy revisions to lose their workspace association.
Add workspace field to AtomicPolicyRevisionWrite and thread it through
both backend INSERT statements, matching the non-atomic put_policy_revision
path which already included it.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(proxy): skip ancestor walk when socket owner is the entrypoint

collect_ancestor_identities walked the entire process tree above the
entrypoint when the connecting process was the entrypoint itself,
SHA256-hashing every ancestor binary (IDE, shell, container runtime).
On dev machines with large binaries in the ancestor chain this exceeded
the 30-second test timeout. When start_pid == stop_pid there are no
intermediate ancestors to verify, so return an empty list immediately.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(workspace): make provider profile catalog scope-aware

Allow the same profile ID at platform and workspace scopes by
introducing layered catalog entries where workspace profiles shadow
platform profiles. Add source and scope fields to the ProviderProfile
proto and CLI output. Migrate List/Get handlers to the catalog,
fixing divergence with runtime profile resolution.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(e2e): align podman e2e labels with centralized driver constants

The podman driver moved its container labels to the centralized
openshell.ai/ prefix, but the e2e test harness and cleanup script
still referenced the old openshell.sandbox-* keys, causing the
local_driver_token_restart test to fail on container lookup.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(e2e): align python profile isolation test with scope-aware catalog

Platform profiles are now visible in workspace listings as fallbacks
per the layered catalog design. Update the assertion to match.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(workspace): honor profile_workspace in runtime profile resolution

Runtime profile lookups now consult provider.profile_workspace via
get_type_profile_for_scope. Providers created with --global-profile
(profile_workspace="") resolve to the platform profile even when a
workspace profile shadows the same ID. All 6 runtime call sites
updated; type-only call sites remain scope-agnostic.

Signed-off-by: Derek Carr <decarr@redhat.com>

---------

Signed-off-by: Derek Carr <decarr@redhat.com>
2026-07-21 00:59:18 +00:00
Kyle ZhengandJohn Myers ff9af8e320 fix(sandbox): acknowledge initial policy revision; expose SDK labels/selectors (#2170)
* fix(sandbox): acknowledge initial policy revision

The supervisor loaded and enforced a sandbox-scoped policy but never told
the gateway which revision it loaded. The policy poll loop seeded itself
with the initial revision's hash on its first poll, so `policy_changed`
was never true for that revision and `ReportPolicyStatus(LOADED)` — which
only ran in the hot-reload branch — was never called. The revision stayed
`Pending` and `current_policy_version` stayed 0 even though the sandbox was
`Ready` and the policy was effective. This was most visible with sparse
policies that get baseline-enriched into a new revision during startup.

After the OPA engine is constructed, report the exact sandbox revision the
supervisor loaded as LOADED, and seed the poll loop from that revision so
it is not re-reported. Report FAILED with the original construction error
if engine construction or conversion fails. Only sandbox-sourced revisions
(version > 0) whose canonical content matches the loaded policy are
acknowledged; global and local-file policies are untouched. Delivery uses
the shared bounded retry, is non-fatal on transient failure, and a pending
initial acknowledgement is delivered before any newer revision so policy
history is never reordered.

Signed-off-by: Kyle Zheng <kyzheng@nvidia.com>

* feat(python): expose sandbox labels and selectors

The gateway protobuf and CLI already support request-level sandbox labels
(`CreateSandboxRequest.name`/`labels`) and selector-based listing
(`ListSandboxesRequest.label_selector`), but the public Python SDK dropped
them, so Python-created sandboxes could not be found via
`openshell sandbox list --selector ...`.

Add optional, source-compatible `name`/`labels` to `SandboxClient.create`,
`create_session`, and the high-level `Sandbox`, and `label_selector` to
`list`/`list_ids`. `SandboxRef` now carries the gateway labels as an
immutable mapping (default empty, so `SandboxRef(id, name, status)` still
works). Caller-provided label mappings are copied. Attaching the high-level
`Sandbox` to an existing sandbox rejects `name`/`labels` since creation
metadata cannot change on attach. Template labels remain a separate concept.
No protobuf changes are required.

Signed-off-by: Kyle Zheng <kyzheng@nvidia.com>

* fix(python): keep SandboxRef hashable and copy high-level labels

Excluding the new immutable `labels` field from SandboxRef equality/hash
(`compare=False`) preserves the original (id, name, status) identity and keeps
the frozen dataclass hashable — a MappingProxyType field would otherwise make
`hash(SandboxRef(...))` raise. Also defensively copy caller-provided labels in
the high-level `Sandbox` so later caller mutation cannot change what is sent.

Signed-off-by: Kyle Zheng <kyzheng@nvidia.com>

* fix(sandbox): bound initial-policy-ack retries

The poll loop retried a pending initial acknowledgement before processing any
newer revision, but retried unconditionally forever. A permanently
undeliverable ack (e.g. the revision was superseded before it could be
reported) would then stall all later policy hot-reloads and provider-env
refreshes. Cap the retries; after the bound, give up and resume normal polling
so the loop cannot livelock on a stuck acknowledgement.

Signed-off-by: Kyle Zheng <kyzheng@nvidia.com>

* test(sandbox): add sparse-policy revision-2 acknowledgement e2e

Regression for #2159: create a sandbox with the network-only policy-advisor
fixture, which the supervisor enriches with baseline filesystem paths during
startup (creating revision 2, superseding revision 1). Assert the effective
policy reaches revision 2 and no revision remains Pending once the supervisor
acknowledges the load. Adds SandboxGuard::create_keep_with_args to create a
kept sandbox with an initial --policy.

Signed-off-by: Kyle Zheng <kyzheng@nvidia.com>

* fix(ci): correct sandbox checks

Signed-off-by: Kyle Zheng <kyzheng@nvidia.com>

* test(cli): serialize mTLS environment access

Signed-off-by: Kyle Zheng <kyzheng@nvidia.com>

* fix(sandbox): address policy review feedback

Signed-off-by: Kyle Zheng <kyzheng@nvidia.com>

* fix(sandbox): preserve exact policy acknowledgements

Signed-off-by: Kyle Zheng <kyzheng@nvidia.com>

* fix(sandbox): preserve local policy overrides

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: Kyle Zheng <kyzheng@nvidia.com>
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-07-08 17:24:26 -07:00
Mesut Oezdil b689c82075 fix(python): add encoding=utf-8 to file reads and writes in sandbox.py (#1912)
* fix(python): add encoding=utf-8 to all file reads and writes in sandbox.py

read_text(), fdopen() without encoding use the system locale, which can
differ on non-UTF-8 systems. All config files (metadata.json,
oidc_token.json, active_gateway) are UTF-8 text.

* fix(python): catch UnicodeDecodeError in _read_oidc_token_bundle, add encoding tests

Corrupt or non-UTF-8 oidc_token.json now returns None consistently.
Add regression tests for non-ASCII UTF-8 paths in metadata, oidc token,
and active_gateway files.

* fix(python): use ensure_ascii=False in encoding tests, add from_active_cluster UTF-8 bytes test

* fix(python): assert non-ASCII cluster name and endpoint in encoding test

* fix(python): wrap long write_bytes line for ruff format

* fix(python): apply ruff format to sandbox_test.py
2026-06-22 12:35:11 -05:00
Zygmunt Krynicki 4025894a94 chore(snap): remove early snap packaging (#1648) 2026-06-08 22:09:09 -07:00
Taylor Mutch 7cea9d9b2b fix(gateway): align package TLS bootstrap path (#1601)
* fix(gateway): align package TLS bootstrap path

Closes #1593

Default package-managed gateway services to a stable local TLS directory and use that same value for certificate generation and runtime startup.

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* test(packaging): validate package asset paths exist

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* ci(e2e): pin mise in kubernetes job

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

---------

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
2026-06-01 14:56:10 -05:00
Alexander Watson e98ea3ee93 feat(policy): add agentic approval loop (#1528) 2026-05-29 21:11:31 -07:00
Mrunal Patel fb03e3819a feat(python-sdk): support OIDC Bearer auth on SandboxClient (#1621)
* feat(python-sdk): support OIDC Bearer auth on SandboxClient

PR #1596 hardened the gateway side of the OIDC story; the Python SDK
was the remaining gap — it only supported plaintext or mTLS, with no
Bearer metadata anywhere. Deployments with OIDC enabled (the
recommended posture since PR #935 / PR #1404) were unreachable from
the SDK.

Adds:

- `bearer_token: str | Callable[[], str] | None` kwarg on
  `SandboxClient`. Static strings or zero-arg callables (the latter
  is invoked per RPC, so callers can drop in a refresh loop or
  token-file watcher without reconstructing the client). Composes
  with `tls` for OIDC-over-mTLS deployments.
- `_BearerAuthInterceptor` implementing all four
  `grpc.{Unary,Stream}{Unary,Stream}ClientInterceptor` types.
  Appends `authorization: Bearer <token>` to outgoing metadata.
  Implemented as an interceptor (not call credentials) so it works
  on both plaintext (`disableTls=true` dev) and TLS channels without
  `grpc.composite_channel_credentials`.
- `TlsConfig` ergonomics: all three fields (`ca_path`, `cert_path`,
  `key_path`) are now optional with `cert_path` / `key_path`
  required-together-or-not-at-all (enforced in `__post_init__`). This
  unlocks three transport profiles from one dataclass:
    * full mTLS (all three)
    * CA-only trust (`ca_path` only)
    * system roots (`TlsConfig()` — for OIDC gateways behind a
      public CA)
- `from_active_cluster` mirrors `crates/openshell-tui/src/lib.rs`
  `build_oidc_channel`:
    * For any `https://` gateway, always build a secure channel.
      Pick the strongest TLS profile available in `mtls/` (full
      mTLS → CA-only → system roots). No more `insecure_channel`
      fallback for HTTPS.
    * Gate OIDC bearer attachment on
      `metadata.json["auth_mode"] == "oidc"`. Matches
      `crates/openshell-cli/src/main.rs:132` and the TUI; a stale
      `oidc_token.json` next to a non-OIDC gateway no longer causes
      the SDK to attach a bearer.
- `_OidcRefresher` — thread-safe, in-process native OAuth2 refresh
  modeled on `google.oauth2.credentials.Credentials` and
  `botocore.tokens.SSOTokenProvider`. Lazily checks expiry on every
  RPC; when stale, re-reads disk first (the CLI may have rotated
  the bundle), and only then exchanges the refresh_token against
  the IdP's token endpoint discovered via OIDC discovery
  (`/.well-known/openid-configuration`, cached after first call).
  Concurrent RPCs share a single refresh via `threading.Lock` (no
  IdP stampede). Honors refresh-token rotation. Surfaces IdP
  failures as `SandboxError` with the RFC 6749 error body included
  for diagnostics.

  Mirrors the Rust CLI's HTTP-policy posture from
  `crates/openshell-cli/src/oidc_auth.rs`:
    * `follow_redirects=False` so a 3xx during discovery can't
      steer us to an attacker-controlled token endpoint.
    * Discovery `issuer` is validated against the configured
      issuer; a discovery document claiming a different issuer is
      rejected, preventing the SDK from POSTing the refresh_token
      to a malicious endpoint.
    * `insecure: bool` flag plumbed through to httpx's
      `verify=` so self-signed-cert deployments work the same way
      they do in the Rust CLI.

  Built on `httpx` (chosen over `urllib` specifically for
  follow_redirects + verify control as kwargs). The OAuth2
  refresh-token grant itself (RFC 6749 §6) is one form-encoded
  POST — handled inline rather than via a dedicated OAuth library;
  tried `authlib`'s `OAuth2Client` first but it auto-injects an
  Authorization header on every request, which breaks the
  unauthenticated discovery GET.
- `_make_cluster_bearer_provider(..., auto_refresh=True,
  write_back=True, insecure=False)` factory. Defaults to the
  refresher path with write-back enabled; `auto_refresh=False`
  falls back to the read-only fail-closed behavior for callers that
  don't want the SDK to make outbound HTTP calls to the IdP.

  `write_back=True` is the default (changed from the first round of
  review): IdPs with refresh-token rotation (Keycloak with
  rotation, Entra in strict mode) invalidate the old refresh_token
  on each refresh, so an in-memory-only refresh would leave the
  on-disk bundle pointing at an invalidated value — any second
  process starting from disk would `invalid_grant`. With write-back
  enabled by default, the SDK keeps the shared cache consistent
  with the IdP.
- `from_active_cluster` exposes `auto_refresh`, `write_back`, and
  `insecure` kwargs (defaults: True / True / False). The
  high-level `Sandbox` context manager surfaces the same three
  kwargs and forwards them through, so callers using the wrapper
  have parity with `SandboxClient` for OIDC-protected gateways.
- `SandboxClient.close()` chains to a `_bearer_close` hook so the
  `_OidcRefresher`'s underlying `httpx.Client` is released
  deterministically instead of leaking sockets/FDs until GC runs
  `__del__`. Idempotent.
- `_OidcRefresher._write_to_disk` uses `tempfile.mkstemp` (PID +
  random suffix) instead of a fixed `.oidc_token.json.tmp` path,
  so two writers racing on the same gateway directory don't
  trample each other's tmp content. Success path atomically
  replaces; failure path unlinks the orphan.

OAuth2 refresh policy and write-back semantics deliberately mirror
what the major Python SDKs do — see
github.com/googleapis/google-auth-library-python (`Credentials`)
and github.com/boto/botocore (`SSOTokenProvider`):

| Library                       | Native refresh | Writes back |
|-------------------------------|----------------|-------------|
| google-auth Credentials       | yes            | no          |
| botocore SSOTokenProvider     | yes            | yes         |
| openshell SandboxClient (here)| yes (opt-out)  | yes (opt-out)|

OpenShell sits between the two; chose write-back-by-default because
the rotation invariant matters more for our deployments than the
"CLI is the only writer" assumption that fits google-auth.

Adds `httpx>=0.27` as a runtime dependency. No new OAuth2 library —
the refresh grant is a single POST.

Tested:

- 42 sandbox_test.py tests pass (5 pre-existing + 37 new across
  the bearer interceptor, fail-closed provider, refresher
  behavior, TlsConfig validation, from_active_cluster auth ladder,
  security-review regressions, Sandbox-wrapper kwarg forwarding,
  and lifecycle / concurrency probes).
  `mise run test:python` → 47 passed total across the python
  suite.
- `mise run python:lint` (ruff) clean.
- End-to-end against a Keycloak-protected gateway on OpenShift:
    * unauthenticated `Health` bypass works
    * admin + `openshell:all` reaches user-callable methods
    * reader (`sandbox:read`) denied on `CreateSandbox` by scope
    * admin + `openshell:all` denied on PR #1596 sandbox-only
      methods at the router (the new gate is honored from the SDK)
    * full provider CRUD lifecycle via the SDK
    * callable token provider rotates per RPC as expected
- Regression-probed against three pre-review security findings:
    * **Discovery issuer validation** — a discovery document
      claiming a different `issuer` than the configured one is
      rejected with a clear `SandboxError` before any refresh POST
      can reach the attacker-controlled endpoint.
    * **Redirect during discovery** — `follow_redirects=False` on
      the underlying httpx client means a 3xx during discovery
      surfaces as a SandboxError rather than silently chasing the
      redirect.
    * **Cross-process rotation** — a two-process simulation shows
      process B starting from disk and successfully refreshing
      with the rotated refresh_token, because process A's
      write-back updated the shared cache.
- Refresher unit tests cover: cached-fresh fast path, disk-rotated
  re-read before refresh, OAuth2 exchange against the discovered
  token endpoint, refresh-token rotation, atomic write-back at
  0600 mode (default), default-on write_back proven by test,
  concurrent N-thread coordination (one refresh shared across 8
  threads), IdP failure surfaced with the error body, the
  client_credentials / no-refresh_token error path, issuer-
  mismatch rejection, redirect-during-discovery rejection,
  insecure flag plumbing.
- Lifecycle / concurrency regression tests added: `close()`
  invokes the `_bearer_close` hook (idempotent), the refresher's
  `httpx.Client` is marked closed after `SandboxClient.close()`,
  and 16 concurrent writers don't leave orphan tmp files behind
  while producing a valid final bundle. The `Sandbox` wrapper has
  direct forwarding tests proving `auto_refresh`, `write_back`,
  and `insecure` reach `from_active_cluster` (both explicit
  values and defaults).
- End-to-end against a real OpenShift + Keycloak cluster from
  inside a pod: real OIDC discovery against
  `keycloak.keycloak.svc.cluster.local:8080`, refresh-token grant
  POST, atomic write-back of the rotated bundle at 0600, and a
  follow-up RPC reusing the freshly-rotated in-memory token —
  full round-trip in ~170ms.

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(python-sdk): adopt newer on-disk OIDC bundle before refreshing

_OidcRefresher.current_access_token() only adopted the on-disk
oidc_token.json when its access token was still fresh; otherwise it
refreshed using the in-memory bundle. With refresh-token rotation
enabled (Keycloak with rotation, Entra strict mode), this let a process
keep using an invalidated refresh_token:

1. Process A holds a stale in-memory bundle with refresh_token=r1.
2. Process B refreshes first and writes a rotated (r2) but now
   near-expiry bundle to disk.
3. Process A re-reads disk, sees the access token is not fresh, ignores
   the disk bundle, and POSTs the stale r1 — which the IdP has already
   invalidated, yielding invalid_grant.

Fix: when the cached bundle is stale, adopt the on-disk bundle if it was
refreshed more recently than ours, even when its access token is also
stale. "More recently" is decided by expires_at — a refresh mints a new
access token with a forward expiry alongside the rotated refresh_token,
so the later expiry carries the newest refresh_token. Comparing by
expiry (rather than unconditionally preferring disk) preserves the
write_back=False case, where the in-memory bundle has already rotated
past the on-disk copy and must not be clobbered. When the adopted
bundle's issuer differs, the cached token endpoint is reset so the
refresh re-discovers against the new issuer.

Adds regression tests for the cross-process rotation race and the
issuer-change re-discovery path.

* fix(python-sdk): recover from invalid_grant on lost rotation race

The expiry-based disk re-read narrows but does not fully close the
cross-process refresh-token rotation race: two processes sharing a
gateway directory can both enter their refresh window, both POST their
copy of the refresh_token, and with rotation enabled the IdP invalidates
the loser's token (invalid_grant). Neither google-auth nor botocore
close this window without an OS file lock; a Python-only flock would not
coordinate with the Rust CLI/TUI that also write oidc_token.json, so
locking is not worth its cost here.

Recover instead of prevent: distinguish an OAuth2 invalid_grant (the
refresh_token was rejected) from transport/5xx failures via a private
_InvalidGrantError, and on invalid_grant re-read oidc_token.json once. If
a peer wrote a different refresh_token (it won the race), adopt and retry
with it — returning early if it is already fresh — so the loser succeeds
transparently instead of forcing a re-authenticate. If disk offers no new
token, the rejection is genuine and surfaces the re-authenticate hint as
before. The retry is single-shot; a second invalid_grant propagates.

Adds tests for the peer-rotation recovery and the genuine-rejection
(no-retry) paths.

---------

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
2026-05-29 11:44:20 -04:00
Derek Carr d01d10650b refactor(proto): move phase and current_policy_version into status (#1565)
* refactor(proto): move phase and current_policy_version into SandboxStatus

Move phase and current_policy_version from SandboxSpec into
SandboxStatus to correctly model mutable runtime state. Update all
callers in the gateway server, TUI, and Python SDK to read and write
these fields through SandboxStatus accessors.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(server): preserve sandbox status on statusless driver updates

When a driver update arrives without a status payload (e.g. before
Kubernetes populates the status subresource), preserve the stored
phase, conditions, and current policy version instead of resetting
them. Adds a regression test covering the edge case.

Signed-off-by: Derek Carr <decarr@redhat.com>

---------

Signed-off-by: Derek Carr <decarr@redhat.com>
2026-05-29 08:39:12 -07:00
Mesut Oezdil 4848c40951 fix(python): raise SandboxError instead of FileNotFoundError or KeyError (#1547)
* fix(python): raise SandboxError instead of FileNotFoundError or KeyError

* fix(python): suppress exception chaining in SandboxError raises

Add `from None` to both `raise SandboxError(...)` calls inside `except
FileNotFoundError` blocks to satisfy ruff B904.
2026-05-26 08:34:31 -07:00
Mesut Oezdil cd70249629 ci: pin azure/setup-helm and helm/kind-action to commit SHAs (#1544)
* ci: pin azure/setup-helm and helm/kind-action to commit SHAs

* chore(python): add py.typed marker for PEP 561 compliance

* ci: use full semver in pinned action version comments

Signed-off-by: mesutoezdil <mesudozdil@gmail.com>

---------

Signed-off-by: mesutoezdil <mesudozdil@gmail.com>
2026-05-25 14:03:32 -07:00
Taylor Mutch a3b16c18ab feat(auth): per-sandbox authentication to gateway (#1404) 2026-05-21 17:58:55 -07:00
Drew Newberry f257ed0193 refactor(packaging): rely on gateway runtime defaults (#1415)
* fix(packaging): use gateway TOML config in packages

* refactor(packaging): rely on gateway runtime defaults
2026-05-18 14:13:19 -07:00
Taylor MutchandDrew Newberry b61a98dbad feat(gateway): add TOML configuration file (RFC 0003) (#1317)
* feat(gateway): add TOML configuration file (RFC 0003)

Introduces an opt-in --config / OPENSHELL_GATEWAY_CONFIG flag that loads a
TOML file with gateway-wide settings and per-driver tables. Source
precedence is CLI > env > file > built-in default, implemented via clap's
ValueSource so existing flags and env vars keep their priority.

Driver crates (kubernetes, docker, podman, vm) now derive Deserialize on
their config structs. SupervisorSideloadMethod gains Deserialize with
kebab-case rename. A per-driver inheritance allowlist on the loader side
overlays [openshell.gateway] shared defaults (default_image,
supervisor_image, image_pull_policy, guest_tls_*, ssh_handshake_skew_secs,
client_tls_secret_name, host_gateway_ip, enable_user_namespaces) onto
each [openshell.drivers.<name>] table before deserialization.

The Helm chart renders a new gateway-config ConfigMap and mounts it at
/etc/openshell/gateway.toml. The migrated OPENSHELL_* env entries are
dropped from the StatefulSet — only the Secret-backed
OPENSHELL_SSH_HANDSHAKE_SECRET remains. database_url stays on --db-url.

Adds examples/gateway/gateway.example.toml and updates architecture/gateway.md
with the source precedence and inheritance rules.

* docs(gateway): drop ssh_handshake_skew_secs and ssh_handshake_secret from examples

Both fields are scheduled for removal. Remove the example values and the
env-only note so the gateway.toml example and the architecture doc stop
recommending settings that will not exist much longer.

* docs(rfc): correct OPENSHELL_CONFIG to OPENSHELL_GATEWAY_CONFIG in RFC 0003

* docs(gateway): add per-driver TOML example configurations

Adds focused single-driver examples next to the comprehensive
gateway.example.toml: kubernetes, docker, podman, and microvm. Each one
demonstrates the realistic settings for that driver plus how shared
[openshell.gateway] defaults inherit into the driver table.

A new unit test (`checked_in_examples_parse`) loads every example through
the config_file loader so schema drift fails CI rather than silently
shipping a broken example.

* refactor(gateway): drop image_pull_policy from shared inheritance

Kubernetes and Podman use mutually-incompatible vocabularies for the same
TOML key:

  - Kubernetes: `Always | IfNotPresent | Never` (free-form string passed
    verbatim to the K8s API).
  - Podman: `always | missing | never | newer` (strict lowercase enum
    deserialised into `ImagePullPolicy`).

No value means the same thing in both drivers. Sharing the key at
`[openshell.gateway]` scope and inheriting it into every active driver's
table meant any value safe for one driver was either wrong or silently
dropped for the other (`IfNotPresent` → `ImagePullPolicy::Missing` after
`.unwrap_or_default()`). Operators run one driver per gateway, so the
"shared default" never pays for itself.

Make `image_pull_policy` driver-local:

  - Remove the field from `GatewayFileSection` and from
    `inheritable_keys()` for both Kubernetes and Podman.
  - Drop the file→`RunArgs` merge for the gateway-scope key.
  - Stop unconditionally clobbering the driver value with
    `config.sandbox_image_pull_policy` in the runtime wiring — only apply
    the CLI/env override when it was set (and, for Podman, only when it
    parses into the lowercase enum).
  - Move the key under `[openshell.drivers.kubernetes]` and
    `[openshell.drivers.podman]` in every example, the RFC, the
    architecture doc, and the Helm-rendered gateway ConfigMap.

The supervisor pull policy follows the same shape: it is K8s-only and
moves into `[openshell.drivers.kubernetes]` alongside `image_pull_policy`
in the Helm template.

* fix(gateway): address review feedback on TOML configuration

Resolves the P1 and P2 issues raised in PR #1317:

- Helm gateway ConfigMap moves `grpc_endpoint` under
  `[openshell.drivers.kubernetes]` so the default install no longer fails
  the gateway's `deny_unknown_fields` schema check.
- `kubernetes_config_from_file` and `podman_config_from_file` only let
  the gateway-wide CLI/env `grpc_endpoint` overwrite the driver-table
  value when it was actually supplied, preserving file-only configs.
- Kubernetes driver default `image_pull_policy` is now empty (was Podman
  vocabulary "missing"), so default deployments let the Kubernetes API
  apply its own policy instead of being rejected.
- New `disable_tls` gateway field plumbs `.Values.server.disableTls`
  through the TOML ConfigMap instead of relying on env vars dropped from
  the StatefulSet.
- StatefulSet pod template now carries a `checksum/gateway-config`
  annotation so `helm upgrade` rolls pods when the ConfigMap changes.
- Auxiliary listener resolution preserves the full `SocketAddr` from
  `health_bind_address` / `metrics_bind_address`, so a loopback-pinned
  health port is not silently relocated onto the public bind address.
- `ssh_session_ttl_secs` from the file is now applied to `Config` (it
  was previously accepted by the loader but never read).

New regression coverage: cli-level merge tests for the new fields plus
helm-unittest assertions for the ConfigMap shape, checksum annotation,
and `disable_tls` rendering.

* docs(gateway): consolidate gateway TOML examples into docs reference

Replaces the per-driver example files under examples/gateway/ with a
single published reference page at docs/reference/gateway-config.mdx
covering source precedence, layout, the full example, and the four
per-driver examples (Kubernetes, Docker, Podman, microVM). Drew flagged
during PR #1317 review that the examples belong with the user-facing
docs rather than in a sibling examples/ directory.

The cross-references in architecture/gateway.md and RFC 0003 are updated
to point at the new docs page; the round-trip test in config_file.rs is
removed (schema coverage stays on the inline parses_full_example test
and per-field merge tests — doc-snippet drift belongs in a separate
docs-lint, not in a cross-tree Rust unit test).

* refactor(core): move DEFAULT_K8S_NAMESPACE into K8s driver

The constant is Kubernetes-specific (used only by KubernetesComputeConfig's
Default impl) and does not belong in openshell-core. Relocate it to the
driver crate that owns the K8s vocabulary; openshell-core retains only
truly cross-cutting defaults.

* refactor(core): move Podman bridge default into Podman driver

DEFAULT_NETWORK_NAME is Podman vocabulary, consumed only by the Podman
driver. Also drops the unused DEFAULT_IMAGE_PULL_POLICY constant.

* docs(auth): scrub remaining SSH handshake secret references

Sweeps the trailing mentions left after the rebase: the gateway
config-file module doc, the Helm gateway-config ConfigMap header,
the gateway-config.mdx env-only note, and the RPM systemd unit
comment for init-gateway-env.sh.

* docs(gateway): clarify OPENSHELL_GRPC_ENDPOINT applies to all drivers

The previous comment implied the callback endpoint was Kubernetes-only,
but the value is propagated to every compute driver (Kubernetes, Docker,
Podman, VM) and must be reachable from wherever the sandbox runs.

* refactor(gateway): move driver options into config (#1394)

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(e2e): regenerate gateway config via TOML for docker + podman harnesses

The gateway CLI flags moved into TOML config tables in 560550d2 (#1394),
which made every existing e2e/with-{docker,podman}-gateway.sh invocation
fail with "unexpected argument '--sandbox-namespace'" (and a long tail
of similar driver-specific options) before the gateway could even bind.

Replace the obsolete CLI flags with a synthesized
`[openshell.drivers.<driver>]` table written to `${STATE_DIR}/gateway.toml`
and passed via the new `--config` flag. Only the gateway-wide flags
that survived 560550d2 (bind-address, port, drivers, db-url, tls-*,
disable-tls, log-level, health-port) stay on the command line.

Both scripts get a small `toml_string` helper to properly TOML-quote
the values (the previous `%q` printf format produced bash-escape, not
TOML-escape). The Docker harness also corrects two field names that
diverged from the driver schema: `docker_network_name` →
`network_name`, and the supervisor binary/image plumbing now reads
through to `supervisor_bin` / `supervisor_image` in the same table.

The Podman harness drops `--ssh-gateway-port` (deleted in 560550d2 —
gRPC + SSH are multiplexed on the same port now) and substitutes
`network_name` + `gateway_port` for the obsolete `--sandbox-namespace`
(which the Podman driver never had as a typed field).

* fix(core): swap bind-only 0.0.0.0 SSH gateway host for cluster URL host

CLI's resolve_ssh_gateway treated 0.0.0.0 as a loopback "keep as-is"
when the cluster URL was also loopback, so the SSH proxy connected to
0.0.0.0:port. The unspecified address is never a valid connect target
and is not present in any TLS cert SAN, which produced BadCertificate
TLS handshake failures during `openshell sandbox create -- ...` in
docker/podman e2e (e.g. bypass_detection).

Resolution: when the server returns 0.0.0.0 or :: as the gateway host
and both endpoints are loopback, fall back to the cluster URL's host
(which the CLI is already using to reach the gateway, so it must
resolve and match the cert).

* fix(e2e): repair podman harness on macOS

Podman 5.x with the applehv/libkrun provider no longer creates the legacy
~/.local/share/containers/podman/machine/podman.sock symlink, and
`podman system service` is a Linux-only subcommand — the macOS client
delegates the API service to the VM. Both assumptions in the harness
were stale, so the script tried to start a temporary service that podman
rejected with "unknown flag: --time".

- Discover the macOS socket via `podman machine inspect` instead of the
  hardcoded path.
- On Darwin, fail fast with a "start podman machine" message rather than
  attempting the Linux-only `podman system service` fallback.
- Write socket_path into [openshell.drivers.podman] so the in-process
  driver picks up the discovered socket; the driver reads TOML only
  after the config refactor (560550d2), so OPENSHELL_PODMAN_SOCKET alone
  was no longer enough.

* fix(server): use clone_from for TLS client CA assignment

clippy 1.95.0 rejects assigning the result of `Clone::clone()` to an
existing variable under `-D warnings` (`assigning_clones`). Switch to
`clone_from(&...)` to satisfy the lint and avoid the redundant
allocation.

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
2026-05-15 12:43:48 -07:00
Drew Newberry 0471c6d2ab fix(gateway): keep vm driver opt-in (#1375) 2026-05-13 22:32:58 -07:00