Commit Graph
355 Commits
Author SHA1 Message Date
Jim Meyer dbe8a0854b chore(agents): resolve contributor guidance review gaps
Signed-off-by: Jim Meyer <jimeyer@nvidia.com>
2026-10-02 08:41:33 -07:00
Oliver Calder 6e865df349 feat(snap): ship the standalone prover binary in the snap (#3717)
Signed-off-by: Oliver Calder <oliver.calder@canonical.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-10-02 11:51:56 +00:00
Matthew GrossmanandEvan Lezar 6048bed368 fix(ci): retry Nix shell and app dependency preparation (#4066)
* fix(ci): prepare Nix development shells in setup-nix

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(ci): retry Nix builds before executing apps once

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Evan Lezar <elezar@nvidia.com>
Co-authored-by: Evan Lezar <elezar@nvidia.com>
2026-10-02 11:27:20 +00:00
John T. MyersandJohn Myers 6e369f2396 chore(agents): simplify contributor instructions and workflows (#3987)
* chore(agents): simplify contributor instructions and workflows

Closes #3980

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* docs(contributing): scope verification to affected components

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* docs(contributing): standardize issue branch naming

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <johntmyers@users.noreply.github.com>
2026-10-01 16:20:35 +00:00
Evan Lezar cb193ef1c2 ci: use package installers consistently in integration tests (#4056)
* test(tmachine): add Fedora RPM package installer

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* ci: qualify Ubuntu branch installs with DEB packages

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* ci: align package installers across integration matrices

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-10-01 15:05:35 +00:00
Evan Lezar 82e889374f test(tmachine): add Fedora RPM package installer (#4025)
Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-10-01 14:29:09 +00:00
Simon ScattonandMrunal Patel 5698c4f746 fix(ci): qualify protobuf compatibility by release train (#4049)
* feat(ci): detect breaking protobuf changes

Compare the proto module against the PR or merge-group base and report Buf violations in Branch Checks. Add local reproduction and fixture coverage.

Closes #3794

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(ci): pin protobuf check container image

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(ci): qualify protobuf compatibility by release train

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* refactor(ci): reuse protobuf compatibility action

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* refactor(ci): run protobuf checks as a Nix app with one ref

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

---------

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
Signed-off-by: Simon Scatton <sscatton@nvidia.com>
Co-authored-by: Mrunal Patel <mrunalp@gmail.com>
2026-10-01 14:00:45 +00:00
Simon Scatton fde79f1aa6 fix(ci): align integration inputs with release candidate source (#4048)
Signed-off-by: Simon Scatton <sscatton@nvidia.com>
2026-10-01 13:01:20 +00:00
Simon Scatton 21fea95935 test(tmachine): add K3s conformance scenario (#3848)
* test(tmachine): add K3s conformance scenario

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* refactor(tmachine): use Helm values file for K3s installer

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* ci(tmachine): run K3s conformance in integration jobs

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* ci(tmachine): verify installer scripts and document version baseline

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

---------

Signed-off-by: Simon Scatton <sscatton@nvidia.com>
2026-09-30 14:49:09 +00:00
Piotr Mlocek c0eb3dbd30 fix(ci): restore repository permission vetters (#3875)
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-09-29 18:05:21 +00:00
krishicks cf1bbb965d docs: remove the architecture directory (#3799)
Remove architecture/. It was a constant source of merge conflicts, became
an effectively append-only log of the project, and was of dubious value.
Design records live in rfc/, crate details in crate READMEs, and user
documentation in docs/.

Move the git-ignored plans directory from architecture/plans to plans/,
keeping the old .gitignore entry. Remove the arch-doc-writer agents and
update AGENTS.md, CONTRIBUTING.md, skills, the feature request template,
and links in proto/, rfc/, and examples/ that pointed into architecture/.

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-09-29 17:35:22 +00:00
Evan Lezar eef8bec0c9 test(e2e): run podman suite with tmachine (#3637)
* test(e2e): remove superseded podman userns coverage

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(tmachine): run podman e2e archive

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(tmachine): generate podman e2e archive inventory

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-09-28 11:15:33 +00:00
Johnny Greco 87929ad130 docs: keep page URLs aligned with file paths (#3713)
* docs: align page file names and nav labels with published URLs

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs: redirect moved dev pages and fix agent guide redirects

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* ci(docs): check that page URLs match their file paths

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(docs): align navigation checks with repo conventions

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs: omit historical redirect aliases

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(docs): preserve published overview and TypeScript URLs

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(docs): redirect unversioned overview and TypeScript URLs

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

---------

Signed-off-by: Johnny Greco <jogreco@nvidia.com>
2026-09-28 02:08:54 +00:00
Drew Newberry c9da461a58 fix(ci): keep snap canary sandbox name within limit (#3751)
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-27 19:20:30 +00:00
Drew Newberry a67567e583 fix(snap): require mTLS for the snap gateway (#3726)
* fix(snap): require mTLS for the snap gateway

Replace the installer opt-in with an authenticated snap gateway. The wrapper
no longer forces plaintext, so the gateway serves TLS from the bundle it
already generates in $SNAP_COMMON/tls. The install hook writes a config that
enables mTLS user auth instead of unauthenticated access, and a new
post-refresh hook migrates the exact legacy default on existing installs.

install.sh waits for the gateway, detects whether it serves TLS, copies the
client bundle into the target user's snap state directory, and registers the
gateway over HTTPS. Older plaintext snap revisions still register over HTTP
with a warning. The release canary asserts mTLS auth and HTTPS registration.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(snap): pass config preflight and detect the mTLS gateway reliably

An explicit [openshell.gateway.mtls_auth] table fails config preflight,
which validates mTLS auth before the local TLS bundle supplies the client
CA. Write a default that pins the Docker driver instead; with the wrapper's
TLS bundle the gateway requires client certificates and enables mTLS user
auth automatically, as the native packages do.

The mTLS gateway rejects TLS handshakes without a client certificate, and
it still answers plaintext loopback HTTP for sandbox service routing, so the
installer could misdetect it as a legacy plaintext gateway. Probe HTTPS with
the root-owned client bundle, and treat a gateway as legacy only when a
plaintext gRPC Health call succeeds.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* chore(snap): simplify install hook comment

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(snap): migrate insecure gateway configs on refresh

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(snap): simplify mTLS detection and config migration

Detect the mTLS snap from the installed revision's post-refresh hook instead
of probing plaintext gRPC, and drop the scheme global. Remove the installer's
pre-hook config fallback, which is dead now that every channel ships the
install hook and which wrote the insecure default. Give the install hook a
single write path with a simple backup name, and shorten the manual client
certificate steps.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(snap): stop keeping a copy of replaced insecure configs

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(install): make the snap an opt-in install method

Stop selecting the OpenShell snap just because the snap command exists. Linux
installs default to the Debian or RPM package; OPENSHELL_INSTALL_METHOD=snap
(or deb, rpm) selects the package explicitly. Hosts that already have the
OpenShell snap keep refreshing it rather than gaining a second gateway on the
same port. The release canary and snap repro script opt in explicitly.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(snap): let the gateway auto-detect its compute driver

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(snap): restart the gateway after refresh

Published revisions use refresh-mode: endure, and snapd honors the old
revision's setting during a refresh, so the plaintext gateway kept running
with the migrated config unused until a manual restart. Restart the gateway
from the post-refresh hook so the mTLS config takes effect immediately.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(snap): drop refresh notes from the snap description

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(snap): trim snap refresh notes from installation docs

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-26 01:50:02 +00:00
Piotr Mlocek 6f596828ad docs(fern): publish ordered version snapshots (#3721)
* docs(fern): publish and order versioned release docs

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(fern): remove version availability badges

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(fern): keep version badges optional

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

---------

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-09-25 19:56:05 +00:00
Polite_realismandDrew Newberry d376c90755 test(podman): close rootful userns, resource-limit, and daemon-failure CI gaps (#3690)
* test(podman): run driver-podman userns suite against rootful Podman too

The driver-specific-integration job only ran the driver-podman testsuite
(default/auto/keep-id/private userns reference checks) against
fedora-podman-rootless, leaving rootful behavior for this scenario
unverified even though the compute driver auto-detects and explicitly
supports rootful Podman.

The default-userns-baseline and userns-profile playbooks hard-asserted a
rootless tmachine gateway user, so pointing them at a rootful environment
would have failed that assertion immediately rather than exercising
anything. They now detect rootful vs. rootless via the existing
tmachine_container_runtime role and branch the reference-capture user
accordingly, while keeping the captured reference file itself owned by
tmachine, since the archived test binary that reads it back always runs
unprivileged as tmachine regardless of daemon mode.

Signed-off-by: politerealism <burdcat17@gmail.com>

* test(podman): add real-daemon coverage for resource limits and daemon failure

Neither the Podman driver's resource-limit enforcement nor its behavior
when the Podman daemon is unreachable had any test coverage against a
real daemon; both were only exercised through unit tests against a
mocked Podman client.

podman_resource_limits.rs creates a sandbox with --cpu/--memory flags and
reads /sys/fs/cgroup/memory.max and cpu.max from inside the sandbox
itself, verifying the limit is actually enforced rather than just echoed
back by the template API. Expected values are cross-checked against the
driver's own parse_cpu_to_microseconds/parse_memory_to_bytes and against
a real local `podman run --cpus/--memory` container.

podman_preflight.rs spawns the standalone openshell-driver-podman binary
against a guaranteed-nonexistent Podman socket and asserts it exits
non-zero within its bounded retry window with an actionable error naming
the socket path, rather than hanging or failing silently.

Signed-off-by: politerealism <burdcat17@gmail.com>

* test(podman): make rootful userns and cgroup checks pass

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(podman): match lifecycle containers by isolation role label

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(podman): accept non-expiring bootstrap tokens

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: politerealism <burdcat17@gmail.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
2026-09-24 20:16:29 -07:00
Jim Meyer 48725c5fcc ci(release): move CodeQL, Trivy, and Zizmor to advisory (#3693)
* ci(release): Move CodeQL, Trivy, and Zizmor to advisory

* docs(ci): describe advisory static findings for tagged releases

Signed-off-by: Jim Meyer <jimeyer@nvidia.com>

---------

Signed-off-by: Jim Meyer <jimeyer@nvidia.com>
2026-09-25 01:26:59 +00:00
Jim Meyer a00ea31c66 ci: restrict copy-pr-bot manual vetters (#3678)
Signed-off-by: Jim Meyer <jimeyer@nvidia.com>
2026-09-24 22:42:51 +00:00
Drew Newberry 52cb8ecee7 fix(kubernetes): remove NetworkPolicy acknowledgement (#3677)
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-24 22:41:45 +00:00
Oliver Calder 6c864ec9ab fix(install): avoid installing incompatible docker snap (#3666)
* fix(install): avoid installing incompatible docker snap

The work to land RFC-0012 added new restrictions when interacting with
Docker by setting `NoNewPrivs`. This prevents the `docker` snap from
transitioning its AppArmor profile from `snap.docker.dockerd` to
`docker-default` when it tries to launch a container. Thus, the `docker`
snap is currently incompatible with OpenShell.

This commit prevents `install.sh` from installing the `docker` snap
before installing the `openshell` snap, and instead requires the user to
install a non-snap Docker daemon before proceeding with installing the
snap. Systems without the `snap` command are unaffected, since they
install native packages without checking for the presence of Docker or
other compute providers.

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* fix(snap): only snapd 2.76 for openshell snap since store installs work

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* fixup! fix(install): avoid installing incompatible docker snap

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* fixup! fix(snap): only snapd 2.76 for openshell snap since store installs work

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

---------

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>
2026-09-24 19:46:00 +00:00
grs 8369bc11a5 fix(podman): restore host gateway alias mediation (#3606)
* fix(podman): restore host gateway alias mediation

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(podman-e2e-tests): enable broader test podman e2e coverage

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(tests): make test more reliable

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(podman): fix macos linting error

Signed-off-by: Gordon Sim <gsim@redhat.com>

---------

Signed-off-by: Gordon Sim <gsim@redhat.com>
2026-09-24 18:41:59 +00:00
Oliver Calder e60098d748 fix(snap): install openshell snap via install.sh when snap available (#3656)
* fix(snap): update stale snap docs and tests

Previously, the `openshell` snap required the `docker` snap. Now, it
works with any Docker daemon running on the system. Furthermore, the
`snap-declaration` assertion on the `openshell` snap when installed from
the Snap Store causes the `openshell` snap to always connect to the
system `:docker` slot, rather than a slot provided by the `docker` snap.

This commit updates the documentation, including the `description` field in
`snapcraft.yaml`, to ensure that all information is correct and
up-to-date.

Additionally, some tests connected the `openshell:docker` plug to the
`docker` snap's `docker:docker-daemon` slot, which is inconsistent with
how the `openshell` snap operates when installed from the store. Update
those tests to connect to the system `:docker` slot as well.

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* fix(snap): require snapd 2.76 for openshell snap

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* fix(nix): align snap gateway reproducer timeout with release-canary

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* fix(snap): require snapd 2.77 for openshell snap

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* fixup! fix(snap): require snapd 2.77 for openshell snap

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* feat(snap): install openshell snap via install.sh when snap available

Change `install.sh` to install the `openshell` snap by default when
snapd is installed on the host. This installs the snap from the
`latest/stable` channel, which should match the most up-to-date release
tag on github.

If `OPENSHELL_VERSION=dev` is set for `install.sh`, then it will install
the `openshell` snap from the `latest/edge` channel, which matches the
latest dev release available on github.

The `openshell` snap currently requires Docker in order to function. If
a Docker daemon is already installed on the system, it will be used by
the `openshell` snap. Otherwise, `install.sh` will install the `docker`
snap first, wait for the Docker daemon to be ready, and then install the
`openshell` snap.

Also, update the `release-canary.yml` to split the `ubuntu-snap` job
into `ubuntu-snap-system-docker` and `ubuntu-snap-provisions-docker`,
which test the two aforementioned scenarios. Previously, `ubuntu-snap`
manually installed a given snap artifact as built from CI, but with
these new jobs, it instead uses `install.sh` to install the published
`openshell` snap from the `latest/edge` track, thus matching the
behavior of the other release canary jobs.

Make corresponding changes to the `nix` guest reproducer, documentation,
and tests.

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* fix(snap): configure local gateway authentication

The `openshell` snap runs the gateway as a systemd system service, which
runs as root. Thus, the mTLS certs are generated by root and stored in a
root-owned directory to which non-root users do not have access. For
this reason, the snap's `openshell-gateway-wrapper` script sets
`OPENSHELL_DISABLE_TLS=true`.

This commit ensures that the `openshell` snap's gateway allows
unauthenticated local access by writing a default `gateway.toml`
configuration file during the install hook, which runs after the snap is
first installed but before services are started. The config file
contains sets `allow_unauthenticated_users = true`.

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* fix(install): configure snap gateway authentication

Recently, a new install hook was added which writes a default config
file for the `openshell` snap to allow unauthenticated local access to
the gateway. This is because the gateway service runs as root and the
mTLS certificates are not accessible to non-root users.

However, the `install.sh` script installs the `openshell` snap from the
snap store, and the published version may not yet have that new install
hook. Or, the user may already have the snap installed, in which case
the install hook does not run. In either case, we need `install.sh` to
ensure that the config file is written to set the gateway auth to
`allow_unauthenticated_users = true`, and then restart the openshell
gateway service.

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* fixup! feat(snap): install openshell snap via install.sh when snap available

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

---------

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>
2026-09-24 15:10:08 +00:00
Mrunal Patel a408f5dd08 chore(kubernetes): update Agent Sandbox to v1.0.3 (#3578)
Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
2026-09-23 18:30:24 +00:00
krishicks c02683688f ci(e2e): run the Kubernetes HA and credential-driver suites on test:e2e-kubernetes (#3626)
PRs labeled test:e2e-kubernetes now run the Kubernetes HA suite and the
Kubernetes credential-driver suite (Kubernetes Secrets and Vault). Both stay
optional and off in merge groups.

- Read the test:e2e-kubernetes label for the HA and credential-driver
  lanes in Branch E2E Checks.
- Point gator at test:e2e for Helm and Kubernetes coverage and at
  test:e2e-kubernetes for gateway high availability and credential
  driver storage.

Refs #3481

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-09-23 15:47:35 +00:00
Jim Meyer fd49df4f41 fix(ci): use approved setup-oras revision (#3625)
Signed-off-by: Jim Meyer <jimeyer@nvidia.com>
2026-09-23 15:35:08 +00:00
Evan Lezar 907f894ebc ci(release): publish prereleases with qualification summary (#3593)
Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-09-23 13:18:30 +00:00
Drew Newberry c8b20bf0a2 ci(windows): seed caches on windows branch (#3576)
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-22 21:12:05 -07:00
Evan Lezar df88bedb31 fix(podman): support rootless user namespace configurations (#3527)
* test(podman): cover user namespace configurations

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(podman): support keep-id runtime groups

Signed-off-by: Evan Lezar <elezar@nvidia.com>

refactor(podman): generalize keep-id group handling

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* ci(podman): run driver integration tests

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-09-23 00:30:29 +00:00
Evan Lezar 3107ff1f82 ci(security): stage release finding enforcement (#3552)
Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-09-22 17:21:23 +00:00
50230616d5 refactor(runtime): retire Community image dependencies (#3386)
* feat(sandbox): default to official Alpine sandbox image

default_sandbox_image() now returns docker.io/library/alpine:3.22, a generic
version-qualified official image, so a fresh install no longer depends on the
community sandbox image catalog. All compute drivers (docker, podman,
kubernetes, vm) inherit this fallback.

Part of #3116.

Signed-off-by: Akram
Signed-off-by: Akram <akram.benaissi@gmail.com>

* feat(deploy): default deployment configs to the official Alpine sandbox image

Update the shared gateway default_image, Helm chart values, the standalone
Kubernetes manifest, and the dev gateway task scripts to use
docker.io/library/alpine:3.22 instead of the community base image, consistent
with default_sandbox_image(). GPU e2e image-build base is left unchanged (CUDA
needs a glibc base).

Part of #3116.

Signed-off-by: Akram
Signed-off-by: Akram <akram.benaissi@gmail.com>

* feat(driver): default to numeric non-root identity for USER-less images

With the default sandbox image now Alpine, images that declare no OCI USER
must start instead of being rejected. When the image declares no USER and
the policy requests none, the Podman and Docker drivers now supply a numeric
non-root identity (DEFAULT_SANDBOX_UID/GID = 1000) instead of rejecting,
matching the numeric-identity behavior of the Kubernetes and VM drivers. The
supervisor's resolved-identity path runs the sandbox as a synthesized
non-root account without the account existing in the image. Images that
declare a USER keep the OCI resolution path unchanged.

Part of #3116.

Signed-off-by: Akram <akram.benaissi@gmail.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(conformance): use Alpine workload image

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* refactor(policy): drop community image /app path from default policy

The restrictive default policy granted read-only access to /app, a directory
that only existed in the community base image. A generic Alpine default has no
/app, so remove it. Landlock best-effort already ignores absent paths; this
just stops advertising a community-specific layout in the default.

Part of #3116.

Signed-off-by: Akram
Signed-off-by: Akram <akram.benaissi@gmail.com>

* docs(config): document Alpine default images

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(podman): report early sandbox termination

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(podman): initialize rootless workspace ownership

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(sandbox): qualify NVIDIA Ubuntu default

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): initialize rootful default workspace

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(sftp): add native sandbox adapter

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sftp): gate runtime helper support to Linux

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sftp): support standard OpenSSH file operations

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sftp): harden rename and special file handling

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(runtime): remove community image dependencies

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(e2e): build provider readiness tool fixture

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(e2e): use a dedicated Noble fixture for Docker tests

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Akram
Signed-off-by: Akram <akram.benaissi@gmail.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Evan Lezar <elezar@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
2026-09-22 14:43:51 +02:00
Simon ScattonandEvan Lezar 24706c175b ci(rust): parallelize branch checks (#3462)
* ci(rust): parallelize branch checks

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* ci(rust): acknowledge trusted sccache environment

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* ci(rust): remove ineffective sccache backend

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* ci(rust): retain dependency boundary checks

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Simon Scatton <sscatton@nvidia.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>
Co-authored-by: Evan Lezar <elezar@nvidia.com>
2026-09-22 08:35:42 +00:00
0a770d9173 feat(kubernetes): support HA gateway rebalancing (#1868)
* feat(kubernetes): support HA gateway rebalancing

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* perf(server): cache peer connections, tokens, and owner lookups

Every forwarded relay rebuilt its setup from scratch: an owner lookup, a
blocking read of the peer token, a TLS connect to the owning replica, and
a TokenReview plus Pod GET on the receiving side. Sandbox service routing
does this per HTTP request, so the apiserver calls scaled with traffic.

Cache all of it on ServerState:

- peer channels pooled per endpoint, so relays multiplex over one
  connection instead of redialing
- peer tokens keyed by SHA-256, expiring at min(ttl, token exp) so a hit
  cannot accept an expired token
- owner records for 3s against a 45s ownership TTL, still freshness
  checked before use

Entries are evicted when a relay fails. Also raise HTTP/2
max_concurrent_streams to 1024, since pooling funnels every relay between
two replicas onto one connection and hyper's default of 200 sits below
the 256 pending-relay budget.

Signed-off-by: divesh <dgude@nvidia.com>

* perf(server): pool upstream connections for sandbox services

Each HTTP request to a sandbox service opened its own supervisor relay,
paying a new TCP connection and HTTP/1 handshake every time. Worse, it
counted against the 32 in-flight relay cap, so a service handling more
than 32 concurrent requests failed outright.

Pool idle upstreams per endpoint and port, up to 8 each for 15s. Reuse is
safe because the pool only returns a connection hyper reports as ready,
and HTTP/1 cannot start a request until the previous body has drained.
Upgrades are never pooled since they take the connection over, and a
failed send evicts that endpoint. Pruning is bounded per key, with the
full sweep limited to once per 30s.

Signed-off-by: divesh <dgude@nvidia.com>

* fix(server): address HA gateway review findings (#3449)

- Let a gateway own supervisor sessions without a peer endpoint. Requiring
  one whenever the store is PostgreSQL broke every single-instance
  PostgreSQL deployment, because no sandbox supervisor could connect.
  A cross-replica request to an owner that advertises no endpoint now fails
  immediately naming the cause, instead of retrying until the wait timeout.
- Close a supervisor session on heartbeat only when another replica owns it,
  or after renewals fail for the ownership TTL. A database error no longer
  drops every session heartbeating during an outage.
- Clamp owner record ages at zero so a skewed or corrupt stored timestamp
  cannot produce a negative age.
- Bound the cross-object advisory lock with a lock timeout, so a stuck holder
  fails instead of blocking every mutation in the fleet.
- Refuse to start when a peer endpoint is configured on a multi-replica
  backend but peer authentication is unavailable, and warn when a
  multi-replica backend has no peer endpoint at all.
- Reject a plaintext peer endpoint when the gateway serves TLS.
- Skip the sandbox watch poller on single-replica backends, where the local
  update bus already sees every write.
- Rate-limit the peer owner cache sweep so an insert no longer scans the
  whole map under the lock.
- Retry GET and HEAD on a pooled upstream the sandbox closed, instead of
  returning 502, and drop an emptied endpoint from the pool right away.
- Document the gateway peer environment variables and the post-rollout
  ownership skew operators should expect.

Signed-off-by: divesh <dgude@nvidia.com>

* fix(server): harden HA supervisor ownership

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Signed-off-by: divesh <dgude@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: divesh <dgude@nvidia.com>
Co-authored-by: Divesh Chowdary <47188680+FrostGod@users.noreply.github.com>
2026-09-21 21:22:12 +00:00
Evan Lezar 251f77e2b8 ci(security): gate tagged releases on scans (#3523)
Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-09-21 18:51:37 +00:00
Simon ScattonandEvan Lezar 65eb9167d1 test(tmachine): add Debian installer profile (#3461)
* test(tmachine): add Debian installer profile

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* refactor(ci): centralize conformance matrix

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* test(ci): prefer packaged conformance artifacts

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* test(tmachine): use packaged Debian gateway service

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* test(tmachine): isolate Debian qualification config

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Simon Scatton <sscatton@nvidia.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>
Co-authored-by: Evan Lezar <elezar@nvidia.com>
2026-09-21 12:17:11 +02:00
Drew Newberry 17ce738bfb fix(ci)!: remove gateway callback listener dependency (#3365)
* fix(ci): repair post-merge release canary

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(packaging): bootstrap canary runtime prerequisites

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci(canary): collect macOS VM diagnostics

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci(canary): pin libkrun-compatible macOS runner

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci(canary): limit macOS smoke test to package startup

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(compute)!: remove gateway callback listeners

Run Docker supervisors on host networking so they use the operator-configured primary gateway endpoint. Remove the unused compute-driver callback listener negotiation and listener-scoped routing machinery.

BREAKING CHANGE: The ComputeDriver API no longer exposes GetGatewayListenerRequirements or GatewayListenerRequirement. External drivers must regenerate bindings and connect supervisors to the configured primary gateway endpoint.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(docker): use sandbox runtime image in launcher

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(podman): exercise production endpoint selection

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(e2e): route supervisors to reachable gateways

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(e2e): align Podman endpoint fixtures

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): preserve host aliases for supervisors

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): align sandbox host gateway pin

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(e2e): address Docker fixtures by bridge IP

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(e2e): serialize sandbox lifecycle cases

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(e2e): host Docker TCP fixture with gateway

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(e2e): use loopback for host-network supervisor

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-18 20:55:55 +00:00
Simon Scatton 8bd3dcc565 refactor(tmachine): separate installers from environments (#3419)
Signed-off-by: Simon Scatton <sscatton@nvidia.com>
2026-09-18 12:59:20 +00:00
Evan Lezar 2263685cf3 test(tmachine): migrate Keycloak provider refresh coverage (#3404)
* test(tmachine): add Keycloak provider refresh suite

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* refactor(tmachine): share container runtime detection

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* ci(tmachine): run feature suites in GitHub Actions

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* ci(tmachine): run conformance with Podman tests

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* ci(tmachine): cover provider refresh with Podman

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* ci(integration): split input preparation from runners

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-09-18 13:56:07 +02:00
Drew NewberryandPiotr Mlocek 8b77925eb7 feat(installer): support prerelease installations (#3364)
* feat(installer): support prerelease installations

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(readme): clarify 0.1.0 production readiness

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(readme): refine production readiness message

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(release): publish rolling prerelease channels

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(release): keep prereleases out of GitHub releases

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(readme): refine 0.1.0 notice

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(readme): expand 0.1.0 notice

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(readme): simplify 0.1.0 notice

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(installer): simplify prerelease alias

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(installer): require successful prerelease runs

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(release): reuse exact prerelease artifacts

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(installer): include prover in prerelease bundle

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci: remove prerelease post-publish validation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(fern): sync announcement configuration

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs: remove stale prerelease canary claim

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(fern): defer announcement synchronization

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(fern): simplify release announcement

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Co-authored-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-09-18 00:53:36 +00:00
Piotr Mlocek c31ff5743f fix(ci): use renamed conformance suite in release dev (#3425)
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-09-17 19:52:50 +00:00
Piotr Mlocek 769273f096 feat(middleware): add a hook to inspect HTTP responses (#3074)
* feat(middleware): implement HTTP response processing

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): allow one-byte response stream units

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* feat(examples): separate content guard from middleware protocol demos

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): address HTTP response review findings

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* refactor(examples): defer protocol demo to a separate PR

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* refactor(middleware): keep response body timeout local

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): honor fail-open for unrepresentable responses

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(middleware): trim runtime docs and extract troubleshooting reference

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* chore(middleware): split build fix and simplify test and skill guidance

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* refactor(middleware): isolate HTTP response processing

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): hide response credential headers

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): end invalid preflight streams

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(middleware): remove hard-wrapped prose

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): use protobuf request timeouts

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

---------

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-09-17 18:54:38 +00:00
Johnny Greco 58b5f8f976 feat(prover): add standalone policy boundary checker (#3289)
* feat(prover): add standalone policy maximum checker

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* refactor(prover): simplify check scope schema

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(prover): align containment and cancellation with runtime

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(prover): stabilize containment checks in CI

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* test(prover): avoid solver in fast-path guard test

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(prover): align string containment with runtime

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(prover): reject ambiguous z3 string escapes

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(prover): align containment with runtime boundaries

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(prover-cli): harden cancellation and invalid input

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* feat(packaging): install policy prover with OpenShell

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(prover): clarify installation and check results

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(build): describe prover distribution directly

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* refactor(prover): rename maximum policy to boundary

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* refactor(prover): localize fail-closed validation

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* test(prover): cover fail-closed CLI surfaces

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* test(prover): allow CI load for REST solver proof

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(prover): use canonical policy schema for containment

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(prover): preserve uncertainty for runtime binary globs

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(prover): bound policy validation work

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(prover): document validation resource limits

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(prover): make containment API extensible

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(prover): define containment API contract

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(ci): integrate prover with consolidated builds

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(ci): declare release packaging dependency

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

---------

Signed-off-by: Johnny Greco <jogreco@nvidia.com>
2026-09-17 17:30:03 +00:00
Simon Scatton 0a0a563dd3 ci(conformance): run tmachine suites in release dev (#3382)
Signed-off-by: Simon Scatton <sscatton@nvidia.com>
2026-09-17 17:13:56 +02:00
Simon Scatton fc03bffead ci: consolidate multi-platform image builds (#3408)
* ci: remove release smoke tests

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* ci: consolidate image builds

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

---------

Signed-off-by: Simon Scatton <sscatton@nvidia.com>
2026-09-17 13:46:31 +00:00
Simon Scatton 44afc321ca fix(ci): restore release tag push authentication (#3410)
Signed-off-by: Simon Scatton <sscatton@nvidia.com>
2026-09-17 13:46:22 +00:00
Simon Scatton 7d5b2e4fb0 ci: consolidate release binary builds (#3405)
Signed-off-by: Simon Scatton <sscatton@nvidia.com>
2026-09-17 12:09:42 +00:00
krishicks 12a7a35910 chore(tools): upgrade mise to 2026.9.9 (#3385)
Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-09-16 09:04:50 -07:00
Simon Scatton 4a3f2678d9 ci(release): advance seeded prereleases daily at Zurich time (#3239)
* ci(release): advance seeded prereleases daily at Zurich time

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* ci(release): schedule prereleases on weekdays

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

---------

Signed-off-by: Simon Scatton <sscatton@nvidia.com>
2026-09-16 08:48:18 +00:00
Prekshi Vyas 314c73343c fix(ci): restore prebuilt Z3 on Windows (#3353)
Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>
2026-09-16 01:41:53 +00:00
Drew Newberry c1f2e7189f feat(isolation): implement the RFC 0012 sandbox architecture (#2942)
* feat(isolation): add RFC 0012 backend contract

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>

* refactor(isolation): name the interface crate explicitly

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): expose trusted host gateway

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(agents): inventory the MXC driver

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): add mediated DNS transport

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): tighten interface error and digest contracts

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(isolation): remove unrelated driver inventory

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): define capability-free launch contract

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): seal confirmed boundary state

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): validate confirmation for external backend implementations

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): clarify mediated DNS identity

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(isolation): generalize loopback connector

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(isolation): unify typed network mediation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): bind launches to sandbox sessions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(mxc): initialize extended sandbox status

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): add boundary protocol and Linux primitives

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): harden signals and separate process status from transport

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): validate remote confirmation through public contract

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): validate wire state and propagate snapshot failures

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(isolation): import owned agent specification explicitly

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(isolation): describe mediated DNS channel

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): bound mediation attach without nested retries

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(isolation): generalize loopback protocol

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): add transport-neutral session authentication

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(isolation): separate sandbox backend protocol

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): harden runtime boundary controls

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): add terminal boundary operation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): split supervisor and sandbox runtimes

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): harden boundary isolation and lifecycle ownership

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): reject private root redirects and adopt typed errors

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): preserve accept thread ownership on musl

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(sandbox): isolate credential probes from filtered threads

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): return retained exec exit status to independent waiters

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): bound network mediation and preserve socket authorization

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): bound control admission and retire stale mediation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci(e2e): select migrated drivers per stack layer

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(sandbox): implement loopback connector

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): authenticate the Sandbox Protocol

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(supervisor): rotate launch-scoped authentication

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(sandbox): consume dedicated backend crate

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(sandbox): align topology session fixture

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): align projected bootstrap bundle

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(auth): validate refreshed credentials before rotation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): fail closed across supervisor disconnects

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): repair rebased sandbox CI

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* build(runtime): publish separate sandbox and supervisor images

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(config): configure the sandbox runtime image

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(ci): validate sandbox binary linkage

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(isolation): use backend and runtime terminology

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(sandbox): use a scratch runtime image

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(ci): refresh schema and dependency policy

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): bind reconnects to supervisor process

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs: align runtime split operational guidance

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* chore(security): document Kubernetes runtime RBAC

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): enforce runtime lifecycle invariants

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(compute): identify sandbox start generations

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(server): restore sandbox launch sessions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): support authenticated runtime replacement

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(auth): bind sandbox session successors

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(auth): retry pending sandbox successors

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(vm): run the supervisor outside the guest workload

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(vm): use sandbox backend protocol

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): repair rebase integration

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): use unified build toolchain

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(vm): use sandbox runtime terminology

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(vm): own guest network bootstrap

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): expose guest init version

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): select native supervisor artifacts

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): guard guest init Linux symbols

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): scope Linux test imports

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): avoid guest interface casts

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): reconcile admitted sandbox identity

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): share resolved sandbox identity

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): surface host supervisor failures

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): include guest logs on supervisor exit

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): rotate and clean runtime generations

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): make sandbox starts generation-aware

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): rotate restored sandbox sessions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(vm): keep shared paths in the base layer

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): bind sandbox session lineage

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(docker): isolate workloads behind the host supervisor

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(docker): rotate launch-scoped authentication

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(docker): use sandbox backend protocol

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(docker): use host networking for supervisor

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): preserve host gateway alias resolution

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(docker): use separate sandbox and supervisor images

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): restore startup validation after rebase

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(docker): name the sandbox runtime directly

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): narrow supervisor CA runtime storage

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): close companion isolation gaps

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(docker): align mediated network expectations

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(docker): exercise mediated network paths

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): attach supervisor to managed network

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): defer supervisor recovery until gateway is ready

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): make sandbox starts generation-aware

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): rotate restored sandbox sessions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): preserve workloads during session rotation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(docker): remove unrelated configuration RFC changes

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): bind sandbox session lineage

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(kubernetes): add proxy-pod isolation topology

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(kubernetes): use sandbox backend protocol

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): use stable sandbox service authority

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(kubernetes): split sandbox and supervisor images

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): adapt proxy pods to current runtime APIs

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(kubernetes): describe the single runtime placement

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(kubernetes): simplify sandbox orchestration

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): validate deployment prerequisites

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): update Trivy Helm profile inventory

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(kubernetes): update Trivy scan inventory count

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): reuse preloaded runtime images in e2e

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): type and clean runtime resources

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): make sandbox restarts recoverable

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): rotate restored sandbox sessions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): preserve supervisor egress

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): bind sandbox session lineage

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(podman): adopt isolated sandbox and supervisor containers

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): stage bootstrap archives at named volume destinations

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(podman): rotate launch-scoped authentication

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(podman): use sandbox backend protocol

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(podman): use host networking for supervisor

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(podman): split sandbox and supervisor images

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): repair rebase integration

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(podman): name the sandbox runtime directly

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): provision supervisor CA runtime storage

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): address isolation review findings

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): inspect Debian supervisor provenance

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): use libpod-compatible tmpfs options

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): bind verified sandbox runtime binary

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): provide external driver data directory

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): start sandbox before joining user namespace

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): separate supervisor user namespace

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): make sandbox starts generation-aware

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): rotate restored sandbox sessions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): bind sandbox session lineage

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* perf(isolation): add TCP and DNS benchmark harnesses

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(perf): align benchmark timing and supported protocols

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(perf): report TCP benchmark metrics accurately

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(perf): cancel failed worker startup

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): build matching local supervisor image

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): make local sandbox smoke test runnable

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): wire local sandbox runtime image

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): narrow sandbox service RBAC

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(ci): validate split runtime artifacts

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): harden runtime session handling

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(supervisor): add standalone network proxy role

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(rfc): remove implementation companion notes

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(vm): standardize runtime release name

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): pin renamed runtime artifacts

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(auth): persist sandbox runtime identity

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(runtime): restore branch validation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): reconcile main after rebase

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(network): close unframed HTTP 1.0 responses

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* chore(isolation): preserve upstream OCSF updates

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(security): close credential and TLS replay paths

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(auth): make sandbox refresh retries idempotent

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-16 00:49:14 +00:00