Commit Graph
1497 Commits
Author SHA1 Message Date
Evan Lezar ee1528d4f0 fix(upgrade): preserve pre.4 supervisor compatibility
Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-09-23 15:57:47 +02:00
Evan Lezar 50f274fcd0 refactor(upgrade): share fence format detection
Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-09-23 15:51:14 +02:00
Evan Lezar 4c37f92a7d fix(upgrade): preserve legacy driver fence bootstraps
Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-09-23 15:51:14 +02:00
Evan Lezar 61548c714d test(upgrade): capture runtime logs on failure
Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-09-23 15:50:29 +02:00
Evan Lezar 1a589111f8 fix(storage): migrate pre.4 endpoint modes on upgrade
Rewrite legacy NetworkEndpoint string fields before protobuf decode so DEB and RPM upgrades from v0.1.0-pre.4 can load persisted sandboxes.

Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-09-23 15:50:29 +02:00
Evan Lezar 4818025be2 ci: share package upgrade source resolution
Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-09-23 15:50:29 +02:00
Evan Lezar 3ea2bd7219 ci: always build deb and rpm packages
Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-09-23 15:50:29 +02:00
Evan Lezar e09cd36d39 test(tmachine): stabilize package upgrade scenarios
Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-09-23 15:50:29 +02:00
Evan Lezar 41211de089 test(tmachine): qualify RPM package upgrades
Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-09-23 15:50:29 +02:00
Evan Lezar bedd141c06 test(ci): add label-gated upgrade qualification
Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-09-23 15:50:29 +02:00
Evan Lezar dab9744768 test(release): run Debian upgrade qualification
Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-09-23 15:50:29 +02:00
Evan Lezar f47e23f469 test(tmachine): add Debian upgrade scenario
Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-09-23 15:50:29 +02:00
Evan Lezar 907f894ebc ci(release): publish prereleases with qualification summary (#3593)
Signed-off-by: Evan Lezar <elezar@nvidia.com>
v0.1.0-pre.9
2026-09-23 13:18:30 +00:00
Florent BENOIT d3480d2a7e fix(sandbox-backend): use String for CA cert/bundle in boundary protocol (#3456)
CA certificates are PEM-encoded text. Using String instead of Vec<u8>
avoids base64 overhead in the JSON wire format, keeping large CA bundles
within the 1 MiB control frame limit without needing to raise it.

Signed-off-by: Florent Benoit <fbenoit@redhat.com>
v0.1.0-pre.8
2026-09-23 09:15:34 +00:00
Drew Newberry 069ae6bd96 fix(vm): scope GPU filesystem enrichment to assigned workloads (#3580)
* fix(vm): scope GPU filesystem enrichment to assigned workloads

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(supervisor): clarify proxy baseline enrichment

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-23 07:14:19 +00:00
Divesh Chowdary d02ebe2c4b fix(kubernetes): prevent false sandbox suspension (#3567)
Signed-off-by: divesh <dgude@nvidia.com>
2026-09-22 21:57:33 -07:00
Drew Newberry c8b20bf0a2 ci(windows): seed caches on windows branch (#3576)
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-22 21:12:05 -07:00
Drew Newberry b671171091 fix(sandbox): qualify task memory against workload child (#3574)
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-23 03:36:56 +00:00
Evan Lezar df88bedb31 fix(podman): support rootless user namespace configurations (#3527)
* test(podman): cover user namespace configurations

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(podman): support keep-id runtime groups

Signed-off-by: Evan Lezar <elezar@nvidia.com>

refactor(podman): generalize keep-id group handling

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* ci(podman): run driver integration tests

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-09-23 00:30:29 +00:00
Drew Newberry 84960e70a3 fix(kubernetes): scope resource admission RBAC (#3571)
* fix(kubernetes): scope resource admission RBAC

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(helm): gate PVC admission reads

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-23 00:02:10 +00:00
Drew Newberry feff897968 fix(sandbox): await SFTP writes before acknowledging (#3568)
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-22 23:04:15 +00:00
Evan Lezar 49df4d7ca7 fix(server): drain supervisor ownership cleanup on shutdown (#3547)
Closes #3546

Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-09-22 22:29:54 +00:00
Mrunal Patel bdffa102c3 feat(api): add durable exec launch admission (#3324)
* feat(api): add durable exec launch admission

Fence duplicate exec launches with keyed durable admission and producer-owned terminal completion. Keep uncertain launches unresolved and never replay output or interactive input.

Part of #3051 (phase 4a).

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(api): fence exec identity across authorization lookups

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

---------

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
2026-09-22 21:52:19 +00:00
Drew Newberry 1e34e8c576 fix(drivers): require admission labels for external resources (#3538)
* fix(drivers): require admission labels for external resources

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(drivers): address resource admission review findings

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(core): reserve driver-owned admission labels

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(core): clarify workspace admission label

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(drivers): clarify resource admission failures

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(e2e): configure resource admission fixtures

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): retry forbidden admission lookups

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(e2e): preserve external driver admission defaults

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-22 21:22:30 +00:00
krishicks f8002d19ad fix(e2e): repair the credential driver test (#3565)
The credential driver e2e has not passed end to end, and the disabled
kubernetes-credential-drivers CI lane hid three problems.

The test broke when the JSON format of provider list changed.
Continuation-token pagination (#3249) changed
provider list --output json from a bare array of providers to an object
with next_page_token and a providers array. The test still parsed the
output as an array, so it failed before checking either storage
backend. Read the providers array from the new object instead.

Its sandbox name was about 58 characters, but sandbox names are
DNS-routable and limited to 19, so sandbox creation was rejected. Build
a short unique name instead.

The sandbox guard deletes its sandbox from a detached thread on drop, so
the test deleted the provider while the sandbox still existed. The
gateway rejects deleting a provider that is attached to a sandbox, the
test ignored that error, and the credential Secret remained. Delete the
sandbox explicitly before returning from the sandbox check.

Signed-off-by: Kris Hicks <khicks@nvidia.com>
v0.1.0-pre.7
2026-09-22 20:55:04 +00:00
Akram Ben Aissi 293fab75d4 fix(sandbox): support kernels < 5.19 via seccomp WAIT_KILLABLE_RECV fallback (#3420)
* fix(sandbox): fall back to plain seccomp listener when WAIT_KILLABLE_RECV is unavailable

SECCOMP_FILTER_FLAG_WAIT_KILLABLE_RECV was added in Linux 5.19. On older
kernels (for example RHEL 9.x / 5.14 nodes such as RHCOS on OpenShift) the
flag is rejected with EINVAL, which made the capability-free sandbox fail to
start during the notification probe with "notification launcher disappeared".

Install the notification listener with WAIT_KILLABLE_RECV when the kernel
supports it and fall back to a plain NEW_LISTENER on EINVAL. The fallback
listener records wait_killable_recv = false: its notification receive is
uninterruptible, but the sandbox is otherwise fully functional.

Signed-off-by: Akram <akram.benaissi@gmail.com>

* fix(sandbox): address review — legacy read-only listener mode + cancellation invariant

Follow-up to the PR review (GATOR-de00bfcc-01 / mrunalp): make the < 5.19
fallback cancellation-safe instead of racing broker writes.

- Record an explicit ListenerMode (Killable vs LegacyReadOnly); add
  writes_disabled()/mode() and emit the selected mode in qualification output
  (seccomp_listener_mode).
- Centralize task-memory output writes behind NotificationListener::
  write_task_output; in LegacyReadOnly mode getpeername, accept/accept4 with a
  non-null address, and sendmmsg length write-backs fail closed with EOPNOTSUPP.
  accept with a null address, socket/connect/bind/listen/sendto/sendmsg keep
  working (copied inputs, scalar responses, atomic ADDFD_SEND).
- Enforce the launch invariant `cancellation || task_memory_writes_disabled` in
  SandboxConfirmEvidence::validate() rather than dropping cancellation
  unconditionally; add task_memory_writes_disabled to SeccompEvidence.
- Add per-path fail-closed tests (write_task_output, write_socket_addr, a real
  plain listener installed on a modern kernel) and confirmation-invariant tests.
- Correct the flag-semantics comments and document both modes plus the reduced
  legacy syscall compatibility in architecture/sandbox.md.

Signed-off-by: Akram <akram.benaissi@gmail.com>

* docs: document legacy read-only sandbox mode on kernels before 5.19

Kernels without SECCOMP_FILTER_FLAG_WAIT_KILLABLE_RECV (< 5.19, e.g. RHEL 9.x /
RHCOS 5.14) run the sandbox in a legacy read-only cancellation mode where the
broker fails closed with EOPNOTSUPP on the mediated operations that write results
back into workload memory (getpeername, accept/accept4 with a non-null address,
sendmmsg length write-backs). Document this observable behavior and its syscall
limitations in the public Fern docs: the support-matrix kernel requirements and
the OpenShift runtime guidance.

Signed-off-by: Akram <akram.benaissi@gmail.com>

* style(sandbox): satisfy rustfmt and clippy doc_markdown

Match the pinned rustfmt (Rust 1.95.0) line-wrapping for the write_task_output
call, and backtick `legacy_read_only` in the qualification-report doc comment
so clippy::doc_markdown (-D warnings) passes.

Signed-off-by: Akram <akram.benaissi@gmail.com>

* fix(sandbox): migrate task_memory_writes_disabled into backend protocol

Add the `task_memory_writes_disabled` field to `SeccompEvidence` in the
backend protocol contract and relax the validation from requiring
`cancellation` to accepting `cancellation || task_memory_writes_disabled`.

This completes the rebase migration missed by the isolation-interface
refactor: the sandbox reports this field but the backend struct lacked
it, and the validator rejected every pre-5.19 legacy listener.

Signed-off-by: Akram <akram.benaissi@gmail.com>

---------

Signed-off-by: Akram <akram.benaissi@gmail.com>
2026-09-22 20:37:19 +00:00
Piotr Mlocek 5a81d2b37b fix(network): bound chunked relay memory (#3537)
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-09-22 20:03:52 +00:00
Artem LytvynandJohn Myers 470a34635d fix(api): make WatchSandbox loss-aware and resumable (#3209)
* fix(api): emit warning on WatchSandbox broadcast lag instead of terminating

Broadcast lag on the status, log, and platform receivers was converted to a RESOURCE_EXHAUSTED status that terminated the whole watch stream. Lag is recoverable: the receiver resumes at the oldest surviving message. Emit a SandboxStreamWarning and continue streaming instead; keep terminating on Closed. Add helpers and unit tests covering the warning payload and receiver recovery after lag.

Partially addresses #3055 (cursor/resume follow up separately).

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* refactor(server): group per-sandbox log bus state and stamp sequence numbers

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* feat(proto): add resume cursor fields to sandbox watch API

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* feat(server): stamp watch cursors from a shared per-sandbox sequence

Allocate cursors from a single SeqAllocator shared by the log and
platform event buses, so a sandbox's merged watch stream carries
unique, strictly increasing cursors. A single resume_after_cursor can
then unambiguously locate a client's position across both sources.

Rewrite both publish paths to allocate the sequence, stamp
event.cursor, send, and append to the tail under one lock. This
removes the previous get_mut().expect() TOCTOU race where a concurrent
remove() between the two lock sections could panic.

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* feat(server): serve WatchSandbox resume from cursor with gap detection

Add tail_after() to the log and platform event buses, returning every
buffered event newer than a client's resume cursor. Each PerSandbox now
tracks last_trimmed_seq (the highest seq it has evicted) so a resume is
reported as an unrecoverable ResumeGap only when this bus dropped an
event the client still needs.

Judging gaps by evictions, not by the tail's oldest seq, is required
under the shared cursor space: each bus's tail is non-contiguous in the
global sequence because the other bus owns the missing seqs, so
comparing against tail.front() would flag false gaps.

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* feat(server): resume WatchSandbox from cursor across log and platform buses

Wire resume_after_cursor into the watch producer. On a non-zero cursor,
replay events strictly after it from both the log and platform buses,
merge by shared cursor, and emit in order before entering the live loop.
A trimmed range on either bus is an unrecoverable gap and terminates the
stream with OUT_OF_RANGE carrying the requested and earliest-available
cursors, distinct from recoverable lag which warns and continues.

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* test(server): cover WatchSandbox cursor resume paths

Add handler-level tests for the resumable watch stream: replay strictly
after the client cursor, merge log and platform events in shared-cursor
order, suppress duplicates when resuming at the latest cursor, and
terminate with OUT_OF_RANGE when the requested cursor has been trimmed.

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* docs(api): document WatchSandbox loss-awareness and resume

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* fix(server): deliver watch events once and harden cursor teardown

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* feat(sdk): add loss-aware resumable watch_logs to Rust SDK client

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* fix(server): keep watch cursors monotonic across teardown and restart

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* fix(server): merge live watch sources by cursor before emission

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* fix(api): bind watch cursors to a cursor space and merge tail sources

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* fix(server): revalidate the watch cursor space after collecting replay

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* test(server): update the public RPC schema fingerprint for the string cursor

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* fix(server): hold watch events above the publication watermark and emit the watch lag warning before its batch

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* test(server): synchronize the watch live-order test with the end of initialization

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* fix(sdk): use canonical sandbox name in watch_logs

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* test(sdk): guard canonical-name addressing in watch_logs

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* fix(server): fix public rpc schema

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* fix(api): reconcile watch resume rebase

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(server): bound interactive relay cleanup

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-09-22 19:45:55 +00:00
Jim Meyer 88357775ad docs(windows): align Z3 pin with z3-sys 0.13 (#3561)
z3-sys 0.13 generates bindings for the Z3 5.x series and resolves
GH_RELEASE_VERSION to 5.1.0 for the prebuilt-release path. The Windows
wrapper still pinned Z3_SYS_Z3_VERSION to 4.16.0, so a prebuilt Windows
build would fetch an archive that does not match the generated bindings.

Move the wrapper pin to 5.1.0 and update the contributor, architecture,
and skill references that still named 4.16.0.

Signed-off-by: Jim Meyer <jimeyer@nvidia.com>
2026-09-22 19:29:24 +00:00
Piotr Mlocek e367d47394 fix(sandbox): reclaim socket descriptors before exhaustion (#3532)
* fix(sandbox): reclaim socket descriptors before exhaustion

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(sandbox): separate socket and descriptor limits

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(sandbox): account for existing broker descriptors

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

---------

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-09-22 19:15:51 +00:00
Piotr Mlocek 4b1c09de28 fix(network): preserve chunked request boundaries (#3530)
* fix(network): preserve chunked request boundaries

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(network): bound chunk framing amplification

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(supervisor): box sandbox runtime future

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(network): buffer chunked relay read-ahead

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(network): flush completed chunks promptly

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

---------

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-09-22 19:04:59 +00:00
Yuedong Wu 718dba3430 fix(policy)!: reject removed tls endpoint values (#3414)
Signed-off-by: Yuedong Wu <dwcn22@outlook.com>
2026-09-22 17:58:19 +00:00
Philippe Martin 35e0a68e4a feat(kubernetes): support corporate proxy CA bundle (#3447)
* feat(kubernetes): support corporate proxy CA bundle

The Kubernetes driver had no way to supply a CA bundle for the corporate
egress proxy, so an `https://` proxy with a private CA, or a TLS-intercepting
proxy, could not be used. Podman and VM already expose `proxy_ca_bundle`.

Add `proxy_ca_bundle` to `[openshell.drivers.kubernetes]` as a path the
gateway Pod reads. The gateway stages the PEM into the existing
per-generation supervisor bootstrap Secret and passes
`--upstream-proxy-ca-bundle` on the supervisor argv. That Secret is already
immutable, owner-referenced and garbage-collected, and its volume mounts
every key at /.openshell/supervisor with no items filter, so this needs no
new object kind, volume, mount, or RBAC verb, and works in shared, managed
and operator workspace modes.

The bundle is deliberately read from the gateway's filesystem rather than
referenced as an object in the sandbox namespace. It becomes a trust anchor
for every upstream the sandbox reaches, so it must stay in the gateway's
trust domain; the immutable staging Secret also keeps the anchor from
changing underneath a running sandbox.

Bound the staged bundle at 256 KiB. The shared reader's limit is exactly the
apiserver's own Secret limit and the bootstrap Secret carries four other
keys, so a bundle between the two would pass gateway startup and then fail
every sandbox create with an opaque `data: Too long`.

Delegate the URL, no_proxy, connect_by_hostname and ca_bundle rules to the
shared validate_upstream_proxy_settings, keeping the Secret-specific
credential block local: this driver accepts an explicit
`proxy_auth_allow_insecure = false` without credentials, which the shared
rules reject. This also fixes the acknowledgement being demanded for an
`https://` proxy, where the credential travels inside the verified TLS
session. Add auth_setting_label so the inline-credential diagnostic names
the Secret keys instead of proxy_auth_file, which this driver rejects as an
unknown key.

Document that the bundle should carry only the CA that signs the proxy's
certificate, or that an intercepting proxy re-signs upstream certificates
with. Public roots already reach the sandbox through the supervisor image and
its TLS stack, and the bundle is concatenated with that system store into a
single boundary control frame, so a full merged trust bundle spends the frame
budget on duplicated roots. The frame, not the apiserver Secret limit, is the
tighter of the two ceilings in practice; raising the staging bound requires
checking it.

Closes #3443

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(helm): quote proxy CA ConfigMap references

Signed-off-by: Philippe Martin <phmartin@redhat.com>

---------

Signed-off-by: Philippe Martin <phmartin@redhat.com>
2026-09-22 17:57:40 +00:00
Evan Lezar 3107ff1f82 ci(security): stage release finding enforcement (#3552)
Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-09-22 17:21:23 +00:00
krishicks f8b1fd8b57 fix(supervisor): restore OCSF schema downgrade (#3554)
The RFC-0012 architecture migration preserved sandbox OCSF JSON enablement but
introduced a regression: ocsf_schema_version was no longer honored.

This fixes the regression so that schema downgrades occur again.

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-09-22 16:42:15 +00:00
Varsha 7139df8ca5 fix(exec): preserve output after stdin EOF and verify stream completion (#3359)
* fix(exec): preserve output after stdin EOF and verify stream completion

Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>

* test(exec): cover fair duplex progress and live SDK completion

Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>

* test(sdk): compare large exec buffers with native equality

Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>

* test(exec): use workspace-scoped sandbox name in EOF regression

Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>

---------

Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>
2026-09-22 14:11:06 +00:00
Artem Lytvyn ca4573588d fix(driver-vm): resolve lifecycle requests on sandbox_id alone (#3305)
Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>
2026-09-22 14:10:10 +00:00
50230616d5 refactor(runtime): retire Community image dependencies (#3386)
* feat(sandbox): default to official Alpine sandbox image

default_sandbox_image() now returns docker.io/library/alpine:3.22, a generic
version-qualified official image, so a fresh install no longer depends on the
community sandbox image catalog. All compute drivers (docker, podman,
kubernetes, vm) inherit this fallback.

Part of #3116.

Signed-off-by: Akram
Signed-off-by: Akram <akram.benaissi@gmail.com>

* feat(deploy): default deployment configs to the official Alpine sandbox image

Update the shared gateway default_image, Helm chart values, the standalone
Kubernetes manifest, and the dev gateway task scripts to use
docker.io/library/alpine:3.22 instead of the community base image, consistent
with default_sandbox_image(). GPU e2e image-build base is left unchanged (CUDA
needs a glibc base).

Part of #3116.

Signed-off-by: Akram
Signed-off-by: Akram <akram.benaissi@gmail.com>

* feat(driver): default to numeric non-root identity for USER-less images

With the default sandbox image now Alpine, images that declare no OCI USER
must start instead of being rejected. When the image declares no USER and
the policy requests none, the Podman and Docker drivers now supply a numeric
non-root identity (DEFAULT_SANDBOX_UID/GID = 1000) instead of rejecting,
matching the numeric-identity behavior of the Kubernetes and VM drivers. The
supervisor's resolved-identity path runs the sandbox as a synthesized
non-root account without the account existing in the image. Images that
declare a USER keep the OCI resolution path unchanged.

Part of #3116.

Signed-off-by: Akram <akram.benaissi@gmail.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(conformance): use Alpine workload image

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* refactor(policy): drop community image /app path from default policy

The restrictive default policy granted read-only access to /app, a directory
that only existed in the community base image. A generic Alpine default has no
/app, so remove it. Landlock best-effort already ignores absent paths; this
just stops advertising a community-specific layout in the default.

Part of #3116.

Signed-off-by: Akram
Signed-off-by: Akram <akram.benaissi@gmail.com>

* docs(config): document Alpine default images

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(podman): report early sandbox termination

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(podman): initialize rootless workspace ownership

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(sandbox): qualify NVIDIA Ubuntu default

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): initialize rootful default workspace

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(sftp): add native sandbox adapter

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sftp): gate runtime helper support to Linux

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sftp): support standard OpenSSH file operations

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sftp): harden rename and special file handling

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(runtime): remove community image dependencies

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(e2e): build provider readiness tool fixture

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(e2e): use a dedicated Noble fixture for Docker tests

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Akram
Signed-off-by: Akram <akram.benaissi@gmail.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Evan Lezar <elezar@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
2026-09-22 14:43:51 +02:00
Simon ScattonandEvan Lezar 24706c175b ci(rust): parallelize branch checks (#3462)
* ci(rust): parallelize branch checks

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* ci(rust): acknowledge trusted sccache environment

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* ci(rust): remove ineffective sccache backend

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* ci(rust): retain dependency boundary checks

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Simon Scatton <sscatton@nvidia.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>
Co-authored-by: Evan Lezar <elezar@nvidia.com>
v0.1.0-pre.6
2026-09-22 08:35:42 +00:00
Evan Lezar 551a81c729 test(tmachine): add interactive shell testsuite (#3522)
* test(tmachine): add interactive shell testsuite

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(tmachine): add no-install profile

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* docs(tmachine): document interactive shell usage

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-09-22 07:01:56 +00:00
John T. MyersandDrew Newberry c8a4ff5f19 fix(kubernetes): bind bootstrap to runtime identity (#3531)
* fix(kubernetes): bind bootstrap to runtime identity

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(compute): compensate runtime binding failures

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(compute): clean up backend on store failure

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(compute): merge runtime binding after start

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(auth): bind restarted sandbox sessions

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* refactor(compute): fold runtime binding into authentication

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
2026-09-22 05:08:37 +00:00
Emilien Macchi 99ed6a9df0 chore(build): remove stale static-supervisor leftovers (#3520)
The opt-in glibc-static supervisor variant from #2682 is gone: the Nix
release builds (#2977) removed it from CI and the supervisor Dockerfile,
and the RFC 0012 sandbox split (#2942) removed it from the staging script.
The split also moved the binary that runs inside workload images to
openshell-sandbox; the supervisor now runs from its own image and is
dynamically linked. A few places still describe the old model:

- skills/debug-openshell-cluster told operators to check that
  supervisor_image contains a static /openshell-supervisor. Ask instead for
  a supervisor from the matching release whose loader and shared libraries
  are available inside that image, and give a `--version` check that shows
  both. The static requirement for /openshell-sandbox is unchanged.
- verify-static-binary.sh justified its check with the supervisor and the
  glibc-static variant. Describe the property it enforces instead, with
  the musl sandbox runtime as the main example.
- stage-prebuilt-binaries.sh kept an unreachable gnu-static case in
  target_triple.

No behavior change: resolve_component only ever selects gnu or musl.

Signed-off-by: Emilien Macchi <emacchi@redhat.com>
2026-09-22 00:19:37 +00:00
Drew Newberry 96c08f111b refactor(isolation)!: make confirmation backend-neutral (#3366)
* refactor(isolation): make confirmation backend-neutral

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): validate confirmation evidence at host boundary

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(isolation): keep fence evidence driver-owned

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation)!: validate explicit fence projections

Require each compute driver to map its native evidence to individual outer-fence guarantees, and reject incomplete projections before a boundary becomes ready. Exercise the assembled remote confirmation path for invalid audit, property, generation, and digest evidence.

BREAKING CHANGE: BoundaryConfig and SandboxRuntimeDescriptor use outer_fence projections rather than the earlier driver_fence representation. State written by earlier builds cannot be decoded; operators must stop and recreate affected sandboxes after upgrading.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(isolation): clarify outer fence ownership

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): update confirmation test fixtures

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
vm-runtime
2026-09-21 15:27:16 -07:00
0a770d9173 feat(kubernetes): support HA gateway rebalancing (#1868)
* feat(kubernetes): support HA gateway rebalancing

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* perf(server): cache peer connections, tokens, and owner lookups

Every forwarded relay rebuilt its setup from scratch: an owner lookup, a
blocking read of the peer token, a TLS connect to the owning replica, and
a TokenReview plus Pod GET on the receiving side. Sandbox service routing
does this per HTTP request, so the apiserver calls scaled with traffic.

Cache all of it on ServerState:

- peer channels pooled per endpoint, so relays multiplex over one
  connection instead of redialing
- peer tokens keyed by SHA-256, expiring at min(ttl, token exp) so a hit
  cannot accept an expired token
- owner records for 3s against a 45s ownership TTL, still freshness
  checked before use

Entries are evicted when a relay fails. Also raise HTTP/2
max_concurrent_streams to 1024, since pooling funnels every relay between
two replicas onto one connection and hyper's default of 200 sits below
the 256 pending-relay budget.

Signed-off-by: divesh <dgude@nvidia.com>

* perf(server): pool upstream connections for sandbox services

Each HTTP request to a sandbox service opened its own supervisor relay,
paying a new TCP connection and HTTP/1 handshake every time. Worse, it
counted against the 32 in-flight relay cap, so a service handling more
than 32 concurrent requests failed outright.

Pool idle upstreams per endpoint and port, up to 8 each for 15s. Reuse is
safe because the pool only returns a connection hyper reports as ready,
and HTTP/1 cannot start a request until the previous body has drained.
Upgrades are never pooled since they take the connection over, and a
failed send evicts that endpoint. Pruning is bounded per key, with the
full sweep limited to once per 30s.

Signed-off-by: divesh <dgude@nvidia.com>

* fix(server): address HA gateway review findings (#3449)

- Let a gateway own supervisor sessions without a peer endpoint. Requiring
  one whenever the store is PostgreSQL broke every single-instance
  PostgreSQL deployment, because no sandbox supervisor could connect.
  A cross-replica request to an owner that advertises no endpoint now fails
  immediately naming the cause, instead of retrying until the wait timeout.
- Close a supervisor session on heartbeat only when another replica owns it,
  or after renewals fail for the ownership TTL. A database error no longer
  drops every session heartbeating during an outage.
- Clamp owner record ages at zero so a skewed or corrupt stored timestamp
  cannot produce a negative age.
- Bound the cross-object advisory lock with a lock timeout, so a stuck holder
  fails instead of blocking every mutation in the fleet.
- Refuse to start when a peer endpoint is configured on a multi-replica
  backend but peer authentication is unavailable, and warn when a
  multi-replica backend has no peer endpoint at all.
- Reject a plaintext peer endpoint when the gateway serves TLS.
- Skip the sandbox watch poller on single-replica backends, where the local
  update bus already sees every write.
- Rate-limit the peer owner cache sweep so an insert no longer scans the
  whole map under the lock.
- Retry GET and HEAD on a pooled upstream the sandbox closed, instead of
  returning 502, and drop an emptied endpoint from the pool right away.
- Document the gateway peer environment variables and the post-rollout
  ownership skew operators should expect.

Signed-off-by: divesh <dgude@nvidia.com>

* fix(server): harden HA supervisor ownership

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Signed-off-by: divesh <dgude@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: divesh <dgude@nvidia.com>
Co-authored-by: Divesh Chowdary <47188680+FrostGod@users.noreply.github.com>
2026-09-21 21:22:12 +00:00
Emilien Macchiands2cube cb6e88acb7 fix(vm): unpack registry images correctly and validate prepared disks (#3524)
* fix(vm): unpack registry images correctly and validate prepared disks

The registry image-prep path expected `umoci raw unpack` to produce a
bundle-style rootfs/ subdirectory, but it extracts the image filesystem
directly into the target. Every registry prep therefore failed after a
successful unpack. Guest init exit codes do not survive the libkrun
boundary, so the failure looked like success and the broken disk was
cached, making every later sandbox for that image fail with "prepared
image disk missing /image-rootfs".

VM E2E started hitting this after the bootstrap image moved to
nvcr.io/nvidia/base/ubuntu:24.04: `--from base` no longer matches the
bootstrap image, so it now goes through registry prep.

- Accept umoci's direct extraction layout in the guest prep script.
- Build the image rootfs under a partial directory and rename it to
  /image-rootfs only after every prep step succeeds.
- Check the prepared disk for /image-rootfs before caching it. On
  failure, leave the cache untouched and report the prep console tail.
- Size the prep disk to hold the payload and the unpacked rootfs at the
  same time. The community base image needs 1.40 GB + 3.32 GB, which
  did not fit in the old payload*3 + 512 MiB.

Fixes #2358

Co-authored-by: s2cube <26961336+s2cube@users.noreply.github.com>
Signed-off-by: Emilien Macchi <emacchi@redhat.com>

* test(e2e): run tool-dependent VM tests from the community base image

The host_gateway_alias and vm_corporate_proxy workloads run curl and
python3. The VM driver now defaults to nvcr.io/nvidia/base/ubuntu:24.04,
which ships neither, so these tests fail in VM E2E with "command not
found". Request the community base image explicitly with `--from base`.

Docker, Podman, and Kubernetes E2E already default to that image, so
their behavior is unchanged.

Signed-off-by: Emilien Macchi <emacchi@redhat.com>

---------

Signed-off-by: Emilien Macchi <emacchi@redhat.com>
Co-authored-by: s2cube <26961336+s2cube@users.noreply.github.com>
2026-09-21 13:25:02 -07:00
Seth Jennings 2493d415c2 feat(extensions)!: normalize protocol negotiation (#3352)
* feat(extensions)!: normalize protocol negotiation

Closes #3057

Introduce a shared extension handshake, enforce protocol and capability compatibility across extension families, and expose immutable negotiated snapshots through gateway info and the Go SDK.

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(credentials): fail fast on negotiation errors

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(extensions): validate gateway handshake metadata

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(go-sdk): re-export extension kind constants

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(extensions): fail fast on credential handshake rejection

Signed-off-by: Seth Jennings <sjenning@redhat.com>

---------

Signed-off-by: Seth Jennings <sjenning@redhat.com>
2026-09-21 13:19:39 -07:00
Evan Lezar 251f77e2b8 ci(security): gate tagged releases on scans (#3523)
Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-09-21 18:51:37 +00:00
Shiju fa8f6d3949 feat(cli): promote profile commands to top level (#3258)
* feat(cli): promote profile commands to top level

Add profile discovery and management commands with shared handlers for
the existing provider entry points. List a flat catalog across scopes
and follow continuation tokens through full and short pages.

Describe metadata, credentials, endpoints, TLS inspection, and MCP access
settings while preserving complete JSON/YAML definitions. Cover parser
equivalence, scope forwarding, pagination, and inspection settings with
focused unit and compiled-CLI integration tests.

Update docs, public skills, examples, and E2E command invocations.

Refs #2588

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(cli): remove redundant workspace selector qualification

Use the imported WorkspaceSelector in the provider integration helper so
the target passes Clippy with warnings denied.

Signed-off-by: Shiju <shiju@nvidia.com>

---------

Signed-off-by: Shiju <shiju@nvidia.com>
2026-09-21 17:18:55 +00:00
krishicks cbf026366d fix(ocsf): require network activity endpoints (#3355)
Previously, Network Activity could be constructed without a source or
destination endpoint, allowing connection, accept, relay, and configuration
events to violate the OCSF 1.8 endpoint constraint.

Now, NetworkActivityBuilder requires a source or destination endpoint at
compile time. Connection failures identify the workload peer or genuine
transparent destination, listener failures identify the listening endpoint,
and mediation-lane failures use Application Lifecycle rather than fabricated
network endpoints. Malformed forward requests use HTTP Activity with a
method-only request, generated 400 response, and workload peer.

Additionally, Unix relay-channel events use Base Event, policy-validation
warnings use Config State Change, and the unused bypass monitor is removed
because the current isolation architecture no longer uses it.

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-09-21 17:13:41 +00:00
Drew Newberry 484f0768fc fix(server): serialize sandbox restart authentication (#3485)
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
v0.1.0-pre.5
2026-09-21 10:23:17 +00:00