The credential driver e2e has not passed end to end, and the disabled
kubernetes-credential-drivers CI lane hid three problems.
The test broke when the JSON format of provider list changed.
Continuation-token pagination (#3249) changed
provider list --output json from a bare array of providers to an object
with next_page_token and a providers array. The test still parsed the
output as an array, so it failed before checking either storage
backend. Read the providers array from the new object instead.
Its sandbox name was about 58 characters, but sandbox names are
DNS-routable and limited to 19, so sandbox creation was rejected. Build
a short unique name instead.
The sandbox guard deletes its sandbox from a detached thread on drop, so
the test deleted the provider while the sandbox still existed. The
gateway rejects deleting a provider that is attached to a sandbox, the
test ignored that error, and the credential Secret remained. Delete the
sandbox explicitly before returning from the sandbox check.
Signed-off-by: Kris Hicks <khicks@nvidia.com>
* fix(sandbox): fall back to plain seccomp listener when WAIT_KILLABLE_RECV is unavailable
SECCOMP_FILTER_FLAG_WAIT_KILLABLE_RECV was added in Linux 5.19. On older
kernels (for example RHEL 9.x / 5.14 nodes such as RHCOS on OpenShift) the
flag is rejected with EINVAL, which made the capability-free sandbox fail to
start during the notification probe with "notification launcher disappeared".
Install the notification listener with WAIT_KILLABLE_RECV when the kernel
supports it and fall back to a plain NEW_LISTENER on EINVAL. The fallback
listener records wait_killable_recv = false: its notification receive is
uninterruptible, but the sandbox is otherwise fully functional.
Signed-off-by: Akram <akram.benaissi@gmail.com>
* fix(sandbox): address review — legacy read-only listener mode + cancellation invariant
Follow-up to the PR review (GATOR-de00bfcc-01 / mrunalp): make the < 5.19
fallback cancellation-safe instead of racing broker writes.
- Record an explicit ListenerMode (Killable vs LegacyReadOnly); add
writes_disabled()/mode() and emit the selected mode in qualification output
(seccomp_listener_mode).
- Centralize task-memory output writes behind NotificationListener::
write_task_output; in LegacyReadOnly mode getpeername, accept/accept4 with a
non-null address, and sendmmsg length write-backs fail closed with EOPNOTSUPP.
accept with a null address, socket/connect/bind/listen/sendto/sendmsg keep
working (copied inputs, scalar responses, atomic ADDFD_SEND).
- Enforce the launch invariant `cancellation || task_memory_writes_disabled` in
SandboxConfirmEvidence::validate() rather than dropping cancellation
unconditionally; add task_memory_writes_disabled to SeccompEvidence.
- Add per-path fail-closed tests (write_task_output, write_socket_addr, a real
plain listener installed on a modern kernel) and confirmation-invariant tests.
- Correct the flag-semantics comments and document both modes plus the reduced
legacy syscall compatibility in architecture/sandbox.md.
Signed-off-by: Akram <akram.benaissi@gmail.com>
* docs: document legacy read-only sandbox mode on kernels before 5.19
Kernels without SECCOMP_FILTER_FLAG_WAIT_KILLABLE_RECV (< 5.19, e.g. RHEL 9.x /
RHCOS 5.14) run the sandbox in a legacy read-only cancellation mode where the
broker fails closed with EOPNOTSUPP on the mediated operations that write results
back into workload memory (getpeername, accept/accept4 with a non-null address,
sendmmsg length write-backs). Document this observable behavior and its syscall
limitations in the public Fern docs: the support-matrix kernel requirements and
the OpenShift runtime guidance.
Signed-off-by: Akram <akram.benaissi@gmail.com>
* style(sandbox): satisfy rustfmt and clippy doc_markdown
Match the pinned rustfmt (Rust 1.95.0) line-wrapping for the write_task_output
call, and backtick `legacy_read_only` in the qualification-report doc comment
so clippy::doc_markdown (-D warnings) passes.
Signed-off-by: Akram <akram.benaissi@gmail.com>
* fix(sandbox): migrate task_memory_writes_disabled into backend protocol
Add the `task_memory_writes_disabled` field to `SeccompEvidence` in the
backend protocol contract and relax the validation from requiring
`cancellation` to accepting `cancellation || task_memory_writes_disabled`.
This completes the rebase migration missed by the isolation-interface
refactor: the sandbox reports this field but the backend struct lacked
it, and the validator rejected every pre-5.19 legacy listener.
Signed-off-by: Akram <akram.benaissi@gmail.com>
---------
Signed-off-by: Akram <akram.benaissi@gmail.com>
* fix(api): emit warning on WatchSandbox broadcast lag instead of terminating
Broadcast lag on the status, log, and platform receivers was converted to a RESOURCE_EXHAUSTED status that terminated the whole watch stream. Lag is recoverable: the receiver resumes at the oldest surviving message. Emit a SandboxStreamWarning and continue streaming instead; keep terminating on Closed. Add helpers and unit tests covering the warning payload and receiver recovery after lag.
Partially addresses #3055 (cursor/resume follow up separately).
Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>
* refactor(server): group per-sandbox log bus state and stamp sequence numbers
Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>
* feat(proto): add resume cursor fields to sandbox watch API
Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>
* feat(server): stamp watch cursors from a shared per-sandbox sequence
Allocate cursors from a single SeqAllocator shared by the log and
platform event buses, so a sandbox's merged watch stream carries
unique, strictly increasing cursors. A single resume_after_cursor can
then unambiguously locate a client's position across both sources.
Rewrite both publish paths to allocate the sequence, stamp
event.cursor, send, and append to the tail under one lock. This
removes the previous get_mut().expect() TOCTOU race where a concurrent
remove() between the two lock sections could panic.
Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>
* feat(server): serve WatchSandbox resume from cursor with gap detection
Add tail_after() to the log and platform event buses, returning every
buffered event newer than a client's resume cursor. Each PerSandbox now
tracks last_trimmed_seq (the highest seq it has evicted) so a resume is
reported as an unrecoverable ResumeGap only when this bus dropped an
event the client still needs.
Judging gaps by evictions, not by the tail's oldest seq, is required
under the shared cursor space: each bus's tail is non-contiguous in the
global sequence because the other bus owns the missing seqs, so
comparing against tail.front() would flag false gaps.
Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>
* feat(server): resume WatchSandbox from cursor across log and platform buses
Wire resume_after_cursor into the watch producer. On a non-zero cursor,
replay events strictly after it from both the log and platform buses,
merge by shared cursor, and emit in order before entering the live loop.
A trimmed range on either bus is an unrecoverable gap and terminates the
stream with OUT_OF_RANGE carrying the requested and earliest-available
cursors, distinct from recoverable lag which warns and continues.
Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>
* test(server): cover WatchSandbox cursor resume paths
Add handler-level tests for the resumable watch stream: replay strictly
after the client cursor, merge log and platform events in shared-cursor
order, suppress duplicates when resuming at the latest cursor, and
terminate with OUT_OF_RANGE when the requested cursor has been trimmed.
Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>
* docs(api): document WatchSandbox loss-awareness and resume
Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>
* fix(server): deliver watch events once and harden cursor teardown
Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>
* feat(sdk): add loss-aware resumable watch_logs to Rust SDK client
Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>
* fix(server): keep watch cursors monotonic across teardown and restart
Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>
* fix(server): merge live watch sources by cursor before emission
Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>
* fix(api): bind watch cursors to a cursor space and merge tail sources
Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>
* fix(server): revalidate the watch cursor space after collecting replay
Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>
* test(server): update the public RPC schema fingerprint for the string cursor
Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>
* fix(server): hold watch events above the publication watermark and emit the watch lag warning before its batch
Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>
* test(server): synchronize the watch live-order test with the end of initialization
Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>
* fix(sdk): use canonical sandbox name in watch_logs
Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>
* test(sdk): guard canonical-name addressing in watch_logs
Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>
* fix(server): fix public rpc schema
Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>
* fix(api): reconcile watch resume rebase
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
* fix(server): bound interactive relay cleanup
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
---------
Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <9696606+johntmyers@users.noreply.github.com>
z3-sys 0.13 generates bindings for the Z3 5.x series and resolves
GH_RELEASE_VERSION to 5.1.0 for the prebuilt-release path. The Windows
wrapper still pinned Z3_SYS_Z3_VERSION to 4.16.0, so a prebuilt Windows
build would fetch an archive that does not match the generated bindings.
Move the wrapper pin to 5.1.0 and update the contributor, architecture,
and skill references that still named 4.16.0.
Signed-off-by: Jim Meyer <jimeyer@nvidia.com>
* fix(sandbox): reclaim socket descriptors before exhaustion
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
* fix(sandbox): separate socket and descriptor limits
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
* fix(sandbox): account for existing broker descriptors
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
---------
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
* feat(kubernetes): support corporate proxy CA bundle
The Kubernetes driver had no way to supply a CA bundle for the corporate
egress proxy, so an `https://` proxy with a private CA, or a TLS-intercepting
proxy, could not be used. Podman and VM already expose `proxy_ca_bundle`.
Add `proxy_ca_bundle` to `[openshell.drivers.kubernetes]` as a path the
gateway Pod reads. The gateway stages the PEM into the existing
per-generation supervisor bootstrap Secret and passes
`--upstream-proxy-ca-bundle` on the supervisor argv. That Secret is already
immutable, owner-referenced and garbage-collected, and its volume mounts
every key at /.openshell/supervisor with no items filter, so this needs no
new object kind, volume, mount, or RBAC verb, and works in shared, managed
and operator workspace modes.
The bundle is deliberately read from the gateway's filesystem rather than
referenced as an object in the sandbox namespace. It becomes a trust anchor
for every upstream the sandbox reaches, so it must stay in the gateway's
trust domain; the immutable staging Secret also keeps the anchor from
changing underneath a running sandbox.
Bound the staged bundle at 256 KiB. The shared reader's limit is exactly the
apiserver's own Secret limit and the bootstrap Secret carries four other
keys, so a bundle between the two would pass gateway startup and then fail
every sandbox create with an opaque `data: Too long`.
Delegate the URL, no_proxy, connect_by_hostname and ca_bundle rules to the
shared validate_upstream_proxy_settings, keeping the Secret-specific
credential block local: this driver accepts an explicit
`proxy_auth_allow_insecure = false` without credentials, which the shared
rules reject. This also fixes the acknowledgement being demanded for an
`https://` proxy, where the credential travels inside the verified TLS
session. Add auth_setting_label so the inline-credential diagnostic names
the Secret keys instead of proxy_auth_file, which this driver rejects as an
unknown key.
Document that the bundle should carry only the CA that signs the proxy's
certificate, or that an intercepting proxy re-signs upstream certificates
with. Public roots already reach the sandbox through the supervisor image and
its TLS stack, and the bundle is concatenated with that system store into a
single boundary control frame, so a full merged trust bundle spends the frame
budget on duplicated roots. The frame, not the apiserver Secret limit, is the
tighter of the two ceilings in practice; raising the staging bound requires
checking it.
Closes#3443
Signed-off-by: Philippe Martin <phmartin@redhat.com>
* fix(helm): quote proxy CA ConfigMap references
Signed-off-by: Philippe Martin <phmartin@redhat.com>
---------
Signed-off-by: Philippe Martin <phmartin@redhat.com>
The RFC-0012 architecture migration preserved sandbox OCSF JSON enablement but
introduced a regression: ocsf_schema_version was no longer honored.
This fixes the regression so that schema downgrades occur again.
Signed-off-by: Kris Hicks <khicks@nvidia.com>
* feat(sandbox): default to official Alpine sandbox image
default_sandbox_image() now returns docker.io/library/alpine:3.22, a generic
version-qualified official image, so a fresh install no longer depends on the
community sandbox image catalog. All compute drivers (docker, podman,
kubernetes, vm) inherit this fallback.
Part of #3116.
Signed-off-by: Akram
Signed-off-by: Akram <akram.benaissi@gmail.com>
* feat(deploy): default deployment configs to the official Alpine sandbox image
Update the shared gateway default_image, Helm chart values, the standalone
Kubernetes manifest, and the dev gateway task scripts to use
docker.io/library/alpine:3.22 instead of the community base image, consistent
with default_sandbox_image(). GPU e2e image-build base is left unchanged (CUDA
needs a glibc base).
Part of #3116.
Signed-off-by: Akram
Signed-off-by: Akram <akram.benaissi@gmail.com>
* feat(driver): default to numeric non-root identity for USER-less images
With the default sandbox image now Alpine, images that declare no OCI USER
must start instead of being rejected. When the image declares no USER and
the policy requests none, the Podman and Docker drivers now supply a numeric
non-root identity (DEFAULT_SANDBOX_UID/GID = 1000) instead of rejecting,
matching the numeric-identity behavior of the Kubernetes and VM drivers. The
supervisor's resolved-identity path runs the sandbox as a synthesized
non-root account without the account existing in the image. Images that
declare a USER keep the OCI resolution path unchanged.
Part of #3116.
Signed-off-by: Akram <akram.benaissi@gmail.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>
* test(conformance): use Alpine workload image
Signed-off-by: Evan Lezar <elezar@nvidia.com>
* refactor(policy): drop community image /app path from default policy
The restrictive default policy granted read-only access to /app, a directory
that only existed in the community base image. A generic Alpine default has no
/app, so remove it. Landlock best-effort already ignores absent paths; this
just stops advertising a community-specific layout in the default.
Part of #3116.
Signed-off-by: Akram
Signed-off-by: Akram <akram.benaissi@gmail.com>
* docs(config): document Alpine default images
Signed-off-by: Evan Lezar <elezar@nvidia.com>
* fix(podman): report early sandbox termination
Signed-off-by: Evan Lezar <elezar@nvidia.com>
* fix(podman): initialize rootless workspace ownership
Signed-off-by: Evan Lezar <elezar@nvidia.com>
* fix(sandbox): qualify NVIDIA Ubuntu default
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* fix(podman): initialize rootful default workspace
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* feat(sftp): add native sandbox adapter
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* fix(sftp): gate runtime helper support to Linux
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* fix(sftp): support standard OpenSSH file operations
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* fix(sftp): harden rename and special file handling
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* refactor(runtime): remove community image dependencies
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* test(e2e): build provider readiness tool fixture
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* fix(e2e): use a dedicated Noble fixture for Docker tests
Signed-off-by: Evan Lezar <elezar@nvidia.com>
---------
Signed-off-by: Akram
Signed-off-by: Akram <akram.benaissi@gmail.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Evan Lezar <elezar@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>