Files
Drew Newberry 12cec59bf4 fix(sandbox): accept local connections natively on loopback-confined sockets (#4150)
* fix(sandbox): accept local connections natively on loopback-confined sockets

On kernels without SECCOMP_FILTER_FLAG_WAIT_KILLABLE_RECV (RHEL 9 / RHCOS
5.14), the broker cannot safely write an accepted peer address into
workload memory, so accept/accept4 with a peer-address buffer failed with
EOPNOTSUPP. Static binaries and Go servers, which issue the raw syscall,
could not accept connections at all.

Move local acceptance into the kernel and replace per-accept inspection
with standing kernel confinement:

- Bind every broker-created TCP/UDP socket to the loopback device before
  injection and verify the binding. Accepted sockets inherit it, so they
  can neither receive routed ingress nor emit routed egress.
- Stop notifying accept/accept4 and remove the accept workers, the
  SIGUSR2 accept-interrupt monitor, and the 64-accept ceiling.
- Continue getpeername natively for every descriptor except relayed
  connections.
- Deny interface-selection socket options and MSG_FASTOPEN sends from
  scalar syscall arguments, so confinement does not depend on the
  capability state of the user namespace that owns the network namespace
  and covers unregistered descriptors.
- Drop loopback-interface ingress on the non-loopback TCP control
  listener before it listens, so a workload cannot reach it by
  reconnecting a natively accepted socket.
- Mark descriptors above stdio close-on-exec in every workload pre_exec
  path.
- Require an active confinement probe at qualification and carry it as
  required authenticated audit evidence.

Document OpenShift 4.19 as the minimum release: RHCOS kernels for 4.16
through 4.18 are built without Landlock.

Closes #4058

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): close review gaps in native local accept confinement

- Reject a TCP boundary control listener on a loopback address,
  including IPv4-mapped loopback. Production drivers bind the
  unspecified address; a loopback listener would be reachable from
  workload sockets the broker does not track.
- Extend the confinement probe to IPv6 and prove that an accepted
  socket keeps its loopback binding after an AF_UNSPEC disconnect.
- Pin the accepted-peer contract with tests: a loopback-bound listener
  admits only clients in the sandbox network namespace, and workload
  sockets cannot bind a non-loopback source address.
- Correct the support matrix: accepted connections are bounded by
  per-process descriptor limits, the runtime PID limit, and the sandbox
  memory limit, not by a broker limit.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(sandbox): replace unsafe socket calls with socket2 and rustix

Socket confinement now uses socket2 for device binding and the ingress
filter, and rustix for interface lookup and AF_UNSPEC disconnect, so the
module contains no unsafe code. The control-listener filter is no longer
locked: no safe API exposes SO_LOCK_FILTER, and the listener descriptor
never leaves the trusted sandbox process, which marks every descriptor
above stdio close-on-exec before running workload code.

Tests added for native accept use rustix and socket2 instead of raw libc
calls. rustix issues accept4 and getpeername as raw syscalls, so the
direct-syscall test keeps its meaning. The close-on-exec sweep test runs
in a re-executed test process instead of a forked child.

The close_range(CLOSE_RANGE_CLOEXEC) syscall in the pre_exec hook remains
the only unsafe added by this branch; neither rustix nor nix wraps it.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): allowlist workload socket families

The workload seccomp filter denied only AF_PACKET, AF_BLUETOOTH, and
AF_VSOCK, and the broker continued socket() for every non-INET family.
Several protocol families, including AF_RXRPC, AF_SMC, and AF_KCM, carry
traffic over kernel-owned sockets that the broker never creates and that
are not bound to loopback.

Allow only AF_UNIX, AF_NETLINK (still limited to NETLINK_ROUTE), and the
brokered AF_INET/AF_INET6 families, and restrict socketpair(2) to
AF_UNIX. Both decisions use the scalar domain argument. The broker
independently refuses non-INET families other than AF_UNIX and
AF_NETLINK with EAFNOSUPPORT.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(sandbox): remove legacy read-only mode

After native local accept, only two broker paths still wrote into
workload memory: getpeername on relayed connections and the per-message
lengths of a first sendmmsg to the DNS relay. Legacy mode existed only to
refuse those writes on kernels without WAIT_KILLABLE_RECV, and the
getpeername refusal broke CPython TLS on RHEL 9.

The broker now never writes workload memory:
- getpeername is no longer mediated and reports the kernel peer on every
  kernel; relayed connections report the loopback relay address.
- A first send to the DNS relay pins the broker's socket copy to the
  relay and continues the syscall, so the kernel performs the send and
  writes any per-message results.

Remove the listener mode, the task-memory write path and probe, and the
task_memory_write, cancellation, and task_memory_writes_disabled audit
evidence fields. WAIT_KILLABLE_RECV is still used when available and is
reported for diagnostics, but no mediation decision depends on it.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(sandbox): drop WAIT_KILLABLE_RECV and make handlers restart-safe

The broker no longer writes workload memory, so killable notification
waits only reduced how often a signal restarts a notified syscall. RHEL 9
kernels never had them, so the handlers must tolerate restarts anyway.
Install the listener without the flag on every kernel instead of
special-casing newer ones.

Handlers now check that the notification is still live immediately
before each side effect (bind, connect of the retained socket, listen,
and the first DNS send), and answer a restarted operation the broker
already completed as the kernel would: a repeated TCP connect returns
EISCONN, repeating the same UDP association succeeds, a repeat of a
completed bind succeeds, and a failed relay reports its errno.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): finish teardown only when every workload descendant has exited

Teardown waited only for registered process groups to disappear. A root
is unregistered once reaped, so a descendant that ignored SIGTERM could
outlive it while termination reported success and never sent SIGKILL,
violating the bounded-termination requirement.

The sandbox now becomes a child subreaper when it is not PID 1, so
orphaned descendants stay in its tree and are reaped. Termination waits
until no live descendant remains, and every scanned process is signalled
through a pidfd after confirming its start time, so a reused PID is never
signalled.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): keep a frozen workload from resuming itself

Freezing stops every workload process with SIGSTOP. tgkill and
rt_tgsigqueueinfo were not mediated, and mediated kill, rt_sigqueueinfo,
and tkill passed SIGCONT through, so a workload process that was not yet
stopped could resume the others while the supervisor recovered.

Mediate tgkill and rt_tgsigqueueinfo like tkill: refuse targets in the
sandbox thread group, report a thread outside the named group as
missing, and continue otherwise. The boundary marks the broker frozen
before stopping the workload and clears it after resuming; while frozen,
workload requests to send SIGCONT fail with EPERM.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): send mediated DNS datagrams from the broker socket

The native UDP send path re-ran the workload's own sendmsg, so
per-message ancillary data such as IP_PKTINFO rode along; only the
loopback destination contained a routing override.

The broker now reads the datagram and sends it from its retained,
relay-connected socket, which it builds without ancillary data, so a
per-message override cannot redirect the packet. It writes nothing back
into workload memory. A send carrying control data is refused with
EOPNOTSUPP rather than silently stripped.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): keep sandbox control variables out of the workload environment

The canonical process inherits the sandbox environment so the image's
own ENV reaches the workload, then removed only a denylist of credential
variables. Other variables in the reserved OPENSHELL_ namespace, such as
the serialized user environment and the log level, still reached the
workload.

Remove every inherited OPENSHELL_ variable before applying the declared
environment, and restore OPENSHELL_SANDBOX=1. The image's ordinary ENV
and declared variables are unaffected; declared variables cannot use the
reserved namespace.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): keep /proc read-only for GPU workloads

GPU mode granted read-write access to all of /proc so CUDA's cuInit could
write thread names to /proc/<pid>/task/<tid>/comm.

Keep /proc read-only. Open mediation now serves a writable open of the
caller's own thread comm file: the broker opens it and injects the
descriptor, with no syscall continued. The kernel accepts a comm write
only from the target's own thread group, so a substituted path or reused
thread ID cannot rename another process's thread. Any other /proc write
is left to Landlock, which denies it.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): keep the broker socket for connected UDP DNS sockets

Sending DNS datagrams from the broker's socket required the broker to
keep its copy, but the connect paths to the relay and to loopback still
released it. glibc connects the resolver socket and then sends A and
AAAA together with sendmmsg, which is mediated, so resolution failed.

Release the copy on those connect paths only for TCP. Add a regression
test for connect followed by sendmmsg that fails fast rather than
hanging.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): keep slow loopback connects from stalling mediation

The loopback connect path held the registry lock while polling a TCP
connect for up to five seconds on the single notification dispatcher.
A workload connecting to a busy local listener stalled every other
mediated syscall, including opens and signals.

Connect a duplicate of the retained socket without holding the lock. A
nonblocking socket gets the native EINPROGRESS and the kernel completes
the handshake on the shared socket; a blocking socket waits on a bounded
worker thread.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): contain network broker handler panics

A panic in a notification handler unwound the single broker thread while
the health flag still read healthy, so every later blocked workload
syscall hung until the sandbox was killed.

Run each dispatch under catch_unwind: a panicking handler fails only that
syscall with EIO and the broker keeps mediating. If the broker thread
ever exits, mark it unhealthy so dependent operations fail closed instead
of blocking.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): reject ancillary data on DNS sends instead of broker-sending

Sending mediated DNS datagrams from the broker's own socket required
keeping a broker handle for each DNS socket's whole life, which regressed
real name resolution through reclaim and resource accounting that the
mock-based unit tests did not exercise.

Revert to continuing the kernel send, but refuse a send that carries
ancillary control data (msg_controllen != 0) at read time, so a
per-message routing override such as IP_PKTINFO cannot ride a mediated
DNS send. A loopback destination contains an override that races the
check. Full broker-side UDP mediation is left to a separate change with
deployment e2e.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): address review findings across the hardening changes

Fixes:
- The blocking local-connect worker no longer toggles O_NONBLOCK on the
  open file it shares with the workload; it waits with a plain connect.
- tgkill and rt_tgsigqueueinfo are notified only when they send SIGCONT,
  so ordinary thread signals such as Go preemption stay in the kernel.
  The frozen flag is re-read immediately before delivery.
- A shell redirect to the caller's own thread comm file works again;
  O_CREAT and O_TRUNC are no-ops there and O_EXCL returns EEXIST.
- The teardown scan is a linear walk and fails closed when /proc cannot
  be read.

Simplifications:
- Remove tests that need infrastructure outside the repo or prove
  nothing: the topology harness tests, the static-server benchmark, the
  direct-syscall accept test, the comm rename test without Landlock, a
  serde-default test, and the test-only panic hook in setsockopt.
- Drop the unused Failed connect outcome, the redundant read-back after
  binding to loopback, redundant OPENSHELL_SANDBOX settings, and stale or
  duplicated comments; reuse the reserved environment prefix constant.
- Tighten the support matrix and OpenShift wording.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(cli): wait for input when piped exec stdin is nonblocking

Processes that inherit the same stdin share one open file description,
so any of them can make it nonblocking for all. `sandbox exec` then read
no input yet, got EAGAIN, and failed with "Resource temporarily
unavailable (os error 11)". Parallel e2e tests inherit the runner's
stdin, which made credential_gating fail intermittently.

The stdin reader now blocks in poll() until input or end of file arrives
instead of treating EAGAIN as fatal, and it leaves the shared descriptor's
flags alone. The e2e exec helper also stops inheriting the runner's
stdin, since its commands take no input.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): fix macOS lint and older-glibc test linking

Import HashSet only in the Linux-only process scan, and call gettid
through the raw syscall in a test, since the libc wrapper needs glibc
2.30 or newer.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(sandbox): clarify socket peer address reporting

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-10-06 04:08:04 +00:00
..