mirror of
https://github.com/NVIDIA/OpenShell.git
synced 2026-10-06 17:34:26 +08:00
* fix(sandbox): accept local connections natively on loopback-confined sockets On kernels without SECCOMP_FILTER_FLAG_WAIT_KILLABLE_RECV (RHEL 9 / RHCOS 5.14), the broker cannot safely write an accepted peer address into workload memory, so accept/accept4 with a peer-address buffer failed with EOPNOTSUPP. Static binaries and Go servers, which issue the raw syscall, could not accept connections at all. Move local acceptance into the kernel and replace per-accept inspection with standing kernel confinement: - Bind every broker-created TCP/UDP socket to the loopback device before injection and verify the binding. Accepted sockets inherit it, so they can neither receive routed ingress nor emit routed egress. - Stop notifying accept/accept4 and remove the accept workers, the SIGUSR2 accept-interrupt monitor, and the 64-accept ceiling. - Continue getpeername natively for every descriptor except relayed connections. - Deny interface-selection socket options and MSG_FASTOPEN sends from scalar syscall arguments, so confinement does not depend on the capability state of the user namespace that owns the network namespace and covers unregistered descriptors. - Drop loopback-interface ingress on the non-loopback TCP control listener before it listens, so a workload cannot reach it by reconnecting a natively accepted socket. - Mark descriptors above stdio close-on-exec in every workload pre_exec path. - Require an active confinement probe at qualification and carry it as required authenticated audit evidence. Document OpenShift 4.19 as the minimum release: RHCOS kernels for 4.16 through 4.18 are built without Landlock. Closes #4058 Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): close review gaps in native local accept confinement - Reject a TCP boundary control listener on a loopback address, including IPv4-mapped loopback. Production drivers bind the unspecified address; a loopback listener would be reachable from workload sockets the broker does not track. - Extend the confinement probe to IPv6 and prove that an accepted socket keeps its loopback binding after an AF_UNSPEC disconnect. - Pin the accepted-peer contract with tests: a loopback-bound listener admits only clients in the sandbox network namespace, and workload sockets cannot bind a non-loopback source address. - Correct the support matrix: accepted connections are bounded by per-process descriptor limits, the runtime PID limit, and the sandbox memory limit, not by a broker limit. Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(sandbox): replace unsafe socket calls with socket2 and rustix Socket confinement now uses socket2 for device binding and the ingress filter, and rustix for interface lookup and AF_UNSPEC disconnect, so the module contains no unsafe code. The control-listener filter is no longer locked: no safe API exposes SO_LOCK_FILTER, and the listener descriptor never leaves the trusted sandbox process, which marks every descriptor above stdio close-on-exec before running workload code. Tests added for native accept use rustix and socket2 instead of raw libc calls. rustix issues accept4 and getpeername as raw syscalls, so the direct-syscall test keeps its meaning. The close-on-exec sweep test runs in a re-executed test process instead of a forked child. The close_range(CLOSE_RANGE_CLOEXEC) syscall in the pre_exec hook remains the only unsafe added by this branch; neither rustix nor nix wraps it. Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): allowlist workload socket families The workload seccomp filter denied only AF_PACKET, AF_BLUETOOTH, and AF_VSOCK, and the broker continued socket() for every non-INET family. Several protocol families, including AF_RXRPC, AF_SMC, and AF_KCM, carry traffic over kernel-owned sockets that the broker never creates and that are not bound to loopback. Allow only AF_UNIX, AF_NETLINK (still limited to NETLINK_ROUTE), and the brokered AF_INET/AF_INET6 families, and restrict socketpair(2) to AF_UNIX. Both decisions use the scalar domain argument. The broker independently refuses non-INET families other than AF_UNIX and AF_NETLINK with EAFNOSUPPORT. Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(sandbox): remove legacy read-only mode After native local accept, only two broker paths still wrote into workload memory: getpeername on relayed connections and the per-message lengths of a first sendmmsg to the DNS relay. Legacy mode existed only to refuse those writes on kernels without WAIT_KILLABLE_RECV, and the getpeername refusal broke CPython TLS on RHEL 9. The broker now never writes workload memory: - getpeername is no longer mediated and reports the kernel peer on every kernel; relayed connections report the loopback relay address. - A first send to the DNS relay pins the broker's socket copy to the relay and continues the syscall, so the kernel performs the send and writes any per-message results. Remove the listener mode, the task-memory write path and probe, and the task_memory_write, cancellation, and task_memory_writes_disabled audit evidence fields. WAIT_KILLABLE_RECV is still used when available and is reported for diagnostics, but no mediation decision depends on it. Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(sandbox): drop WAIT_KILLABLE_RECV and make handlers restart-safe The broker no longer writes workload memory, so killable notification waits only reduced how often a signal restarts a notified syscall. RHEL 9 kernels never had them, so the handlers must tolerate restarts anyway. Install the listener without the flag on every kernel instead of special-casing newer ones. Handlers now check that the notification is still live immediately before each side effect (bind, connect of the retained socket, listen, and the first DNS send), and answer a restarted operation the broker already completed as the kernel would: a repeated TCP connect returns EISCONN, repeating the same UDP association succeeds, a repeat of a completed bind succeeds, and a failed relay reports its errno. Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): finish teardown only when every workload descendant has exited Teardown waited only for registered process groups to disappear. A root is unregistered once reaped, so a descendant that ignored SIGTERM could outlive it while termination reported success and never sent SIGKILL, violating the bounded-termination requirement. The sandbox now becomes a child subreaper when it is not PID 1, so orphaned descendants stay in its tree and are reaped. Termination waits until no live descendant remains, and every scanned process is signalled through a pidfd after confirming its start time, so a reused PID is never signalled. Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): keep a frozen workload from resuming itself Freezing stops every workload process with SIGSTOP. tgkill and rt_tgsigqueueinfo were not mediated, and mediated kill, rt_sigqueueinfo, and tkill passed SIGCONT through, so a workload process that was not yet stopped could resume the others while the supervisor recovered. Mediate tgkill and rt_tgsigqueueinfo like tkill: refuse targets in the sandbox thread group, report a thread outside the named group as missing, and continue otherwise. The boundary marks the broker frozen before stopping the workload and clears it after resuming; while frozen, workload requests to send SIGCONT fail with EPERM. Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): send mediated DNS datagrams from the broker socket The native UDP send path re-ran the workload's own sendmsg, so per-message ancillary data such as IP_PKTINFO rode along; only the loopback destination contained a routing override. The broker now reads the datagram and sends it from its retained, relay-connected socket, which it builds without ancillary data, so a per-message override cannot redirect the packet. It writes nothing back into workload memory. A send carrying control data is refused with EOPNOTSUPP rather than silently stripped. Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): keep sandbox control variables out of the workload environment The canonical process inherits the sandbox environment so the image's own ENV reaches the workload, then removed only a denylist of credential variables. Other variables in the reserved OPENSHELL_ namespace, such as the serialized user environment and the log level, still reached the workload. Remove every inherited OPENSHELL_ variable before applying the declared environment, and restore OPENSHELL_SANDBOX=1. The image's ordinary ENV and declared variables are unaffected; declared variables cannot use the reserved namespace. Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): keep /proc read-only for GPU workloads GPU mode granted read-write access to all of /proc so CUDA's cuInit could write thread names to /proc/<pid>/task/<tid>/comm. Keep /proc read-only. Open mediation now serves a writable open of the caller's own thread comm file: the broker opens it and injects the descriptor, with no syscall continued. The kernel accepts a comm write only from the target's own thread group, so a substituted path or reused thread ID cannot rename another process's thread. Any other /proc write is left to Landlock, which denies it. Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): keep the broker socket for connected UDP DNS sockets Sending DNS datagrams from the broker's socket required the broker to keep its copy, but the connect paths to the relay and to loopback still released it. glibc connects the resolver socket and then sends A and AAAA together with sendmmsg, which is mediated, so resolution failed. Release the copy on those connect paths only for TCP. Add a regression test for connect followed by sendmmsg that fails fast rather than hanging. Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): keep slow loopback connects from stalling mediation The loopback connect path held the registry lock while polling a TCP connect for up to five seconds on the single notification dispatcher. A workload connecting to a busy local listener stalled every other mediated syscall, including opens and signals. Connect a duplicate of the retained socket without holding the lock. A nonblocking socket gets the native EINPROGRESS and the kernel completes the handshake on the shared socket; a blocking socket waits on a bounded worker thread. Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): contain network broker handler panics A panic in a notification handler unwound the single broker thread while the health flag still read healthy, so every later blocked workload syscall hung until the sandbox was killed. Run each dispatch under catch_unwind: a panicking handler fails only that syscall with EIO and the broker keeps mediating. If the broker thread ever exits, mark it unhealthy so dependent operations fail closed instead of blocking. Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): reject ancillary data on DNS sends instead of broker-sending Sending mediated DNS datagrams from the broker's own socket required keeping a broker handle for each DNS socket's whole life, which regressed real name resolution through reclaim and resource accounting that the mock-based unit tests did not exercise. Revert to continuing the kernel send, but refuse a send that carries ancillary control data (msg_controllen != 0) at read time, so a per-message routing override such as IP_PKTINFO cannot ride a mediated DNS send. A loopback destination contains an override that races the check. Full broker-side UDP mediation is left to a separate change with deployment e2e. Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): address review findings across the hardening changes Fixes: - The blocking local-connect worker no longer toggles O_NONBLOCK on the open file it shares with the workload; it waits with a plain connect. - tgkill and rt_tgsigqueueinfo are notified only when they send SIGCONT, so ordinary thread signals such as Go preemption stay in the kernel. The frozen flag is re-read immediately before delivery. - A shell redirect to the caller's own thread comm file works again; O_CREAT and O_TRUNC are no-ops there and O_EXCL returns EEXIST. - The teardown scan is a linear walk and fails closed when /proc cannot be read. Simplifications: - Remove tests that need infrastructure outside the repo or prove nothing: the topology harness tests, the static-server benchmark, the direct-syscall accept test, the comm rename test without Landlock, a serde-default test, and the test-only panic hook in setsockopt. - Drop the unused Failed connect outcome, the redundant read-back after binding to loopback, redundant OPENSHELL_SANDBOX settings, and stale or duplicated comments; reuse the reserved environment prefix constant. - Tighten the support matrix and OpenShift wording. Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(cli): wait for input when piped exec stdin is nonblocking Processes that inherit the same stdin share one open file description, so any of them can make it nonblocking for all. `sandbox exec` then read no input yet, got EAGAIN, and failed with "Resource temporarily unavailable (os error 11)". Parallel e2e tests inherit the runner's stdin, which made credential_gating fail intermittently. The stdin reader now blocks in poll() until input or end of file arrives instead of treating EAGAIN as fatal, and it leaves the shared descriptor's flags alone. The e2e exec helper also stops inheriting the runner's stdin, since its commands take no input. Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): fix macOS lint and older-glibc test linking Import HashSet only in the Linux-only process scan, and call gettid through the raw syscall in a test, since the libc wrapper needs glibc 2.30 or newer. Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(sandbox): clarify socket peer address reporting Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com>