* fix(cli): preserve existing directories during single-file upload
Inspect remote destination types before choosing archive entry names.
Keep directory and source symlinks intact, recheck detected type changes,
and report the resulting path. Add real tar and live conformance coverage.
Closes#4175
Signed-off-by: Shiju <shiju@nvidia.com>
* fix(cli): box upload future in sandbox_create
Boxing the upload future keeps the sandbox_create future under the clippy::large_futures threshold after sandbox_upload_planned began returning the uploaded path.
Signed-off-by: Shiju <shiju@nvidia.com>
---------
Signed-off-by: Shiju <shiju@nvidia.com>
* fix(supervisor): report rejected OCI working directories as WorkspaceValidationFailed
The RFC 0012 split dropped the supervisor exit status that drivers map to
the WorkspaceValidationFailed condition. An image WORKDIR the sandbox
identity cannot use then surfaced as ControlSupervisorStartFailed on
Docker and as a signal kill on Podman.
The supervisor now exits with the reserved status when the sandbox rejects
the image working directory. Docker maps that supervisor exit during
readiness and monitoring, and Podman maps the supervisor companion's exit
instead of the workload's.
Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>
* feat(podman): honor OCI image working directories
Podman sandboxes now use a custom OCI image WORKDIR as the workspace and
as the working directory for agent commands, matching Docker. Empty, /,
and /sandbox values keep the managed /sandbox workspace volume.
A custom workspace stays in the image's container filesystem: no
workspace volume, no archive upload, and no root setup step. Resolve the
image ID, user, environment, and working directory from one pinned
inspection, validate the workdir with the shared OCI rules, and reject
Podman control-path overlaps and image volumes or driver mounts that
cover it. The runtime starts from / and passes the resolved path to agent
commands with --workdir.
Add podman_oci_identity to the Podman CI test list.
Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>
* fix(drivers): preserve custom workspace failures across supervisor cleanup
Check the Podman custom workspace rather than the absent managed volume in the OCI identity E2E. Classify Bollard wait errors by exit code and record Docker workspace failures before removing the supervisor so readiness retains the specific failure reason.
Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>
* fix(docker): use sandbox ID for recorded startup failures
Keep the workload container ID for Docker inspection and use the sandbox ID to retrieve the monitor failure. Test with distinct identifiers so the lookup cannot accidentally pass with a container ID.
Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>
* refactor(podman): keep WORKDIR change focused on workspace behavior
Remove the new cross-driver workspace failure classification path and its monitor/readiness workaround. Preserve Podman WORKDIR resolution, mount safety, and workspace rejection without requiring a distinct condition reason.
Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>
* refactor(podman): simplify inspected image metadata and test fixtures
Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>
* code review
Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>
---------
Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>
* ci(release): add advisory compatibility review
Signed-off-by: Simon Scatton <sscatton@nvidia.com>
* test(ci): remove compatibility reviewer test harness
Signed-off-by: Simon Scatton <sscatton@nvidia.com>
* ci(release): run compatibility SDK in the dev shell
Signed-off-by: Simon Scatton <sscatton@nvidia.com>
* ci(release): trigger a one-shot compatibility review
Signed-off-by: Simon Scatton <sscatton@nvidia.com>
* ci(release): remove one-shot review trigger
Signed-off-by: Simon Scatton <sscatton@nvidia.com>
---------
Signed-off-by: Simon Scatton <sscatton@nvidia.com>
Adds the two stdlib-only helpers that back the maintainer approval gate,
ahead of the workflows that call them.
check_maintainer_approval.py parses the linked handles out of
MAINTAINERS.md, folds a review list to each reviewer's latest decisive
position, and exits non-zero unless a maintainer's latest position is an
approval. It fails closed when the list yields no handles.
alert_maintainer_change.py renders the approver-set delta between two
versions of MAINTAINERS.md, and prints nothing when the set is unchanged.
Landing these first lets the gate workflow, which reads both the list and
the decision logic from main rather than from the pull request ref,
actually execute on the pull request that introduces it.
Signed-off-by: Jim Meyer <jimeyer@nvidia.com>
* feat(server): place supervisor sessions by consistent hash and hand them off on shutdown
Signed-off-by: divesh <dgude@nvidia.com>
* feat: route long-lived sandbox connections to the owning gateway replica
Signed-off-by: divesh <dgude@nvidia.com>
* feat(helm): add opt-in per-replica GRPCRoute routing
Signed-off-by: divesh <dgude@nvidia.com>
---------
Signed-off-by: divesh <dgude@nvidia.com>
* docs(observability): correct supervisor log access and retention
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
* docs(observability): clarify runtime log sources and retention limits
Remove Docker copy guidance for the live supervisor tmpfs mount while retaining the Podman example. Distinguish sandbox-runtime security events from supervisor JSONL and gateway log streams, and qualify independent three-file retention.
Validation: mise run docs, Markdown lint, and git diff --check passed. Fern reported the same three existing unrelated warnings.
Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>
---------
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>
Co-authored-by: Matthew Grossman <mgrossman@nvidia.com>
`openshell forward service` minted an SSH session token before every
forwarded TCP connection and revoked it afterwards: two store commits per
connection. The token added nothing on that path. `ForwardTcp` already
authenticates the caller and authorizes it against the sandbox's workspace
on every stream before it looks at the token, the relay to the supervisor
is opened with the sandbox id and target only, and the token is never
forwarded, audited, or visible to the target service. The mechanism exists
for `openshell sandbox ssh`, where the process that opens the stream is an
ssh ProxyCommand holding nothing but the token. Reusing it per TCP
connection put a store write on the connect path and serialized concurrent
forwards on commit latency: #3494 measured the symptom, and #3543 made the
commits cheaper, but each one still holds SQLite's writer lock for an
fsync, so connection setup under a burst stayed linear in the number of
concurrent connections.
Let `target.tcp` streams omit `authorization_token`. The gateway admits
them on the already-authorized principal, counts them against the same
per-sandbox connection cap, and touches no store. `target.ssh` streams keep
requiring the token. A token supplied with a TCP target is still validated
and counted per token, so an older CLI against a new gateway is unchanged.
The CLI stops minting and revoking a session per forwarded connection;
against a gateway that predates this change it recognizes the
`authorization_token is required` rejection once and falls back to
per-connection tokens for the rest of that forward.
Tests cover token-less TCP admission and slot release, SSH targets still
rejected without a token, a supplied token still validated, and the
per-sandbox cap for token-less forwards. CLI integration tests run
`service_forward_tcp` against a mock gateway: token-less inits echo data
with no CreateSshSession or RevokeSshSession call, and a gateway that
rejects the empty token is detected once, after which every connection in
that forward mints and revokes its own token. Architecture and security
docs describe which targets carry a token, and the per-token connection
limit now reads 3, matching the gateway.
Signed-off-by: Jason T. Greene <jason.greene@redhat.com>
* fix(ssh): persist sandbox host identities
Store each sandbox's Ed25519 host key in the gateway credential store and
deliver it only to the supervisor. Preserve identity across restarts,
delete owned credentials with the sandbox, and expose the public SHA256
fingerprint through sandbox and SSH-session APIs and client SDKs.
Cover credential ownership, cancellation, deletion retries, client
compatibility, and pinned SSH connections through lifecycle transitions.
Closes#3835
Signed-off-by: Mike Nguyen <miken@nvidia.com>
* fix(compute): clean up failed sandbox SSH identity creation
Signed-off-by: Mike Nguyen <miken@nvidia.com>
* test(ssh): wait for sandbox deletion before name reuse
Signed-off-by: Mike Nguyen <miken@nvidia.com>
---------
Signed-off-by: Mike Nguyen <miken@nvidia.com>
* fix(sandbox): accept local connections natively on loopback-confined sockets
On kernels without SECCOMP_FILTER_FLAG_WAIT_KILLABLE_RECV (RHEL 9 / RHCOS
5.14), the broker cannot safely write an accepted peer address into
workload memory, so accept/accept4 with a peer-address buffer failed with
EOPNOTSUPP. Static binaries and Go servers, which issue the raw syscall,
could not accept connections at all.
Move local acceptance into the kernel and replace per-accept inspection
with standing kernel confinement:
- Bind every broker-created TCP/UDP socket to the loopback device before
injection and verify the binding. Accepted sockets inherit it, so they
can neither receive routed ingress nor emit routed egress.
- Stop notifying accept/accept4 and remove the accept workers, the
SIGUSR2 accept-interrupt monitor, and the 64-accept ceiling.
- Continue getpeername natively for every descriptor except relayed
connections.
- Deny interface-selection socket options and MSG_FASTOPEN sends from
scalar syscall arguments, so confinement does not depend on the
capability state of the user namespace that owns the network namespace
and covers unregistered descriptors.
- Drop loopback-interface ingress on the non-loopback TCP control
listener before it listens, so a workload cannot reach it by
reconnecting a natively accepted socket.
- Mark descriptors above stdio close-on-exec in every workload pre_exec
path.
- Require an active confinement probe at qualification and carry it as
required authenticated audit evidence.
Document OpenShift 4.19 as the minimum release: RHCOS kernels for 4.16
through 4.18 are built without Landlock.
Closes#4058
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* fix(sandbox): close review gaps in native local accept confinement
- Reject a TCP boundary control listener on a loopback address,
including IPv4-mapped loopback. Production drivers bind the
unspecified address; a loopback listener would be reachable from
workload sockets the broker does not track.
- Extend the confinement probe to IPv6 and prove that an accepted
socket keeps its loopback binding after an AF_UNSPEC disconnect.
- Pin the accepted-peer contract with tests: a loopback-bound listener
admits only clients in the sandbox network namespace, and workload
sockets cannot bind a non-loopback source address.
- Correct the support matrix: accepted connections are bounded by
per-process descriptor limits, the runtime PID limit, and the sandbox
memory limit, not by a broker limit.
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* refactor(sandbox): replace unsafe socket calls with socket2 and rustix
Socket confinement now uses socket2 for device binding and the ingress
filter, and rustix for interface lookup and AF_UNSPEC disconnect, so the
module contains no unsafe code. The control-listener filter is no longer
locked: no safe API exposes SO_LOCK_FILTER, and the listener descriptor
never leaves the trusted sandbox process, which marks every descriptor
above stdio close-on-exec before running workload code.
Tests added for native accept use rustix and socket2 instead of raw libc
calls. rustix issues accept4 and getpeername as raw syscalls, so the
direct-syscall test keeps its meaning. The close-on-exec sweep test runs
in a re-executed test process instead of a forked child.
The close_range(CLOSE_RANGE_CLOEXEC) syscall in the pre_exec hook remains
the only unsafe added by this branch; neither rustix nor nix wraps it.
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* fix(sandbox): allowlist workload socket families
The workload seccomp filter denied only AF_PACKET, AF_BLUETOOTH, and
AF_VSOCK, and the broker continued socket() for every non-INET family.
Several protocol families, including AF_RXRPC, AF_SMC, and AF_KCM, carry
traffic over kernel-owned sockets that the broker never creates and that
are not bound to loopback.
Allow only AF_UNIX, AF_NETLINK (still limited to NETLINK_ROUTE), and the
brokered AF_INET/AF_INET6 families, and restrict socketpair(2) to
AF_UNIX. Both decisions use the scalar domain argument. The broker
independently refuses non-INET families other than AF_UNIX and
AF_NETLINK with EAFNOSUPPORT.
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* refactor(sandbox): remove legacy read-only mode
After native local accept, only two broker paths still wrote into
workload memory: getpeername on relayed connections and the per-message
lengths of a first sendmmsg to the DNS relay. Legacy mode existed only to
refuse those writes on kernels without WAIT_KILLABLE_RECV, and the
getpeername refusal broke CPython TLS on RHEL 9.
The broker now never writes workload memory:
- getpeername is no longer mediated and reports the kernel peer on every
kernel; relayed connections report the loopback relay address.
- A first send to the DNS relay pins the broker's socket copy to the
relay and continues the syscall, so the kernel performs the send and
writes any per-message results.
Remove the listener mode, the task-memory write path and probe, and the
task_memory_write, cancellation, and task_memory_writes_disabled audit
evidence fields. WAIT_KILLABLE_RECV is still used when available and is
reported for diagnostics, but no mediation decision depends on it.
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* refactor(sandbox): drop WAIT_KILLABLE_RECV and make handlers restart-safe
The broker no longer writes workload memory, so killable notification
waits only reduced how often a signal restarts a notified syscall. RHEL 9
kernels never had them, so the handlers must tolerate restarts anyway.
Install the listener without the flag on every kernel instead of
special-casing newer ones.
Handlers now check that the notification is still live immediately
before each side effect (bind, connect of the retained socket, listen,
and the first DNS send), and answer a restarted operation the broker
already completed as the kernel would: a repeated TCP connect returns
EISCONN, repeating the same UDP association succeeds, a repeat of a
completed bind succeeds, and a failed relay reports its errno.
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* fix(sandbox): finish teardown only when every workload descendant has exited
Teardown waited only for registered process groups to disappear. A root
is unregistered once reaped, so a descendant that ignored SIGTERM could
outlive it while termination reported success and never sent SIGKILL,
violating the bounded-termination requirement.
The sandbox now becomes a child subreaper when it is not PID 1, so
orphaned descendants stay in its tree and are reaped. Termination waits
until no live descendant remains, and every scanned process is signalled
through a pidfd after confirming its start time, so a reused PID is never
signalled.
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* fix(sandbox): keep a frozen workload from resuming itself
Freezing stops every workload process with SIGSTOP. tgkill and
rt_tgsigqueueinfo were not mediated, and mediated kill, rt_sigqueueinfo,
and tkill passed SIGCONT through, so a workload process that was not yet
stopped could resume the others while the supervisor recovered.
Mediate tgkill and rt_tgsigqueueinfo like tkill: refuse targets in the
sandbox thread group, report a thread outside the named group as
missing, and continue otherwise. The boundary marks the broker frozen
before stopping the workload and clears it after resuming; while frozen,
workload requests to send SIGCONT fail with EPERM.
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* fix(sandbox): send mediated DNS datagrams from the broker socket
The native UDP send path re-ran the workload's own sendmsg, so
per-message ancillary data such as IP_PKTINFO rode along; only the
loopback destination contained a routing override.
The broker now reads the datagram and sends it from its retained,
relay-connected socket, which it builds without ancillary data, so a
per-message override cannot redirect the packet. It writes nothing back
into workload memory. A send carrying control data is refused with
EOPNOTSUPP rather than silently stripped.
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* fix(sandbox): keep sandbox control variables out of the workload environment
The canonical process inherits the sandbox environment so the image's
own ENV reaches the workload, then removed only a denylist of credential
variables. Other variables in the reserved OPENSHELL_ namespace, such as
the serialized user environment and the log level, still reached the
workload.
Remove every inherited OPENSHELL_ variable before applying the declared
environment, and restore OPENSHELL_SANDBOX=1. The image's ordinary ENV
and declared variables are unaffected; declared variables cannot use the
reserved namespace.
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* fix(sandbox): keep /proc read-only for GPU workloads
GPU mode granted read-write access to all of /proc so CUDA's cuInit could
write thread names to /proc/<pid>/task/<tid>/comm.
Keep /proc read-only. Open mediation now serves a writable open of the
caller's own thread comm file: the broker opens it and injects the
descriptor, with no syscall continued. The kernel accepts a comm write
only from the target's own thread group, so a substituted path or reused
thread ID cannot rename another process's thread. Any other /proc write
is left to Landlock, which denies it.
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* fix(sandbox): keep the broker socket for connected UDP DNS sockets
Sending DNS datagrams from the broker's socket required the broker to
keep its copy, but the connect paths to the relay and to loopback still
released it. glibc connects the resolver socket and then sends A and
AAAA together with sendmmsg, which is mediated, so resolution failed.
Release the copy on those connect paths only for TCP. Add a regression
test for connect followed by sendmmsg that fails fast rather than
hanging.
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* fix(sandbox): keep slow loopback connects from stalling mediation
The loopback connect path held the registry lock while polling a TCP
connect for up to five seconds on the single notification dispatcher.
A workload connecting to a busy local listener stalled every other
mediated syscall, including opens and signals.
Connect a duplicate of the retained socket without holding the lock. A
nonblocking socket gets the native EINPROGRESS and the kernel completes
the handshake on the shared socket; a blocking socket waits on a bounded
worker thread.
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* fix(sandbox): contain network broker handler panics
A panic in a notification handler unwound the single broker thread while
the health flag still read healthy, so every later blocked workload
syscall hung until the sandbox was killed.
Run each dispatch under catch_unwind: a panicking handler fails only that
syscall with EIO and the broker keeps mediating. If the broker thread
ever exits, mark it unhealthy so dependent operations fail closed instead
of blocking.
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* fix(sandbox): reject ancillary data on DNS sends instead of broker-sending
Sending mediated DNS datagrams from the broker's own socket required
keeping a broker handle for each DNS socket's whole life, which regressed
real name resolution through reclaim and resource accounting that the
mock-based unit tests did not exercise.
Revert to continuing the kernel send, but refuse a send that carries
ancillary control data (msg_controllen != 0) at read time, so a
per-message routing override such as IP_PKTINFO cannot ride a mediated
DNS send. A loopback destination contains an override that races the
check. Full broker-side UDP mediation is left to a separate change with
deployment e2e.
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* fix(sandbox): address review findings across the hardening changes
Fixes:
- The blocking local-connect worker no longer toggles O_NONBLOCK on the
open file it shares with the workload; it waits with a plain connect.
- tgkill and rt_tgsigqueueinfo are notified only when they send SIGCONT,
so ordinary thread signals such as Go preemption stay in the kernel.
The frozen flag is re-read immediately before delivery.
- A shell redirect to the caller's own thread comm file works again;
O_CREAT and O_TRUNC are no-ops there and O_EXCL returns EEXIST.
- The teardown scan is a linear walk and fails closed when /proc cannot
be read.
Simplifications:
- Remove tests that need infrastructure outside the repo or prove
nothing: the topology harness tests, the static-server benchmark, the
direct-syscall accept test, the comm rename test without Landlock, a
serde-default test, and the test-only panic hook in setsockopt.
- Drop the unused Failed connect outcome, the redundant read-back after
binding to loopback, redundant OPENSHELL_SANDBOX settings, and stale or
duplicated comments; reuse the reserved environment prefix constant.
- Tighten the support matrix and OpenShift wording.
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* fix(cli): wait for input when piped exec stdin is nonblocking
Processes that inherit the same stdin share one open file description,
so any of them can make it nonblocking for all. `sandbox exec` then read
no input yet, got EAGAIN, and failed with "Resource temporarily
unavailable (os error 11)". Parallel e2e tests inherit the runner's
stdin, which made credential_gating fail intermittently.
The stdin reader now blocks in poll() until input or end of file arrives
instead of treating EAGAIN as fatal, and it leaves the shared descriptor's
flags alone. The e2e exec helper also stops inheriting the runner's
stdin, since its commands take no input.
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* fix(sandbox): fix macOS lint and older-glibc test linking
Import HashSet only in the Linux-only process scan, and call gettid
through the raw syscall in a test, since the libc wrapper needs glibc
2.30 or newer.
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* docs(sandbox): clarify socket peer address reporting
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
---------
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
When the gateway exports OTLP traces, compute drivers pass the gateway's
endpoint to supervisors as OPENSHELL_OTLP_ENDPOINT, along with TRACEPARENT
from the operation that launched them. The Kubernetes, Docker, Podman, and
VM drivers all receive the endpoint from the gateway; driver TOML cannot
set it. On Podman Machine, the Podman driver points a loopback endpoint
at host.containers.internal, since supervisor loopback is the VM's.
The supervisor exports spans as openshell-supervisor, tagged with
the sandbox ID, and flushes them before exiting.
supervisor.startup joins the sandbox's creation trace and covers image
policy discovery, policy load, and boundary attach, confirm, agent start,
and access start. Supervisor calls to the gateway carry W3C trace context,
so the gateway's server spans nest under them.
Each egress connection emits supervisor.egress.connect with authorize,
resolve, and dial children. These spans use DEBUG level because every
outbound connection starts its own trace. OpenShell spans export at INFO,
or at the sandbox log level when it is debug or trace; spans from other
libraries export at INFO.
Signed-off-by: Kris Hicks <khicks@nvidia.com>
* feat(providers): compose dynamic credential grants per request
Select the most-specific grant independently for each protected header,
preserving each credential's issuer, audience and cache identity. Resolve
all selected grants before rewriting the request, replace agent-supplied
header copies, and fail closed on collisions or acquisition failures.
Allow distinct-header compositions in gateway validation and redact raw
issuer errors. Cover independent cache entries, atomic TLS relay failure,
header replacement and concurrent request isolation; document the profile
contract.
Fixes#3320
Signed-off-by: Shiju <shiju@nvidia.com>
* fix(providers): correct dynamic grant CI checks
Use if-let for a grant result whose error payload is deliberately ignored,
and name the empty query map type in the regression fixture. Preserve grant
acquisition, failure redaction and request atomicity. Update the token-exchange
failure test to require the sanitized error instead of raw issuer text.
Signed-off-by: Shiju <shiju@nvidia.com>
---------
Signed-off-by: Shiju <shiju@nvidia.com>
* refactor(supervisor): move backend setup behind the selected backend
Keep launch decoding, policy discovery and client construction with the
trusted backend while sharing supervisor admission and lifecycle handling.
Preserve VM policy identity validation and live credential handles.
Exercise alternate launch data through shared startup, readiness and
shutdown, including confirmation denial and resource cleanup.
Closes#4173
Signed-off-by: Shiju <shiju@nvidia.com>
* fix(supervisor): resolve backend setup lint failures
Box shared startup futures to keep callers' async state small.
Use the private module boundary for backend setup visibility and keep
test helpers and defaults explicit without changing their inputs.
Signed-off-by: Shiju <shiju@nvidia.com>
* refactor(supervisor): simplify selected backend startup
Signed-off-by: Shiju <shiju@nvidia.com>
* test(supervisor): keep imports before descriptor assertions
Signed-off-by: Shiju <shiju@nvidia.com>
* test(supervisor): preserve discovery failure startup coverage
Exercise backend discovery failure through shared startup before the
existing confirmation denial and successful shutdown cases. Assert that
startup preserves the error without launching a workload or readiness.
Signed-off-by: Shiju <shiju@nvidia.com>
---------
Signed-off-by: Shiju <shiju@nvidia.com>
Revert PR #4107 and restore binary installers for the Podman driver qualification lanes.
This reverts commit 8990656c13.
Signed-off-by: Simon Scatton <sscatton@nvidia.com>
* fix(podman): create managed workspace volumes owned by the workload identity
Podman now creates the managed /sandbox volume with uid/gid options for the
resolved workload identity, so the workload starts directly as that
identity. This fixes rootful sandboxes whose image USER or policy
run_as_user could not write to a root-owned /sandbox, and removes the
root-then-drop workspace chown start path.
Resource admission accepts the managed workspace volume when its options
match the workload container's final identity, or are empty for volumes
created by older gateways. The channel volume still requires empty options.
Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>
* test(podman): cover managed volume reuse and workspace access
Signed-off-by: Evan Lezar <elezar@nvidia.com>
* refactor(podman): clarify managed volume creation and validation
Signed-off-by: Evan Lezar <elezar@nvidia.com>
* test(podman): verify workspace access across user namespaces
Signed-off-by: Evan Lezar <elezar@nvidia.com>
---------
Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>
Co-authored-by: Evan Lezar <elezar@nvidia.com>
* feat(server): add gateway capacity metrics and optional HPA
Expose per-replica supervisor sessions, pending relay capacity, relay
rejections and claim latency, and outbound peer request outcomes and
latency. Use bounded labels and Prometheus histograms for the new latency
metrics while preserving existing summary metrics.
Add an optional Helm HPA with external-database and resource validation,
conservative scale-down defaults, and support for custom metrics. Keep
certificate hook pods outside gateway workload selectors.
Document per-pod scraping, scaling limits, upgrade behavior, and
PostgreSQL connection sizing.
Part of #3528
Signed-off-by: Emilien Macchi <emacchi@redhat.com>
* feat(server): unify routing metrics and clarify replica capacity
Combine local relay setup and outbound peer requests in one counter, labeled by operation, route, target, outcome, and gRPC status. Preserve peer latency metrics.
Rename the rejection reason from global_capacity to replica_capacity to reflect the per-replica relay budget. Keep capacity limits unchanged.
Update tests and documentation.
Signed-off-by: divesh <dgude@nvidia.com>
* fix(server): align routed request metric labels and outcomes
Count a local relay as successful only when its supervisor claims it, the
same event that answers a peer relay on the owner, and rename the outcomes
to local_error and remote_error so they say where an attempt failed. Label
the peer latency histogram by operation, like the routed request counter,
and rename the target label to relay_kind so it does not read as the
Prometheus scrape target.
Rename RelayCapacity.global to per_replica, drop the per-sandbox relay
capacity gauge, which no per-sandbox series can pair with, describe the
latency histogram as peer-only, and restore the note that unavailable
spikes are expected during rollouts.
Part of #3528
Signed-off-by: Emilien Macchi <emacchi@redhat.com>
---------
Signed-off-by: Emilien Macchi <emacchi@redhat.com>
Signed-off-by: divesh <dgude@nvidia.com>
Co-authored-by: divesh <dgude@nvidia.com>
Closes#3502
Resolve trusted OCI runtime images with driver TOML, process environment, and compiled-default precedence across Docker, Podman, and Kubernetes.
Signed-off-by: Evan Lezar <elezar@nvidia.com>
* fix(vm): validate launch credentials before preparation
Negotiate a launch-authentication requirement and reject incomplete gateway
configuration before driver validation, archive consumption or persistence.
Validate replacement VM credentials before stopping active compute and
prefer explicit signer configuration over local discovery.
Closes#3949. Part of #3955.
Signed-off-by: Shiju <shiju@nvidia.com>
* fix(vm): validate saved launch credentials before restore side effects
Signed-off-by: Shiju <shiju@nvidia.com>
* chore(gateway): satisfy launch preflight lint checks
Signed-off-by: Shiju <shiju@nvidia.com>
* docs(gateway): complete the MicroVM launch-signing example
Include the gateway JWT paths in the standalone VM configuration and show
the output directory required by local certificate generation. Explain
the distinction between configuration preflight and launch-key validation.
Refs #3949. Part of #3955.
Signed-off-by: Shiju <shiju@nvidia.com>
* test(vm): account for preparation state in launch credential fixture
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
---------
Signed-off-by: Shiju <shiju@nvidia.com>
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <9696606+johntmyers@users.noreply.github.com>
* feat(server): separate image preparation and admission deadlines
Give sandbox image preparation its own deadline and start the admission
deadline after preparation finishes. Persist both phases across gateway
restarts and show recovery guidance for the phase that expired.
Preserve the current service authorization schema and regenerate the Go
bindings with the preparation timestamps.
Fixes#3952
Related to #3955
Signed-off-by: Shiju <shiju@nvidia.com>
* fix(cli): simplify preparation timeout fallback selection
Use lazy Option fallbacks while preserving timeout messages and retained
sandbox behavior.
Signed-off-by: Shiju <shiju@nvidia.com>
* docs(server): clarify admission timer prerequisites
Signed-off-by: Shiju <shiju@nvidia.com>
* fix(compute): enforce deadlines during initial sandbox create
Release stalled create operations after preparation expires and preserve
the timeout diagnosis across late driver results. Keep failed-create cleanup
bound to its original attempt so another replica can retry safely.
Signed-off-by: Shiju <shiju@nvidia.com>
* test(server): satisfy deadline regression lints
Drop the create-error mutex guard before matching its cloned value and use
idiomatic iteration and timeout matching in the deadline fixtures.
Signed-off-by: Shiju <shiju@nvidia.com>
* fix(server): retain ownership of pending provisioning operations
Keep submitted create and start requests alive after caller cancellation,
monitor failure, or preparation timeout. Persist request ownership before
dispatch and retain staged uploads while the driver response is pending.
Require timeout cleanup to stop compute after driver settlement without
discarding an active cleanup claim. Fence late result handling against newer
operations, preserve failed-start recovery, and defer automatic restart while
another request owns the sandbox. Expose pending ownership in CLI JSON.
Add ordered multi-replica and cancellation regressions for the review findings.
Signed-off-by: Shiju <shiju@nvidia.com>
* fix(openshell): preserve tracing and accept ready create responses
Carry the request span into the detached provisioning worker so compute
driver calls remain attached to their parent trace after task handoff.
Update the compensation regression to require no backend DELETE when the
durable cleanup claim fails, matching the operation ownership requirement.
Accept the gateway's current Ready snapshot when a sandbox becomes ready
before CREATE returns. Do not require the client to observe an earlier
provisioning phase. Cover the Ready-only watch and command attachment.
Signed-off-by: Shiju <shiju@nvidia.com>
---------
Signed-off-by: Shiju <shiju@nvidia.com>
* fix(relay): close outbound relay stream when target closes first
Closes#3724
out_tx was cloned into the target-reading task, so the function's
own copy kept the outbound relay stream open until the client side
also ended. A target that closed a keep-alive connection never
reached the client as EOF, so reused connections hung forever.
Move out_tx into the target-reading task instead of cloning it, so
dropping it on target EOF ends the outbound stream right away.
Signed-off-by: Eric Curtin <eric.curtin@docker.com>
* fix(relay): keep client uploads after target half-close
Signed-off-by: Eric Curtin <eric.curtin@docker.com>
---------
Signed-off-by: Eric Curtin <eric.curtin@docker.com>
* feat(gateway): validate VM filesystem tools during preflight
Check required local VM tools through config preflight and share executable
resolution with VM image operations. Report selected paths and actionable
errors without creating gateway or sandbox state.
Bound probe output and execution time, and clean up probe descendants on
interruption. Preserve pure static validation and skip local tool checks
for remote driver endpoints and unrelated drivers.
Fixes#3951
Related to #3955
Signed-off-by: Shiju <shiju@nvidia.com>
* fix(gateway): stabilize filesystem preflight checks
Combine identical filesystem-tool error arms and normalize rendered
diagnostics in command tests so terminal wrapping preserves assertions.
Describe driver TLS validation without depending on removed guest fields.
Signed-off-by: Shiju <shiju@nvidia.com>
* test(gateway): serialize preflight fixture paths as TOML
Keep temporary paths quoted and escaped through the TOML serializer
instead of relying on Rust Debug formatting.
Signed-off-by: Shiju <shiju@nvidia.com>
---------
Signed-off-by: Shiju <shiju@nvidia.com>
Run image preparation in an owned worker process, reserve its process
identity until cleanup completes, and protect staging with leases so
cancellation and recovery cannot race with another preparation attempt.
Signed-off-by: Shiju <shiju@nvidia.com>
* fix(vm): enforce the configured workload identity
Reject conflicting policy users and groups before VM image preparation
and before guest attach or process startup changes state. Validate
supervisor policy updates against the protected VM workload identity.
Preserve the gateway CA transport and capability-free sandbox launcher.
Signed-off-by: Shiju <shiju@nvidia.com>
* test(sandbox): clarify VM identity rejection fixtures
Name invalid user and group fixtures distinctly and move the final
workload identity into its group mismatch test.
Signed-off-by: Shiju <shiju@nvidia.com>
* fix(supervisor): align VM identity startup with current APIs
Pass the optional rejection-log key for VM identity failures and keep
generic startup-write regressions free of VM identity constraints.
Repair the call sites after the branch rebase so the identity and cleanup
proposals compile against the current startup helpers.
Signed-off-by: Shiju <shiju@nvidia.com>
* fix(vm): restore inactive sandbox workload identity
Recover the persisted overlay owner before publishing stopped and terminal
sandboxes. Keep resources manageable when identity metadata is invalid.
Clarify fixed MicroVM ownership in policy-generation guidance.
Signed-off-by: Shiju <shiju@nvidia.com>
* test(vm): flush identity fixture before restart
Persist the canonical identity file before readiness and report the observed
exec, canonical and file-owner identities before comparing them.
Signed-off-by: Shiju <shiju@nvidia.com>
---------
Signed-off-by: Shiju <shiju@nvidia.com>
The Docker driver passed a bare reference as CreateImageOptions.from_image
with no tag. The daemon interprets a tagless fromImage as a request for
every tag in the repository and pulls them all (issue #4029).
Normalize a pull reference by appending ':latest' when it has neither an
explicit tag nor a digest, matching 'docker pull' and the Podman driver.
Parsing inspects only the final path component so a registry port (e.g.
'registry:5000/team/app') is not mistaken for a tag and a digest-pinned
reference ('...@sha256:...') is left untouched. Applied at both pull
sites (pull_image and pull_runtime_image).
Signed-off-by: Udaya Tejas <udayatejas2004@gmail.com>
* feat(cli): stream non-TTY exec input before EOF
Add --stream-stdin using the existing interactive exec RPC without a PTY. Preserve separate output streams and enforce the existing 4 MiB cumulative input cap while forwarding input.
Require explicit clean stdin EOF and drain the response through its final gRPC status. Cover held-open input, limits, cancellation, trailers, and default finite-input behavior with subprocess and live sandbox regressions.
Signed-off-by: Shiju <shiju@nvidia.com>
* test(cli): distinguish exec cancellation from stdin EOF
Treat transport termination and response cancellation as separate test observations. Verify explicit stdin EOF through the shared frame writer and cover cancellation in the pinned Tonic decoder.
Signed-off-by: Shiju <shiju@nvidia.com>
* style(cli): use lazy optional stdin dispatch
Signed-off-by: Shiju <shiju@nvidia.com>
* test(cli): use imported duration in stdin EOF regression
Signed-off-by: Shiju <shiju@nvidia.com>
---------
Signed-off-by: Shiju <shiju@nvidia.com>