Commit Graph
459 Commits
Author SHA1 Message Date
John T. Myers 67adcf1a3c feat(ocsf): emit full JSON records to supervisor stderr (#4323)
* feat(ocsf): emit full JSON records to supervisor stderr

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(ocsf): preserve console record boundaries

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-10-08 05:36:19 +00:00
Piotr Mlocek 9b5bcdd7a2 fix(ha): keep sandboxes ready across gateway pod rolls (#4321)
* fix(ha): keep sandboxes ready across gateway pod rolls

A gateway pod roll left sandboxes not ready long enough for clients to
fail with "sandbox is not ready", which made the HA e2e test
sandbox_file_sync_survives_gateway_pod_rolls flaky. Two causes:

- The supervisor never reset its reconnect backoff after an accepted
  session, so each gateway restart over a sandbox's lifetime doubled the
  delay until every reconnect waited the full 30s maximum. Reset it
  after an accepted session.
- A stopping gateway replica redirected its supervisors and then
  immediately demoted their sandboxes to Provisioning, before the
  redirected supervisor published its replacement session. Wait up to
  5s for the replacement before demoting; the existing disconnect path
  keeps the sandbox Ready once a replacement exists.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(ha): wait for handoff only when a redirect went out

Address review feedback on the shutdown handoff:

- Wait for a replacement session only when the stopping replica queued a
  SessionRedirect. Single-replica gateways, supervisors without redirect
  support, and full outbound channels no longer pay the grace period.
- Close the session stream in both directions before cleanup so a
  supervisor without a redirect sees EOF and reconnects at once.
- Bound the replacement wait with timeout_at and the shared session-wait
  backoff, so the grace is a hard ceiling.
- Name the supervisor session shutdown timeout and assert at compile
  time that it leaves room after the grace period.
- Demote the sandbox when session setup fails after the owner publish,
  so a failed takeover cannot leave it Ready with no session.
- Move the supervisor reconnect delays into ReconnectBackoff so both
  loop branches reset the same way, and test reconnect sequences.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(ha): refresh gateway membership before shutdown redirects

A stopping replica redirected its supervisors using its cached ring,
which the membership worker refreshes only every 10s. During a rolling
update the replacement for the previous pod is often only seconds old,
so the cached ring missed it and still named the departed pod. The
redirect then went nowhere or to a dead pod, whose 10s connect timeout
outlasted the handoff grace, and the sandbox was demoted mid-roll.

Re-read live membership, bounded to 2s, just before session shutdown
computes redirects.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(ha): hold ownership through the shutdown handoff

The stopping replica released its owner record before waiting for the
replacement session. A reconciliation sweep on another replica in that
window saw no owner, and with the Kubernetes driver leaving readiness
to the gateway it demoted the sandbox to Provisioning anyway. Keep the
record until a different session supersedes it or the grace expires,
then release it only if it is still ours.

Also from review:
- Keep polling the owner record through store errors until the deadline.
- Skip the handoff wait for sandboxes that are not Ready.
- Re-read membership in the membership worker's shutdown branch and
  await the worker before session shutdown, so the worker stays the only
  writer of the ring and peer map.
- Share the remove, release, and demote steps of failed session setup.
- Test redirect queuing directly, including unsupported supervisors and
  a full outbound channel.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

---------

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-10-08 05:18:10 +00:00
Shiju ffb530671f fix(providers): prepare dynamic credentials across HTTP relay paths (#4152)
* fix(providers): bind multi-route grants to admitted endpoints

Carry gateway-derived endpoint owners through policy and credential delivery. Select grants only from owners that admit the current request, refresh provider snapshots per request, and reject superseded installations before forwarding.

Cover persistent routes, overlapping owners, denial, refresh, cache expiry and recovery with focused tests and a Podman regression.

Signed-off-by: Shiju <shiju@nvidia.com>

* refactor(providers): simplify admitted-endpoint grant selection

Share admitted endpoints between forwarding and credential authorization. Select grants from the live snapshot without rebuilding a locked map, reuse profile policy construction, and preserve freshness guards when grants are removed.

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(providers): preserve trusted owners across policy approval

Sanitize authored owner metadata before composing effective policies, preserve derived owners through proposal approval, and normalize advisor endpoint comparisons. Account for gateway-only owner fields in SDK coverage and schema inventory.

Signed-off-by: Shiju <shiju@nvidia.com>

* refactor(providers): share dynamic credential key layout and require a snapshot for admission

Move the endpoint-bound credential key layout into openshell-core. DynamicCredentialKey encodes the key the gateway builds, and the supervisor reads its endpoint selector, credential identity, and revision-scoped form through the same module instead of parsing tab-separated strings in three places. Revision scoping becomes ProviderCredentialSnapshot::scoped_key.

inject_for_admitted_owners now takes the pinned snapshot directly. The L7 admission path never acquires a grant from the live credential map, which production populates only alongside a snapshot. L4 forwarding keeps the selector-only path through inject_if_needed. Both share one acquire-and-rewrite step.

Delete two relay tests whose behavior the Rego admission tests already prove. Add a test that grants for different headers are acquired only for admitting owners.

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(providers): share credential preparation across HTTP relays

Apply endpoint-owned grants after policy and request middleware in REST,
GraphQL, JSON-RPC/MCP and inspected plaintext forwarding. Preserve live
provider and policy checks through guarded writes, reject missing owner
metadata, and reject profiles that cannot inject dynamic HTTP credentials.

Cover production stream dispatch, multiple endpoints, failure handling and
plaintext ownership with focused regressions. Document upgrade ordering.

Closes #3657

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(providers): validate dynamic credential key encoding

Reject control characters at profile import and key construction. Propagate
invalid key errors through gateway environment resolution and reject malformed
key layouts before token acquisition while preserving supported legacy keys.

Consolidate duplicate relay success tests and share forward-proxy setup without
removing ownership, canonicalization, refresh or guarded-write assertions.
Document token-grant endpoint authority.

Signed-off-by: Shiju <shiju@nvidia.com>

---------

Signed-off-by: Shiju <shiju@nvidia.com>
2026-10-07 23:48:40 +00:00
Eric Curtin a4954d3a99 fix(auth): match Bearer scheme case-insensitively (#4187)
* fix(auth): match Bearer scheme case-insensitively

Signed-off-by: Eric Curtin <eric.curtin@docker.com>

* docs(auth): note Bearer scheme is case-insensitive

Signed-off-by: Eric Curtin <eric.curtin@docker.com>

---------

Signed-off-by: Eric Curtin <eric.curtin@docker.com>
2026-10-07 23:00:07 +00:00
Shiju 8aa5846d7a fix(cli): preserve existing directories during single-file upload (#4177)
* fix(cli): preserve existing directories during single-file upload

Inspect remote destination types before choosing archive entry names.
Keep directory and source symlinks intact, recheck detected type changes,
and report the resulting path. Add real tar and live conformance coverage.

Closes #4175

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(cli): box upload future in sandbox_create

Boxing the upload future keeps the sandbox_create future under the clippy::large_futures threshold after sandbox_upload_planned began returning the uploaded path.

Signed-off-by: Shiju <shiju@nvidia.com>

---------

Signed-off-by: Shiju <shiju@nvidia.com>
2026-10-07 22:20:46 +00:00
Matthew Grossman 1fff8b97b9 feat(podman): honor OCI image working directories (#3982)
* fix(supervisor): report rejected OCI working directories as WorkspaceValidationFailed

The RFC 0012 split dropped the supervisor exit status that drivers map to
the WorkspaceValidationFailed condition. An image WORKDIR the sandbox
identity cannot use then surfaced as ControlSupervisorStartFailed on
Docker and as a signal kill on Podman.

The supervisor now exits with the reserved status when the sandbox rejects
the image working directory. Docker maps that supervisor exit during
readiness and monitoring, and Podman maps the supervisor companion's exit
instead of the workload's.

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* feat(podman): honor OCI image working directories

Podman sandboxes now use a custom OCI image WORKDIR as the workspace and
as the working directory for agent commands, matching Docker. Empty, /,
and /sandbox values keep the managed /sandbox workspace volume.

A custom workspace stays in the image's container filesystem: no
workspace volume, no archive upload, and no root setup step. Resolve the
image ID, user, environment, and working directory from one pinned
inspection, validate the workdir with the shared OCI rules, and reject
Podman control-path overlaps and image volumes or driver mounts that
cover it. The runtime starts from / and passes the resolved path to agent
commands with --workdir.

Add podman_oci_identity to the Podman CI test list.

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* fix(drivers): preserve custom workspace failures across supervisor cleanup

Check the Podman custom workspace rather than the absent managed volume in the OCI identity E2E. Classify Bollard wait errors by exit code and record Docker workspace failures before removing the supervisor so readiness retains the specific failure reason.

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* fix(docker): use sandbox ID for recorded startup failures

Keep the workload container ID for Docker inspection and use the sandbox ID to retrieve the monitor failure. Test with distinct identifiers so the lookup cannot accidentally pass with a container ID.

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* refactor(podman): keep WORKDIR change focused on workspace behavior

Remove the new cross-driver workspace failure classification path and its monitor/readiness workaround. Preserve Podman WORKDIR resolution, mount safety, and workspace rejection without requiring a distinct condition reason.

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* refactor(podman): simplify inspected image metadata and test fixtures

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* code review

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

---------

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>
2026-10-07 21:29:58 +00:00
Divesh 834b79a8c2 feat(server): place supervisor sessions by consistent hash (#3661)
* feat(server): place supervisor sessions by consistent hash and hand them off on shutdown

Signed-off-by: divesh <dgude@nvidia.com>

* feat: route long-lived sandbox connections to the owning gateway replica

Signed-off-by: divesh <dgude@nvidia.com>

* feat(helm): add opt-in per-replica GRPCRoute routing

Signed-off-by: divesh <dgude@nvidia.com>

---------

Signed-off-by: divesh <dgude@nvidia.com>
2026-10-07 05:01:32 +00:00
John T. MyersandMatthew Grossman 9a6148fc98 docs(observability): correct supervisor log access and retention (#4243)
* docs(observability): correct supervisor log access and retention

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(observability): clarify runtime log sources and retention limits

Remove Docker copy guidance for the live supervisor tmpfs mount while retaining the Podman example. Distinguish sandbox-runtime security events from supervisor JSONL and gateway log streams, and qualify independent three-file retention.

Validation: mise run docs, Markdown lint, and git diff --check passed. Fern reported the same three existing unrelated warnings.
Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

---------

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>
Co-authored-by: Matthew Grossman <mgrossman@nvidia.com>
2026-10-06 21:53:30 +00:00
Jason T. Greene b3a9bd85fb perf(server): drop per-connection session tokens for TCP forwards (#3734)
`openshell forward service` minted an SSH session token before every
forwarded TCP connection and revoked it afterwards: two store commits per
connection. The token added nothing on that path. `ForwardTcp` already
authenticates the caller and authorizes it against the sandbox's workspace
on every stream before it looks at the token, the relay to the supervisor
is opened with the sandbox id and target only, and the token is never
forwarded, audited, or visible to the target service. The mechanism exists
for `openshell sandbox ssh`, where the process that opens the stream is an
ssh ProxyCommand holding nothing but the token. Reusing it per TCP
connection put a store write on the connect path and serialized concurrent
forwards on commit latency: #3494 measured the symptom, and #3543 made the
commits cheaper, but each one still holds SQLite's writer lock for an
fsync, so connection setup under a burst stayed linear in the number of
concurrent connections.

Let `target.tcp` streams omit `authorization_token`. The gateway admits
them on the already-authorized principal, counts them against the same
per-sandbox connection cap, and touches no store. `target.ssh` streams keep
requiring the token. A token supplied with a TCP target is still validated
and counted per token, so an older CLI against a new gateway is unchanged.
The CLI stops minting and revoking a session per forwarded connection;
against a gateway that predates this change it recognizes the
`authorization_token is required` rejection once and falls back to
per-connection tokens for the rest of that forward.

Tests cover token-less TCP admission and slot release, SSH targets still
rejected without a token, a supplied token still validated, and the
per-sandbox cap for token-less forwards. CLI integration tests run
`service_forward_tcp` against a mock gateway: token-less inits echo data
with no CreateSshSession or RevokeSshSession call, and a gateway that
rejects the empty token is detected once, after which every connection in
that forward mints and revokes its own token. Architecture and security
docs describe which targets carry a token, and the per-token connection
limit now reads 3, matching the gateway.

Signed-off-by: Jason T. Greene <jason.greene@redhat.com>
2026-10-06 21:51:35 +00:00
Eric Curtin 3082a9ad7d fix(cli): plan sandbox uploads before provisioning (#4193)
Reject bad uploads before create. Share the upload transfer helper.

Signed-off-by: Eric Curtin <eric.curtin@docker.com>
2026-10-06 21:48:48 +00:00
Mike Nguyen 9fd41e6bfd fix(ssh): persist sandbox host identities (#4094)
* fix(ssh): persist sandbox host identities

Store each sandbox's Ed25519 host key in the gateway credential store and
deliver it only to the supervisor. Preserve identity across restarts,
delete owned credentials with the sandbox, and expose the public SHA256
fingerprint through sandbox and SSH-session APIs and client SDKs.

Cover credential ownership, cancellation, deletion retries, client
compatibility, and pinned SSH connections through lifecycle transitions.

Closes #3835

Signed-off-by: Mike Nguyen <miken@nvidia.com>

* fix(compute): clean up failed sandbox SSH identity creation

Signed-off-by: Mike Nguyen <miken@nvidia.com>

* test(ssh): wait for sandbox deletion before name reuse

Signed-off-by: Mike Nguyen <miken@nvidia.com>

---------

Signed-off-by: Mike Nguyen <miken@nvidia.com>
2026-10-06 20:09:20 +00:00
alangou 33869a1df5 fix(images): update gateway base and allow supervisor base override (#4236)
Signed-off-by: Adrien Langou <alangou@nvidia.com>
2026-10-06 13:06:12 +00:00
Drew Newberry 12cec59bf4 fix(sandbox): accept local connections natively on loopback-confined sockets (#4150)
* fix(sandbox): accept local connections natively on loopback-confined sockets

On kernels without SECCOMP_FILTER_FLAG_WAIT_KILLABLE_RECV (RHEL 9 / RHCOS
5.14), the broker cannot safely write an accepted peer address into
workload memory, so accept/accept4 with a peer-address buffer failed with
EOPNOTSUPP. Static binaries and Go servers, which issue the raw syscall,
could not accept connections at all.

Move local acceptance into the kernel and replace per-accept inspection
with standing kernel confinement:

- Bind every broker-created TCP/UDP socket to the loopback device before
  injection and verify the binding. Accepted sockets inherit it, so they
  can neither receive routed ingress nor emit routed egress.
- Stop notifying accept/accept4 and remove the accept workers, the
  SIGUSR2 accept-interrupt monitor, and the 64-accept ceiling.
- Continue getpeername natively for every descriptor except relayed
  connections.
- Deny interface-selection socket options and MSG_FASTOPEN sends from
  scalar syscall arguments, so confinement does not depend on the
  capability state of the user namespace that owns the network namespace
  and covers unregistered descriptors.
- Drop loopback-interface ingress on the non-loopback TCP control
  listener before it listens, so a workload cannot reach it by
  reconnecting a natively accepted socket.
- Mark descriptors above stdio close-on-exec in every workload pre_exec
  path.
- Require an active confinement probe at qualification and carry it as
  required authenticated audit evidence.

Document OpenShift 4.19 as the minimum release: RHCOS kernels for 4.16
through 4.18 are built without Landlock.

Closes #4058

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): close review gaps in native local accept confinement

- Reject a TCP boundary control listener on a loopback address,
  including IPv4-mapped loopback. Production drivers bind the
  unspecified address; a loopback listener would be reachable from
  workload sockets the broker does not track.
- Extend the confinement probe to IPv6 and prove that an accepted
  socket keeps its loopback binding after an AF_UNSPEC disconnect.
- Pin the accepted-peer contract with tests: a loopback-bound listener
  admits only clients in the sandbox network namespace, and workload
  sockets cannot bind a non-loopback source address.
- Correct the support matrix: accepted connections are bounded by
  per-process descriptor limits, the runtime PID limit, and the sandbox
  memory limit, not by a broker limit.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(sandbox): replace unsafe socket calls with socket2 and rustix

Socket confinement now uses socket2 for device binding and the ingress
filter, and rustix for interface lookup and AF_UNSPEC disconnect, so the
module contains no unsafe code. The control-listener filter is no longer
locked: no safe API exposes SO_LOCK_FILTER, and the listener descriptor
never leaves the trusted sandbox process, which marks every descriptor
above stdio close-on-exec before running workload code.

Tests added for native accept use rustix and socket2 instead of raw libc
calls. rustix issues accept4 and getpeername as raw syscalls, so the
direct-syscall test keeps its meaning. The close-on-exec sweep test runs
in a re-executed test process instead of a forked child.

The close_range(CLOSE_RANGE_CLOEXEC) syscall in the pre_exec hook remains
the only unsafe added by this branch; neither rustix nor nix wraps it.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): allowlist workload socket families

The workload seccomp filter denied only AF_PACKET, AF_BLUETOOTH, and
AF_VSOCK, and the broker continued socket() for every non-INET family.
Several protocol families, including AF_RXRPC, AF_SMC, and AF_KCM, carry
traffic over kernel-owned sockets that the broker never creates and that
are not bound to loopback.

Allow only AF_UNIX, AF_NETLINK (still limited to NETLINK_ROUTE), and the
brokered AF_INET/AF_INET6 families, and restrict socketpair(2) to
AF_UNIX. Both decisions use the scalar domain argument. The broker
independently refuses non-INET families other than AF_UNIX and
AF_NETLINK with EAFNOSUPPORT.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(sandbox): remove legacy read-only mode

After native local accept, only two broker paths still wrote into
workload memory: getpeername on relayed connections and the per-message
lengths of a first sendmmsg to the DNS relay. Legacy mode existed only to
refuse those writes on kernels without WAIT_KILLABLE_RECV, and the
getpeername refusal broke CPython TLS on RHEL 9.

The broker now never writes workload memory:
- getpeername is no longer mediated and reports the kernel peer on every
  kernel; relayed connections report the loopback relay address.
- A first send to the DNS relay pins the broker's socket copy to the
  relay and continues the syscall, so the kernel performs the send and
  writes any per-message results.

Remove the listener mode, the task-memory write path and probe, and the
task_memory_write, cancellation, and task_memory_writes_disabled audit
evidence fields. WAIT_KILLABLE_RECV is still used when available and is
reported for diagnostics, but no mediation decision depends on it.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(sandbox): drop WAIT_KILLABLE_RECV and make handlers restart-safe

The broker no longer writes workload memory, so killable notification
waits only reduced how often a signal restarts a notified syscall. RHEL 9
kernels never had them, so the handlers must tolerate restarts anyway.
Install the listener without the flag on every kernel instead of
special-casing newer ones.

Handlers now check that the notification is still live immediately
before each side effect (bind, connect of the retained socket, listen,
and the first DNS send), and answer a restarted operation the broker
already completed as the kernel would: a repeated TCP connect returns
EISCONN, repeating the same UDP association succeeds, a repeat of a
completed bind succeeds, and a failed relay reports its errno.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): finish teardown only when every workload descendant has exited

Teardown waited only for registered process groups to disappear. A root
is unregistered once reaped, so a descendant that ignored SIGTERM could
outlive it while termination reported success and never sent SIGKILL,
violating the bounded-termination requirement.

The sandbox now becomes a child subreaper when it is not PID 1, so
orphaned descendants stay in its tree and are reaped. Termination waits
until no live descendant remains, and every scanned process is signalled
through a pidfd after confirming its start time, so a reused PID is never
signalled.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): keep a frozen workload from resuming itself

Freezing stops every workload process with SIGSTOP. tgkill and
rt_tgsigqueueinfo were not mediated, and mediated kill, rt_sigqueueinfo,
and tkill passed SIGCONT through, so a workload process that was not yet
stopped could resume the others while the supervisor recovered.

Mediate tgkill and rt_tgsigqueueinfo like tkill: refuse targets in the
sandbox thread group, report a thread outside the named group as
missing, and continue otherwise. The boundary marks the broker frozen
before stopping the workload and clears it after resuming; while frozen,
workload requests to send SIGCONT fail with EPERM.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): send mediated DNS datagrams from the broker socket

The native UDP send path re-ran the workload's own sendmsg, so
per-message ancillary data such as IP_PKTINFO rode along; only the
loopback destination contained a routing override.

The broker now reads the datagram and sends it from its retained,
relay-connected socket, which it builds without ancillary data, so a
per-message override cannot redirect the packet. It writes nothing back
into workload memory. A send carrying control data is refused with
EOPNOTSUPP rather than silently stripped.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): keep sandbox control variables out of the workload environment

The canonical process inherits the sandbox environment so the image's
own ENV reaches the workload, then removed only a denylist of credential
variables. Other variables in the reserved OPENSHELL_ namespace, such as
the serialized user environment and the log level, still reached the
workload.

Remove every inherited OPENSHELL_ variable before applying the declared
environment, and restore OPENSHELL_SANDBOX=1. The image's ordinary ENV
and declared variables are unaffected; declared variables cannot use the
reserved namespace.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): keep /proc read-only for GPU workloads

GPU mode granted read-write access to all of /proc so CUDA's cuInit could
write thread names to /proc/<pid>/task/<tid>/comm.

Keep /proc read-only. Open mediation now serves a writable open of the
caller's own thread comm file: the broker opens it and injects the
descriptor, with no syscall continued. The kernel accepts a comm write
only from the target's own thread group, so a substituted path or reused
thread ID cannot rename another process's thread. Any other /proc write
is left to Landlock, which denies it.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): keep the broker socket for connected UDP DNS sockets

Sending DNS datagrams from the broker's socket required the broker to
keep its copy, but the connect paths to the relay and to loopback still
released it. glibc connects the resolver socket and then sends A and
AAAA together with sendmmsg, which is mediated, so resolution failed.

Release the copy on those connect paths only for TCP. Add a regression
test for connect followed by sendmmsg that fails fast rather than
hanging.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): keep slow loopback connects from stalling mediation

The loopback connect path held the registry lock while polling a TCP
connect for up to five seconds on the single notification dispatcher.
A workload connecting to a busy local listener stalled every other
mediated syscall, including opens and signals.

Connect a duplicate of the retained socket without holding the lock. A
nonblocking socket gets the native EINPROGRESS and the kernel completes
the handshake on the shared socket; a blocking socket waits on a bounded
worker thread.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): contain network broker handler panics

A panic in a notification handler unwound the single broker thread while
the health flag still read healthy, so every later blocked workload
syscall hung until the sandbox was killed.

Run each dispatch under catch_unwind: a panicking handler fails only that
syscall with EIO and the broker keeps mediating. If the broker thread
ever exits, mark it unhealthy so dependent operations fail closed instead
of blocking.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): reject ancillary data on DNS sends instead of broker-sending

Sending mediated DNS datagrams from the broker's own socket required
keeping a broker handle for each DNS socket's whole life, which regressed
real name resolution through reclaim and resource accounting that the
mock-based unit tests did not exercise.

Revert to continuing the kernel send, but refuse a send that carries
ancillary control data (msg_controllen != 0) at read time, so a
per-message routing override such as IP_PKTINFO cannot ride a mediated
DNS send. A loopback destination contains an override that races the
check. Full broker-side UDP mediation is left to a separate change with
deployment e2e.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): address review findings across the hardening changes

Fixes:
- The blocking local-connect worker no longer toggles O_NONBLOCK on the
  open file it shares with the workload; it waits with a plain connect.
- tgkill and rt_tgsigqueueinfo are notified only when they send SIGCONT,
  so ordinary thread signals such as Go preemption stay in the kernel.
  The frozen flag is re-read immediately before delivery.
- A shell redirect to the caller's own thread comm file works again;
  O_CREAT and O_TRUNC are no-ops there and O_EXCL returns EEXIST.
- The teardown scan is a linear walk and fails closed when /proc cannot
  be read.

Simplifications:
- Remove tests that need infrastructure outside the repo or prove
  nothing: the topology harness tests, the static-server benchmark, the
  direct-syscall accept test, the comm rename test without Landlock, a
  serde-default test, and the test-only panic hook in setsockopt.
- Drop the unused Failed connect outcome, the redundant read-back after
  binding to loopback, redundant OPENSHELL_SANDBOX settings, and stale or
  duplicated comments; reuse the reserved environment prefix constant.
- Tighten the support matrix and OpenShift wording.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(cli): wait for input when piped exec stdin is nonblocking

Processes that inherit the same stdin share one open file description,
so any of them can make it nonblocking for all. `sandbox exec` then read
no input yet, got EAGAIN, and failed with "Resource temporarily
unavailable (os error 11)". Parallel e2e tests inherit the runner's
stdin, which made credential_gating fail intermittently.

The stdin reader now blocks in poll() until input or end of file arrives
instead of treating EAGAIN as fatal, and it leaves the shared descriptor's
flags alone. The e2e exec helper also stops inheriting the runner's
stdin, since its commands take no input.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): fix macOS lint and older-glibc test linking

Import HashSet only in the Linux-only process scan, and call gettid
through the raw syscall in a test, since the libc wrapper needs glibc
2.30 or newer.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(sandbox): clarify socket peer address reporting

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-10-06 04:08:04 +00:00
Jim Meyer fb8f6c0885 chore(agents): resolve contributor guidance review gaps (#4002)
* chore(agents): resolve contributor guidance review gaps

Signed-off-by: Jim Meyer <jimeyer@nvidia.com>

* docs(agents): anonymize contributor branch example

Signed-off-by: Jim Meyer <jimeyer@nvidia.com>

---------

Signed-off-by: Jim Meyer <jimeyer@nvidia.com>
2026-10-06 02:21:04 +00:00
krishicks d1e8f44a06 feat(supervisor): export OTLP traces from sandbox supervisors (#3977)
When the gateway exports OTLP traces, compute drivers pass the gateway's
endpoint to supervisors as OPENSHELL_OTLP_ENDPOINT, along with TRACEPARENT
from the operation that launched them. The Kubernetes, Docker, Podman, and
VM drivers all receive the endpoint from the gateway; driver TOML cannot
set it. On Podman Machine, the Podman driver points a loopback endpoint
at host.containers.internal, since supervisor loopback is the VM's.
The supervisor exports spans as openshell-supervisor, tagged with
the sandbox ID, and flushes them before exiting.

supervisor.startup joins the sandbox's creation trace and covers image
policy discovery, policy load, and boundary attach, confirm, agent start,
and access start. Supervisor calls to the gateway carry W3C trace context,
so the gateway's server spans nest under them.

Each egress connection emits supervisor.egress.connect with authorize,
resolve, and dial children. These spans use DEBUG level because every
outbound connection starts its own trace. OpenShell spans export at INFO,
or at the sandbox log level when it is debug or trace; spans from other
libraries export at INFO.

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-10-06 00:23:14 +00:00
Shiju 8579dfb301 feat(providers): add multiple tokens to one request (#4047)
* feat(providers): compose dynamic credential grants per request

Select the most-specific grant independently for each protected header,
preserving each credential's issuer, audience and cache identity. Resolve
all selected grants before rewriting the request, replace agent-supplied
header copies, and fail closed on collisions or acquisition failures.

Allow distinct-header compositions in gateway validation and redact raw
issuer errors. Cover independent cache entries, atomic TLS relay failure,
header replacement and concurrent request isolation; document the profile
contract.

Fixes #3320

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(providers): correct dynamic grant CI checks

Use if-let for a grant result whose error payload is deliberately ignored,
and name the empty query map type in the regression fixture. Preserve grant
acquisition, failure redaction and request atomicity. Update the token-exchange
failure test to require the sanitized error instead of raw issuer text.

Signed-off-by: Shiju <shiju@nvidia.com>

---------

Signed-off-by: Shiju <shiju@nvidia.com>
2026-10-05 23:15:46 +00:00
Eric Curtin dfef088bc3 fix(docker): generate gateway JWT keys in compose quickstart (#3838)
* fix(docker): generate gateway JWT keys in compose quickstart

Fixes #2891

Signed-off-by: Eric Curtin <eric.curtin@docker.com>

* docs(docker): add compose mTLS steps

Signed-off-by: Eric Curtin <eric.curtin@docker.com>

---------

Signed-off-by: Eric Curtin <eric.curtin@docker.com>
2026-10-05 22:00:39 +00:00
Ignas Baranauskas 3592a482c5 fix(kubernetes): wait for OpenShift SCC annotations (#4054)
Signed-off-by: Ignas Baranauskas <ibaranau@redhat.com>
2026-10-05 16:34:47 +00:00
Emilien Macchianddivesh 8b3cc3fdc0 feat(server): add gateway capacity metrics and optional HPA (#3978)
* feat(server): add gateway capacity metrics and optional HPA

Expose per-replica supervisor sessions, pending relay capacity, relay
rejections and claim latency, and outbound peer request outcomes and
latency. Use bounded labels and Prometheus histograms for the new latency
metrics while preserving existing summary metrics.

Add an optional Helm HPA with external-database and resource validation,
conservative scale-down defaults, and support for custom metrics. Keep
certificate hook pods outside gateway workload selectors.

Document per-pod scraping, scaling limits, upgrade behavior, and
PostgreSQL connection sizing.

Part of #3528

Signed-off-by: Emilien Macchi <emacchi@redhat.com>

* feat(server): unify routing metrics and clarify replica capacity

Combine local relay setup and outbound peer requests in one counter, labeled by operation, route, target, outcome, and gRPC status. Preserve peer latency metrics.

Rename the rejection reason from global_capacity to replica_capacity to reflect the per-replica relay budget. Keep capacity limits unchanged.

Update tests and documentation.

Signed-off-by: divesh <dgude@nvidia.com>

* fix(server): align routed request metric labels and outcomes

Count a local relay as successful only when its supervisor claims it, the
same event that answers a peer relay on the owner, and rename the outcomes
to local_error and remote_error so they say where an attempt failed. Label
the peer latency histogram by operation, like the routed request counter,
and rename the target label to relay_kind so it does not read as the
Prometheus scrape target.

Rename RelayCapacity.global to per_replica, drop the per-sandbox relay
capacity gauge, which no per-sandbox series can pair with, describe the
latency histogram as peer-only, and restore the note that unavailable
spikes are expected during rollouts.

Part of #3528

Signed-off-by: Emilien Macchi <emacchi@redhat.com>

---------

Signed-off-by: Emilien Macchi <emacchi@redhat.com>
Signed-off-by: divesh <dgude@nvidia.com>
Co-authored-by: divesh <dgude@nvidia.com>
2026-10-05 15:38:40 +00:00
Simon Scatton 07a05f856a fix(ci): gate tagged publication on qualification (#4198)
* fix(ci): gate tagged publication on qualification

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* test(ci): remove release publication regression scaffolding

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* refactor(ci): separate snap packaging from publication

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

---------

Signed-off-by: Simon Scatton <sscatton@nvidia.com>
2026-10-05 14:00:00 +00:00
Evan Lezar e9271cb313 feat(gateway): support runtime image env overrides (#3504)
Closes #3502

Resolve trusted OCI runtime images with driver TOML, process environment, and compiled-default precedence across Docker, Podman, and Kubernetes.

Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-10-05 12:23:13 +00:00
ShijuandJohn Myers 71c3cd957a fix(gateway): check launch signing before VM image preparation (#4034)
* fix(vm): validate launch credentials before preparation

Negotiate a launch-authentication requirement and reject incomplete gateway
configuration before driver validation, archive consumption or persistence.
Validate replacement VM credentials before stopping active compute and
prefer explicit signer configuration over local discovery.

Closes #3949. Part of #3955.

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(vm): validate saved launch credentials before restore side effects

Signed-off-by: Shiju <shiju@nvidia.com>

* chore(gateway): satisfy launch preflight lint checks

Signed-off-by: Shiju <shiju@nvidia.com>

* docs(gateway): complete the MicroVM launch-signing example

Include the gateway JWT paths in the standalone VM configuration and show
the output directory required by local certificate generation. Explain
the distinction between configuration preflight and launch-key validation.

Refs #3949. Part of #3955.

Signed-off-by: Shiju <shiju@nvidia.com>

* test(vm): account for preparation state in launch credential fixture

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: Shiju <shiju@nvidia.com>
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-10-03 21:22:33 +00:00
Shiju 8983642e28 fix(gateway): give image preparation its own deadline (#4038)
* feat(server): separate image preparation and admission deadlines

Give sandbox image preparation its own deadline and start the admission
deadline after preparation finishes. Persist both phases across gateway
restarts and show recovery guidance for the phase that expired.

Preserve the current service authorization schema and regenerate the Go
bindings with the preparation timestamps.

Fixes #3952
Related to #3955

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(cli): simplify preparation timeout fallback selection

Use lazy Option fallbacks while preserving timeout messages and retained
sandbox behavior.

Signed-off-by: Shiju <shiju@nvidia.com>

* docs(server): clarify admission timer prerequisites

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(compute): enforce deadlines during initial sandbox create

Release stalled create operations after preparation expires and preserve
the timeout diagnosis across late driver results. Keep failed-create cleanup
bound to its original attempt so another replica can retry safely.

Signed-off-by: Shiju <shiju@nvidia.com>

* test(server): satisfy deadline regression lints

Drop the create-error mutex guard before matching its cloned value and use
idiomatic iteration and timeout matching in the deadline fixtures.

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(server): retain ownership of pending provisioning operations

Keep submitted create and start requests alive after caller cancellation,
monitor failure, or preparation timeout. Persist request ownership before
dispatch and retain staged uploads while the driver response is pending.

Require timeout cleanup to stop compute after driver settlement without
discarding an active cleanup claim. Fence late result handling against newer
operations, preserve failed-start recovery, and defer automatic restart while
another request owns the sandbox. Expose pending ownership in CLI JSON.

Add ordered multi-replica and cancellation regressions for the review findings.

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(openshell): preserve tracing and accept ready create responses

Carry the request span into the detached provisioning worker so compute
driver calls remain attached to their parent trace after task handoff.

Update the compensation regression to require no backend DELETE when the
durable cleanup claim fails, matching the operation ownership requirement.

Accept the gateway's current Ready snapshot when a sandbox becomes ready
before CREATE returns. Do not require the client to observe an earlier
provisioning phase. Cover the Ready-only watch and command attachment.

Signed-off-by: Shiju <shiju@nvidia.com>

---------

Signed-off-by: Shiju <shiju@nvidia.com>
2026-10-03 20:48:18 +00:00
Thota Shashank a2429fcdcd fix(tui): preserve quoted post-create command arguments (#4137)
Signed-off-by: Thota shashank <thotashashank302@gmail.com>
2026-10-03 20:32:54 +00:00
Shiju 1c123a4542 feat(gateway): check VM host tools during config preflight (#4037)
* feat(gateway): validate VM filesystem tools during preflight

Check required local VM tools through config preflight and share executable
resolution with VM image operations. Report selected paths and actionable
errors without creating gateway or sandbox state.

Bound probe output and execution time, and clean up probe descendants on
interruption. Preserve pure static validation and skip local tool checks
for remote driver endpoints and unrelated drivers.

Fixes #3951
Related to #3955

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(gateway): stabilize filesystem preflight checks

Combine identical filesystem-tool error arms and normalize rendered
diagnostics in command tests so terminal wrapping preserves assertions.
Describe driver TLS validation without depending on removed guest fields.

Signed-off-by: Shiju <shiju@nvidia.com>

* test(gateway): serialize preflight fixture paths as TOML

Keep temporary paths quoted and escaped through the TOML serializer
instead of relying on Rust Debug formatting.

Signed-off-by: Shiju <shiju@nvidia.com>

---------

Signed-off-by: Shiju <shiju@nvidia.com>
2026-10-03 20:20:02 +00:00
Shiju d676e036a4 fix(vm): reject conflicting workload identity selectors (#4036)
* fix(vm): enforce the configured workload identity

Reject conflicting policy users and groups before VM image preparation
and before guest attach or process startup changes state. Validate
supervisor policy updates against the protected VM workload identity.

Preserve the gateway CA transport and capability-free sandbox launcher.

Signed-off-by: Shiju <shiju@nvidia.com>

* test(sandbox): clarify VM identity rejection fixtures

Name invalid user and group fixtures distinctly and move the final
workload identity into its group mismatch test.

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(supervisor): align VM identity startup with current APIs

Pass the optional rejection-log key for VM identity failures and keep
generic startup-write regressions free of VM identity constraints.

Repair the call sites after the branch rebase so the identity and cleanup
proposals compile against the current startup helpers.

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(vm): restore inactive sandbox workload identity

Recover the persisted overlay owner before publishing stopped and terminal
sandboxes. Keep resources manageable when identity metadata is invalid.
Clarify fixed MicroVM ownership in policy-generation guidance.

Signed-off-by: Shiju <shiju@nvidia.com>

* test(vm): flush identity fixture before restart

Persist the canonical identity file before readiness and report the observed
exec, canonical and file-owner identities before comparing them.

Signed-off-by: Shiju <shiju@nvidia.com>

---------

Signed-off-by: Shiju <shiju@nvidia.com>
2026-10-03 20:03:39 +00:00
Shiju 48d9ab3d0d fix(cli): check final exec status and warn on partial input (#3803)
* feat(cli): stream non-TTY exec input before EOF

Add --stream-stdin using the existing interactive exec RPC without a PTY. Preserve separate output streams and enforce the existing 4 MiB cumulative input cap while forwarding input.

Require explicit clean stdin EOF and drain the response through its final gRPC status. Cover held-open input, limits, cancellation, trailers, and default finite-input behavior with subprocess and live sandbox regressions.

Signed-off-by: Shiju <shiju@nvidia.com>

* test(cli): distinguish exec cancellation from stdin EOF

Treat transport termination and response cancellation as separate test observations. Verify explicit stdin EOF through the shared frame writer and cover cancellation in the pinned Tonic decoder.

Signed-off-by: Shiju <shiju@nvidia.com>

* style(cli): use lazy optional stdin dispatch

Signed-off-by: Shiju <shiju@nvidia.com>

* test(cli): use imported duration in stdin EOF regression

Signed-off-by: Shiju <shiju@nvidia.com>

---------

Signed-off-by: Shiju <shiju@nvidia.com>
2026-10-02 22:10:01 +00:00
Eric Curtin a48920ac04 feat(helm): add sandbox UID and GID values (#3947)
* feat(helm): add sandbox UID and GID values

Closes #2697

Signed-off-by: Eric Curtin <eric.curtin@docker.com>

* fix(helm): reject boolean sandbox UID and GID

Signed-off-by: Eric Curtin <eric.curtin@docker.com>

---------

Signed-off-by: Eric Curtin <eric.curtin@docker.com>
2026-10-02 15:36:29 +00:00
Philippe Martin f7273e48f6 fix(providers): restore supervisor-backed GCP metadata discovery (#3973)
* fix(providers): restore supervisor-backed GCP metadata discovery

Relay the reserved metadata endpoint to the supervisor and restore project, account, and placeholder token responses from live provider state. Cover Google SDK discovery and repeated refresh with provider E2E tests.

Closes #3860

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(providers): preserve default metadata account without an email

Use the default account identifier when the optional service account email is missing or empty. Cover repeated SDK refresh for missing, empty, and configured email values with distinct providers for parallel E2E execution.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* test(sandbox): fix metadata relay lint and timeout

Signed-off-by: Philippe Martin <phmartin@redhat.com>

---------

Signed-off-by: Philippe Martin <phmartin@redhat.com>
2026-10-02 15:33:37 +00:00
alangou 36819f476d fix(cli): stop uploads when Git filtering fails or selects no files (#3957)
Signed-off-by: Adrien Langou <alangou@nvidia.com>
2026-10-02 14:24:32 +00:00
Oliver Calder 6e865df349 feat(snap): ship the standalone prover binary in the snap (#3717)
Signed-off-by: Oliver Calder <oliver.calder@canonical.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-10-02 11:51:56 +00:00
krishicks 8091f66877 feat(sandbox): write agent output to the container log (#4005)
Since #2726 the canonical main process's stdout and stderr are captured
in pipes that feed only the in-memory replay buffer used by sandbox
connect. Agent output therefore never reaches the container's own stdout
and stderr, so it is missing from kubectl logs, docker logs, and podman
logs and from anything that collects container logs. Before #2726 the
entrypoint inherited the container's descriptors and its output appeared
there.

Copy the main process's output to the launcher's stdout and stderr in
addition to the replay buffer, restoring the earlier behavior:

- Output is copied byte for byte to the matching stream from a
  forwarder thread per stream, after it is published to the replay
  buffer. When the container runtime falls behind on a stream, that
  stream's reader waits instead of dropping output, so backpressure
  reaches the agent as it did with inherited descriptors, while the
  other stream and attachments keep receiving output.
- Before the main process's exit is published, the output readers
  finish and queued output is drained to the container log, so an
  agent's final lines are not lost at shutdown. A 30 second deadline
  covers both; when it expires, readers waiting on the container log
  are released and drain the pipes into the replay buffer only, so a
  stalled container log cannot block exit reporting.
- PTY-mode processes are not copied. The terminal stream carries escape
  sequences and echoed input, and terminal commands never reached the
  container log before #2726.
- Exec, SSH, and SFTP sessions are not copied.

Launcher log lines keep their existing format and remain in the
container's stderr. They are written as whole lines, and a newline is
inserted first when the agent left stderr mid-line, so launcher and
agent lines do not merge.

The Docker and VM drivers appended the tail of the workload's output to
failure messages: Docker the workload container's log, and the VM driver
the guest console, which carries the launcher's stdout and stderr. Those
messages land in the sandbox's Ready condition and in platform events
that the gateway republishes to the sandbox event stream. With agent
output in that log, those messages would carry arbitrary agent output,
including anything sensitive the agent prints, into gateway status and
events. The supervisor starts its health endpoint only after the agent
starts, so every Docker failure path could include agent output, and the
VM driver reports one whenever the VM or host supervisor exits. Forward
only the supervisor's log tail, matching the Podman driver, which reads
the workload log solely to match fixed launcher markers and never
forwards raw workload output. The workload's output remains available
through docker logs and the VM's rootfs-console.log.

Document where main process output appears in the logging docs and the
cluster debugging skill.

Closes #3928

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-10-01 17:23:09 +00:00
Fede Kamelhar 71440b28f4 fix(policy): refresh pending proposals when the sandbox policy changes (#3923)
* fix(policy): refresh pending proposals when the sandbox policy changes

Approving, removing, or undoing a rule, or updating the sandbox policy,
changes the inputs every other pending proposal was evaluated against.
Only proposals the new policy covered were reconciled; the rest kept
their old prover result and review token. The review surface
(GetDraftPolicy) therefore showed a stale evaluation, and the first
approval of the next proposal refreshed it and failed with
FAILED_PRECONDITION, so approving proposals one after another always
failed once.

Re-evaluate the remaining pending proposals at each policy change,
reusing the cached prover result unless the proposal's inputs changed.
Approval still rejects a review token that does not match the stored
evaluation, so a reviewer holding a pre-refresh evaluation must still
refetch it.

When a refresh does happen at approval time (inputs changed between
fetch and approve), the CLI now explains that the rule was re-evaluated
and how to review it, instead of printing the raw gRPC status.

Closes #3884

Signed-off-by: fede-kamel <fkamelhar@gmail.com>

* fix(policy): make pending proposal refresh race-safe and bounded

Store refreshed evaluations with a compare-and-swap: the store re-reads
the proposal, refuses when its rule name, proposed rule, or review token
changed since the evaluation read it, copies only the evaluation fields
onto the stored record, and updates only if the payload is still the one
it read. A refresh can no longer revert a concurrent edit or observation,
and the edit path uses the same guard against a concurrent refresh.

Bound each refresh to the 32 newest pending proposals; the rest keep the
approval-time recheck, which still refuses a stale review token. Operator
decisions (approve, approve-all, remove, undo) refresh before responding.
UpdateConfig, which holds the gateway-wide sandbox sync guard, and
agent-driven auto-approval refresh in a background task instead, one per
sandbox with later changes coalesced into a single rerun.

Refs #3884

Signed-off-by: fede-kamel <fkamelhar@gmail.com>

* docs(policy): describe proposal rechecks after approvals and approve-all

Explain that approving, removing, or undoing a rule rechecks the other
pending proposals so they can be approved one after another, when the
recheck is deferred or bounded, and what rule approve reports when a
proposal changed after it was listed. Show rule approve-all in Run Your
First Agent with its security-flag behavior.

Refs #3884

Signed-off-by: fede-kamel <fkamelhar@gmail.com>

* fix(policy): refresh pending proposals after a full policy replacement

A full policy UpdateConfig (openshell policy set) re-reads the latest
revision after its atomic write, finds the revision it just committed,
and returns before reaching the pending-proposal refresh at the end of
the handler. Pending proposals kept their stale evaluation, so rule get
showed the old candidate and the next approval failed with the refresh
precondition. Schedule the background refresh right after the commit.

Refs #3884

Signed-off-by: fede-kamel <fkamelhar@gmail.com>

---------

Signed-off-by: fede-kamel <fkamelhar@gmail.com>
2026-10-01 16:24:51 +00:00
Shiju 8d418f1f62 fix(supervisor): bound pending exec stdin and cancel stalled writers (#3846)
* fix(supervisor): bound pending exec stdin and cancel stalled writers

Signed-off-by: Shiju <shiju@nvidia.com>

* docs(supervisor): separate pending stdin guidance from CLI modes

Keep the pending-input limit beside the RPC lifecycle contract so the streaming CLI documentation can merge independently.

Signed-off-by: Shiju <shiju@nvidia.com>

---------

Signed-off-by: Shiju <shiju@nvidia.com>
2026-10-01 16:23:11 +00:00
John T. MyersandJohn Myers 6e369f2396 chore(agents): simplify contributor instructions and workflows (#3987)
* chore(agents): simplify contributor instructions and workflows

Closes #3980

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* docs(contributing): scope verification to affected components

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* docs(contributing): standardize issue branch naming

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <johntmyers@users.noreply.github.com>
2026-10-01 16:20:35 +00:00
Fede Kamelhar ffcbe6280c fix(cli): start sandbox exec without waiting for piped stdin EOF (#4006)
With a non-terminal stdin, sandbox exec read stdin to EOF before it sent
the exec request. A pipe that never closes (CI runners, supervisors, agent
harnesses) blocked the CLI forever in read(2) without the gateway ever
seeing the request, and a slow producer delayed the command until EOF.

Collect piped stdin on a detached reader thread for at most 200 ms. Input
that reaches EOF within that window still travels in the single request
that older gateways need. If the pipe is still open, start the command
through the streaming RPC and forward the collected prefix plus the rest of
stdin as it arrives, closing remote stdin at EOF. The 4 MiB cap covers the
prefix and the streamed remainder together.

Closes #3993

Signed-off-by: Federico Kamelhar <federico.kamelhar@oracle.com>
2026-10-01 16:11:49 +00:00
Shiju f2901393e6 fix(supervisor): wait for repair when the gateway refuses a startup policy write (#3785)
* fix(supervisor): wait for repair when the gateway refuses a startup policy write

Startup writes the sandbox policy to the gateway in two cases: it
uploads a discovered image policy when the gateway has none, and it
writes the policy back after adding the proxy baseline filesystem paths.
When the gateway refused either write with FAILED_PRECONDITION or
INVALID_ARGUMENT, for example because the policy binds a provider that
is not attached, startup treated the refusal as a permanent error and
the supervisor exited. The sandbox never reached the ConfigurationInvalid
repair state that other startup rejections use.

Report such a refusal as a configuration rejection carrying the
gateway's message, log it once per write and error code, and keep
polling, so attaching the provider or replacing the policy completes
startup. Other error codes keep their current handling: transient codes
are retried, and permission, not-found and authentication failures
still end startup.

Skip the baseline-path write-back while a global policy is active. The
gateway refuses every sandbox policy write in that state, so startup
exited whenever a global policy lacked a baseline path. The supervisor
now adds the paths to its own copy of the policy without saving a
revision.

Signed-off-by: Shiju <shiju@nvidia.com>

* test(supervisor): stabilize startup refusal log capture

Keep a second tracing dispatcher alive while capturing startup refusal
logs. With only one dispatcher, a parallel test thread without a default
subscriber can cache Interest::never for the shared OCSF callsite after
the capture thread rebuilds the cache.

Preserve the exact log-count, diagnostic, configuration-generation and
repair assertions. Production startup behavior is unchanged.

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(supervisor): reconcile stale startup rejection reports

Refetch desired configuration immediately when a rejection report is aborted because its generation changed. Preserve acknowledged rejection pacing and all other report error handling.

Signed-off-by: Shiju <shiju@nvidia.com>

* test(supervisor): box startup repair race futures

Keep the repair regressions below the large-future lint threshold without changing their inputs, scheduling, or assertions.

Signed-off-by: Shiju <shiju@nvidia.com>

---------

Signed-off-by: Shiju <shiju@nvidia.com>
2026-10-01 16:11:05 +00:00
Oliver Calder 1ad4e428a6 fix(snap): simplify snap hooks (#3988)
* fix(snap): simplify snap hooks

The `post-refresh` hook runs after initial snap installation as well, so
there is no need to call the `install` hook from within the
`post-refresh` hook; instead, the logic can simply be moved into the
`post-refresh` hook directly, and the `install` hook removed.

Also, the existing `install` hook logic looked for an insecure
configuration, and if found, replaced the entire configuration file with
a minimal default in the current format. But OpenShell does that default
behavior without any config file, so we may as well simply remove the
configuration file entirely to keep up-to-date with the current default
behavior. Let OpenShell create a configuration file if it needs to,
rather than auto-create one via the packaging scripts.

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* fix(snap): remove the connect-plug-docker hook

The `openshell:docker` is auto-connected to the system `:docker` slot,
so there should not be a need to separately restart the gateway service
when the interface is connected.

For locally-built test snaps which were not published to the store, the
autoconnection is not made, but when the snap is installed, the gateway
will attempt to start anyway and fail to find any available compute
driver, so quickly restart until it hits the systemd start-limit, after
which systemd prevents the service from being started again. If a user
tries to manually connect their locally-built `openshell` snap to the
`:docker` slot, then the `connect-plug-docker` hook runs and triggers a
restart of the gateway, which will usually fail because the start limit
has already been hit. An error in the hook will thus cause the interface
connection to be undone, which is undesirable.

Thus, we can remove this hook entirely, and instead allow interface
connections to succeed as intended. The user still needs to manually
restart the gateway service after making a manual connection (as was the
case previously) and probably needs to `systemctl reset-failed` first,
but at least connection will succeed beforehand so they can proceed with
these steps.

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* fix(snap): set refresh-mode: endure again, with manual restart

Return to the previous behavior before commit a67567e58, where the
gateway is not stopped before refreshes. The `post-refresh` hook
now restarts the gateway if the TLS configuration was corrected, so we
don't have to enforce restarting the gateway on every refresh even when
not necessary. Thus, set `refresh-mode: endure`, and let the hook decide
when the gateway needs to be restarted.

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* fix(snap): update docs and tests to reflect snap hook changes

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* docs(snap): remove verbose explanation of snap gateway refresh behavior

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

---------

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>
2026-10-01 15:10:38 +00:00
Simon Scatton e21b7fd8cf chore(build): remove bundled Z3 support (#3275)
* chore(build): remove bundled Z3 support

Signed-off-by: Simon Scatton <sscatton@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(build): preserve vendored Z3 for local gateway artifacts

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

---------

Signed-off-by: Simon Scatton <sscatton@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-10-01 12:08:33 +00:00
Drew Newberry 021400be8a refactor(auth): separate sandbox identity from TLS (#3110)
* refactor(auth): separate sandbox identity from TLS

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(auth): clarify gateway mTLS behavior

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(auth): include workspace scope in TLS authorization checks

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(e2e): bound service auth sandbox names for large PIDs

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-10-01 04:33:25 +00:00
Derek Carr 912a077bd6 feat(service): add bearer authorization passthrough (#3796)
* feat(service): add bearer authorization passthrough

Signed-off-by: Derek Carr <decarr@redhat.com>

* docs(sdk): add service authorization migration guide

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(server): remove stale version import

Signed-off-by: Derek Carr <decarr@redhat.com>

* docs(upgrade): remove service authorization SDK guide

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(e2e): relabel provider readiness TLS mount

Signed-off-by: Derek Carr <decarr@redhat.com>

* test(e2e): stabilize exposed service routing

Signed-off-by: Derek Carr <decarr@redhat.com>

* test(e2e): support HTTPS service routing

Signed-off-by: Derek Carr <decarr@redhat.com>

---------

Signed-off-by: Derek Carr <decarr@redhat.com>
2026-09-30 20:22:46 +00:00
Shiju 374c035962 fix(network): refuse protocol upgrades on GraphQL endpoints (#3841)
* fix(network): refuse protocol upgrades on GraphQL endpoints

Refuse Upgrade headers before forwarding GraphQL-over-HTTP requests.
Share the protocol refusal table with JSON-RPC and MCP, and close
unexpected protocol switches before relaying frames.

Keep GraphQL-over-WebSocket inspection on separate WebSocket endpoints.
Cover upgrade refusal, audit mode, subscription handshakes, and ordinary
HTTP and WebSocket controls. Update the current policy documentation.

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(network): refuse GraphQL upgrades before reading bodies

Validate the HTTP head and endpoint authority before upgrade refusal, then inspect ordinary GraphQL bodies. Preserve missing-authority credential rejection after body inspection.

Signed-off-by: Shiju <shiju@nvidia.com>

---------

Signed-off-by: Shiju <shiju@nvidia.com>
2026-09-30 20:10:04 +00:00
krishicks 7caff12d3c perf(otel): stop exporting spans from steady-state polling (#3915)
Store operation spans and request spans for supervisor-polled RPCs
(GetSandboxConfig, ReportProviderReadiness) use DEBUG level, so the
default INFO filter no longer exports them. The provider credential
refresh worker opens its span only when a state has work.

Refs #2698

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-09-30 15:19:39 +00:00
Shiju 798500ccdb fix(policy): validate raw OPA settings and redact startup errors (#3788)
* test(policy): reproduce raw OPA loading gaps against the typed schema

The supervisor loads a sandbox policy in two ways: through the typed
schema (parse_sandbox_policy, then from_proto) or directly into OPA
(from_strings and from_files). The raw path fills in defaults where the
typed schema is strict, so the same policy text can produce a different
sandbox configuration, or load when it should be rejected.

Add two regression tests that fail on the current code:

- An empty filesystem_policy loads with include_workdir true through raw
  OPA and false through the typed schema. An absent stanza gives true on
  both paths and must keep doing so.
- Raw OPA accepts a string include_workdir, a non-string read_only entry,
  an unknown Landlock compatibility and an explicit null json_rpc, with
  or without a version key. The typed schema rejects each. Every case has
  a valid twin that both paths must accept.

A follow-up change makes raw loading apply the typed schema's rules.

Refs #3092.

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(policy): align raw OPA loading with typed settings

Validate raw filesystem, Landlock, and process settings with the canonical
authored schema before normalization. Preserve the absent filesystem
default while applying the present-stanza default, and canonicalize valid
Landlock enum representations before runtime evaluation.

Reject explicit null JSON-RPC options through the shared parser. Preserve
versionless and runtime OPA data, and keep rejected reloads from replacing
the active policy or advancing its generation.

Add raw-versus-typed, file-loader, and rejected-reload regressions and
document the local loading contract.

Refs #3092.

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(policy): validate raw OPA settings and redact startup errors

Validate raw network fields through the authored schema before
normalization. Preserve custom Rego data and supported runtime forms.
Apply the shared filesystem path checks and non-root identity predicate
to raw static settings.

Discard authored Rego source and nested errors from static configuration
evaluation. Cover malformed inputs, valid controls, file loading, and
rejected reloads retaining active decisions and generation.

Refs #3092.

Signed-off-by: Shiju <shiju@nvidia.com>

* test(policy): satisfy unit-returning assertion lint

Terminate the two error-assertion match arms with semicolons, as required
by Clippy. Preserve the existing checks and runtime behavior.

Refs #3092.

Signed-off-by: Shiju <shiju@nvidia.com>

---------

Signed-off-by: Shiju <shiju@nvidia.com>
2026-09-30 05:59:31 +00:00
Shiju ba16b9f2c7 fix(mcp): explain revision-scoped policy and rejections (#3850)
* fix(mcp): explain revision-scoped policy and rejections

Explain the selected-revision method set in profile output and policy docs.
Distinguish protocol and policy rejection causes and give a next step while
preserving authorization, response statuses, error codes and YAML keys.

Cover CLI serialization, revision selection, exact extension rules, deny
precedence and rejection before forwarding with focused regressions.

Signed-off-by: Shiju <shiju@nvidia.com>

* docs(mcp): correct HTTP cancellation revision support

Limit notifications/cancelled to the three 2025 revisions in the core
method matrix. State that MCP 2026-07-28 HTTP cancellation closes the
response stream, matching the runtime rejection and sessionless docs.

Signed-off-by: Shiju <shiju@nvidia.com>

---------

Signed-off-by: Shiju <shiju@nvidia.com>
2026-09-29 21:50:35 +00:00
Shiju 0ea0d31020 fix(supervisor): restore canonical stdin after connection loss (#3852)
* fix(supervisor): restore canonical stdin after connection loss

Probe idle SSH peers and enforce a receive deadline during transport I/O,
including writes blocked by a stalled relay. Release the dead attachment's
stdin lease through existing handler cleanup.

Retry denied write intent on later ordinary input without displacing a
healthy owner. Preserve explicit read-only, EOF and detach behavior, and
discard control bytes retained while input ownership was denied.

Cover half-open forwarding, blocked writes, healthy idle peers and competing
reconnects through the production supervisor frame bridge and real SSH.

Fixes #3648

Signed-off-by: Shiju <shiju@nvidia.com>

* docs(skills): describe read-only reconnect input retry

Explain what an openshell-cli user sees when automatic recovery reattaches before the supervisor closes the dead connection: the attachment reports read-only, later ordinary input retries stdin acquisition and prints `input enabled`, input typed while read-only is discarded, exit keys still detach, and an explicitly read-only viewer or a healthy owner is never affected.

Signed-off-by: Shiju <shiju@nvidia.com>

---------

Signed-off-by: Shiju <shiju@nvidia.com>
2026-09-29 20:28:55 +00:00
krishicks 33a8eac196 feat(helm): configure gateway OCSF JSONL output (#3876)
Previously, Helm installations could not enable the gateway OCSF JSONL
destination through chart values because generated `gateway.toml` omitted the
`openshell.gateway.ocsf_log` table.

Now, setting `server.ocsfLog.enabled` renders the path, optional schema
version, rotation, retention, and queue limits into gateway configuration.
Output is disabled by default. The default path, `/tmp/gateway-ocsf.jsonl`,
is writable in the gateway container with either the StatefulSet or
Deployment workload, so enabling output does not require persistent storage.
Invalid schema versions, rotation values, non-positive limits, or an empty
path while enabled fail chart rendering.

Additionally, `server.extraVolumes` and `server.extraVolumeMounts` add
operator-supplied volumes to the gateway pod, so operators who want records
to survive restarts can place the OCSF path on persistent storage without
replacing chart-generated configuration.

The gateway pod's default termination grace period rises from 5 to 30
seconds. Gateway shutdown can spend up to 10 seconds on supervisor session
cleanup before allowing 5 seconds to drain queued OCSF records, so the
5-second default risked a SIGKILL before the final records were written.
The grace period is only an upper bound: the gateway exits as soon as its
shutdown completes.

Refs #2762

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-09-29 19:23:43 +00:00
krishicksandEvan Lezar a875add234 feat(server): write gateway OCSF events to JSONL (#3264)
* feat(server): write gateway OCSF events to JSONL

Previously, gateway security activity was available only in diagnostic output,
and events not associated with a sandbox, such as TLS certificate reloads, had
no independent structured record.

Now, configuring `openshell.gateway.ocsf_log` writes every gateway-produced
OCSF record to a bounded JSONL destination independently of `RUST_LOG`. The
destination supports daily or disabled rotation, retention limits, queue bounds,
and optional schema downgrade to OCSF 1.1 or 1.3.

Additionally, existing gateway emitters (TLS reloads, service routing, and
policy approval and auto-approval audits) emit structured events, so they reach
the JSONL destination, console shorthand, and the affected sandbox's log
stream. Records identify the gateway by its configured name in `device.uid` and
`device.name`, shared across replicas, with `device.hostname` identifying the
replica and `device.os` the gateway's operating system. Metrics and warnings
expose known best-effort losses.

Refs #2762

Signed-off-by: Kris Hicks <khicks@nvidia.com>

* fix(mxc): attribute ETW events to gateway

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Kris Hicks <khicks@nvidia.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>
Co-authored-by: Evan Lezar <elezar@nvidia.com>
2026-09-29 16:19:53 +00:00
Philippe MartinandJohn Myers 9cb72baa2e feat(docker): support corporate proxy CA bundles (#3549)
* feat(docker): support corporate proxy CA bundles

Closes #3545

Validate and stage operator-owned proxy CA bundles for Docker supervisors, add corporate proxy E2E coverage, and document the trust contract.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(docker): validate proxy config on startup

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(docker): use the E2E workload image for proxy tests

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(docker): generate strict corporate proxy certificates

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(docker): surface intercepted TLS fixture errors

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(docker): drain buffered TLS proxy data

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(docker): relay intercepted HTTP deterministically

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: Philippe Martin <phmartin@redhat.com>
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-09-29 05:15:02 +00:00
Divesh 2fe5a0e19c perf(kubernetes): use a TCP readiness probe for the supervisor (#3700)
- Kubernetes now checks supervisor readiness by connecting to TCP port 5501
- Stop starting a supervisor process in every sandbox each second
- The supervisor opens the port only while its gateway session is up
- Accept IPv4 and IPv6 probes, even when net.ipv6.bindv6only is set
- Keep the health socket for Docker, Podman, and debugging
- Add tests and update the docs

Signed-off-by: divesh <dgude@nvidia.com>
2026-09-29 04:47:12 +00:00