208 Commits
Author SHA1 Message Date
Drew Newberry e1f3c82caa fix(deps): update noyalib and preserve policy null rejection (#4348)
* fix(deps): update noyalib and preserve policy null rejection

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(providers): preserve quoted YAML duration exports

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(providers): assert YAML duration semantics instead of quoting

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(yaml): preserve authored input and cross-reader string semantics

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(yaml): validate schema objects before standard decoding

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(policy): remove YAML compatibility documentation changes

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-10-09 05:04:49 +00:00
cjagwani a991b8e369 fix(sandbox): preserve exec capacity under failure and load (#4324)
* fix(sandbox): reclaim failed exec waits

Signed-off-by: Charan Jagwani <cjagwani@nvidia.com>

* test(e2e): cover exec recovery after policy denial

Signed-off-by: Charan Jagwani <cjagwani@nvidia.com>

* fix(sandbox): preserve live descriptor headroom

Signed-off-by: Charan Jagwani <cjagwani@nvidia.com>

* fix(sandbox): reclaim descriptors before exec

Signed-off-by: Charan Jagwani <cjagwani@nvidia.com>

* fix(sandbox): separate exec and socket capacity

Signed-off-by: Charan Jagwani <cjagwani@nvidia.com>

* fix(sandbox): stream Landlock baseline entries

Signed-off-by: Charan Jagwani <cjagwani@nvidia.com>

---------

Signed-off-by: Charan Jagwani <cjagwani@nvidia.com>
2026-10-09 01:26:31 +00:00
Gaizka MenendezandTaylor Mutch fe3942f2da feat(helm): migrate gateway configuration to gatewayConfig (#3384)
* feat(helm): add generic gateway TOML serializer

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* docs(architecture): define Helm gateway configuration boundary

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* feat(helm): define default gateway configuration map

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* refactor(helm): render gateway config from values map

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* test(helm): cover generic gateway TOML rendering

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* feat(helm): protect gateway config secret boundary

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* docs(helm): classify legacy gateway configuration values

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* refactor(helm): derive dual-use resources from gateway config

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* test(helm): migrate gateway config scenarios

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* refactor(helm): derive credential resources from gateway config

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* refactor(helm): migrate gateway config overlays

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* refactor(helm): remove gateway configuration shadow values

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* refactor(helm): align runtime config with resources

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* docs(helm): add gateway config migration guide

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* test(helm): validate rendered gateway config

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* test(helm): cover config resource coherence

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* fix(helm): make Kubernetes E2E deployable

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* fix(e2e): pass host aliases through gateway config

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* docs(e2e): reference gateway config host alias

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* fix(helm): enforce gateway resource ownership

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* test(e2e): cover chart host gateway input

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* fix(helm): address gateway config review findings

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* fix(helm): repair RFC 0012 migration integration

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* fix(e2e): align Kubernetes parity with RFC 0012

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* fix(helm): complete gateway config migration

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* chore(helm): add SPDX header to TOML template

* fix(helm): preserve legacy gateway configuration aliases

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* fix(helm): complete legacy gateway config compatibility

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* feat(helm): support gateway config arrays of tables

* fix(helm): preserve sandbox identity aliases

* fix(helm): address gateway config review findings

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* fix(helm): preserve gateway config migration settings

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* test(helm): validate extension arrays in CI values

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(helm): restore unconditional SPDX headers

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(helm): restore route SPDX header placement

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* docs(helm): simplify configuration and migration examples

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* docs(helm): remove redundant flag explanations

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* docs(helm): move chart configuration guidance to Kubernetes setup

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

---------

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>
Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
Co-authored-by: Taylor Mutch <taylormutch@gmail.com>
2026-10-08 18:39:59 +00:00
John T. Myers 67adcf1a3c feat(ocsf): emit full JSON records to supervisor stderr (#4323)
* feat(ocsf): emit full JSON records to supervisor stderr

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(ocsf): preserve console record boundaries

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-10-08 05:36:19 +00:00
Shiju ffb530671f fix(providers): prepare dynamic credentials across HTTP relay paths (#4152)
* fix(providers): bind multi-route grants to admitted endpoints

Carry gateway-derived endpoint owners through policy and credential delivery. Select grants only from owners that admit the current request, refresh provider snapshots per request, and reject superseded installations before forwarding.

Cover persistent routes, overlapping owners, denial, refresh, cache expiry and recovery with focused tests and a Podman regression.

Signed-off-by: Shiju <shiju@nvidia.com>

* refactor(providers): simplify admitted-endpoint grant selection

Share admitted endpoints between forwarding and credential authorization. Select grants from the live snapshot without rebuilding a locked map, reuse profile policy construction, and preserve freshness guards when grants are removed.

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(providers): preserve trusted owners across policy approval

Sanitize authored owner metadata before composing effective policies, preserve derived owners through proposal approval, and normalize advisor endpoint comparisons. Account for gateway-only owner fields in SDK coverage and schema inventory.

Signed-off-by: Shiju <shiju@nvidia.com>

* refactor(providers): share dynamic credential key layout and require a snapshot for admission

Move the endpoint-bound credential key layout into openshell-core. DynamicCredentialKey encodes the key the gateway builds, and the supervisor reads its endpoint selector, credential identity, and revision-scoped form through the same module instead of parsing tab-separated strings in three places. Revision scoping becomes ProviderCredentialSnapshot::scoped_key.

inject_for_admitted_owners now takes the pinned snapshot directly. The L7 admission path never acquires a grant from the live credential map, which production populates only alongside a snapshot. L4 forwarding keeps the selector-only path through inject_if_needed. Both share one acquire-and-rewrite step.

Delete two relay tests whose behavior the Rego admission tests already prove. Add a test that grants for different headers are acquired only for admitting owners.

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(providers): share credential preparation across HTTP relays

Apply endpoint-owned grants after policy and request middleware in REST,
GraphQL, JSON-RPC/MCP and inspected plaintext forwarding. Preserve live
provider and policy checks through guarded writes, reject missing owner
metadata, and reject profiles that cannot inject dynamic HTTP credentials.

Cover production stream dispatch, multiple endpoints, failure handling and
plaintext ownership with focused regressions. Document upgrade ordering.

Closes #3657

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(providers): validate dynamic credential key encoding

Reject control characters at profile import and key construction. Propagate
invalid key errors through gateway environment resolution and reject malformed
key layouts before token acquisition while preserving supported legacy keys.

Consolidate duplicate relay success tests and share forward-proxy setup without
removing ownership, canonicalization, refresh or guarded-write assertions.
Document token-grant endpoint authority.

Signed-off-by: Shiju <shiju@nvidia.com>

---------

Signed-off-by: Shiju <shiju@nvidia.com>
2026-10-07 23:48:40 +00:00
Matthew Grossman 1fff8b97b9 feat(podman): honor OCI image working directories (#3982)
* fix(supervisor): report rejected OCI working directories as WorkspaceValidationFailed

The RFC 0012 split dropped the supervisor exit status that drivers map to
the WorkspaceValidationFailed condition. An image WORKDIR the sandbox
identity cannot use then surfaced as ControlSupervisorStartFailed on
Docker and as a signal kill on Podman.

The supervisor now exits with the reserved status when the sandbox rejects
the image working directory. Docker maps that supervisor exit during
readiness and monitoring, and Podman maps the supervisor companion's exit
instead of the workload's.

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* feat(podman): honor OCI image working directories

Podman sandboxes now use a custom OCI image WORKDIR as the workspace and
as the working directory for agent commands, matching Docker. Empty, /,
and /sandbox values keep the managed /sandbox workspace volume.

A custom workspace stays in the image's container filesystem: no
workspace volume, no archive upload, and no root setup step. Resolve the
image ID, user, environment, and working directory from one pinned
inspection, validate the workdir with the shared OCI rules, and reject
Podman control-path overlaps and image volumes or driver mounts that
cover it. The runtime starts from / and passes the resolved path to agent
commands with --workdir.

Add podman_oci_identity to the Podman CI test list.

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* fix(drivers): preserve custom workspace failures across supervisor cleanup

Check the Podman custom workspace rather than the absent managed volume in the OCI identity E2E. Classify Bollard wait errors by exit code and record Docker workspace failures before removing the supervisor so readiness retains the specific failure reason.

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* fix(docker): use sandbox ID for recorded startup failures

Keep the workload container ID for Docker inspection and use the sandbox ID to retrieve the monitor failure. Test with distinct identifiers so the lookup cannot accidentally pass with a container ID.

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* refactor(podman): keep WORKDIR change focused on workspace behavior

Remove the new cross-driver workspace failure classification path and its monitor/readiness workaround. Preserve Podman WORKDIR resolution, mount safety, and workspace rejection without requiring a distinct condition reason.

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* refactor(podman): simplify inspected image metadata and test fixtures

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* code review

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

---------

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>
2026-10-07 21:29:58 +00:00
Mike Nguyen 9fd41e6bfd fix(ssh): persist sandbox host identities (#4094)
* fix(ssh): persist sandbox host identities

Store each sandbox's Ed25519 host key in the gateway credential store and
deliver it only to the supervisor. Preserve identity across restarts,
delete owned credentials with the sandbox, and expose the public SHA256
fingerprint through sandbox and SSH-session APIs and client SDKs.

Cover credential ownership, cancellation, deletion retries, client
compatibility, and pinned SSH connections through lifecycle transitions.

Closes #3835

Signed-off-by: Mike Nguyen <miken@nvidia.com>

* fix(compute): clean up failed sandbox SSH identity creation

Signed-off-by: Mike Nguyen <miken@nvidia.com>

* test(ssh): wait for sandbox deletion before name reuse

Signed-off-by: Mike Nguyen <miken@nvidia.com>

---------

Signed-off-by: Mike Nguyen <miken@nvidia.com>
2026-10-06 20:09:20 +00:00
Simon Scatton 8760396975 fix(compute): revalidate late exit events after restart (#4235)
Signed-off-by: Simon Scatton <sscatton@nvidia.com>
2026-10-06 15:25:53 +00:00
Drew Newberry 12cec59bf4 fix(sandbox): accept local connections natively on loopback-confined sockets (#4150)
* fix(sandbox): accept local connections natively on loopback-confined sockets

On kernels without SECCOMP_FILTER_FLAG_WAIT_KILLABLE_RECV (RHEL 9 / RHCOS
5.14), the broker cannot safely write an accepted peer address into
workload memory, so accept/accept4 with a peer-address buffer failed with
EOPNOTSUPP. Static binaries and Go servers, which issue the raw syscall,
could not accept connections at all.

Move local acceptance into the kernel and replace per-accept inspection
with standing kernel confinement:

- Bind every broker-created TCP/UDP socket to the loopback device before
  injection and verify the binding. Accepted sockets inherit it, so they
  can neither receive routed ingress nor emit routed egress.
- Stop notifying accept/accept4 and remove the accept workers, the
  SIGUSR2 accept-interrupt monitor, and the 64-accept ceiling.
- Continue getpeername natively for every descriptor except relayed
  connections.
- Deny interface-selection socket options and MSG_FASTOPEN sends from
  scalar syscall arguments, so confinement does not depend on the
  capability state of the user namespace that owns the network namespace
  and covers unregistered descriptors.
- Drop loopback-interface ingress on the non-loopback TCP control
  listener before it listens, so a workload cannot reach it by
  reconnecting a natively accepted socket.
- Mark descriptors above stdio close-on-exec in every workload pre_exec
  path.
- Require an active confinement probe at qualification and carry it as
  required authenticated audit evidence.

Document OpenShift 4.19 as the minimum release: RHCOS kernels for 4.16
through 4.18 are built without Landlock.

Closes #4058

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): close review gaps in native local accept confinement

- Reject a TCP boundary control listener on a loopback address,
  including IPv4-mapped loopback. Production drivers bind the
  unspecified address; a loopback listener would be reachable from
  workload sockets the broker does not track.
- Extend the confinement probe to IPv6 and prove that an accepted
  socket keeps its loopback binding after an AF_UNSPEC disconnect.
- Pin the accepted-peer contract with tests: a loopback-bound listener
  admits only clients in the sandbox network namespace, and workload
  sockets cannot bind a non-loopback source address.
- Correct the support matrix: accepted connections are bounded by
  per-process descriptor limits, the runtime PID limit, and the sandbox
  memory limit, not by a broker limit.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(sandbox): replace unsafe socket calls with socket2 and rustix

Socket confinement now uses socket2 for device binding and the ingress
filter, and rustix for interface lookup and AF_UNSPEC disconnect, so the
module contains no unsafe code. The control-listener filter is no longer
locked: no safe API exposes SO_LOCK_FILTER, and the listener descriptor
never leaves the trusted sandbox process, which marks every descriptor
above stdio close-on-exec before running workload code.

Tests added for native accept use rustix and socket2 instead of raw libc
calls. rustix issues accept4 and getpeername as raw syscalls, so the
direct-syscall test keeps its meaning. The close-on-exec sweep test runs
in a re-executed test process instead of a forked child.

The close_range(CLOSE_RANGE_CLOEXEC) syscall in the pre_exec hook remains
the only unsafe added by this branch; neither rustix nor nix wraps it.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): allowlist workload socket families

The workload seccomp filter denied only AF_PACKET, AF_BLUETOOTH, and
AF_VSOCK, and the broker continued socket() for every non-INET family.
Several protocol families, including AF_RXRPC, AF_SMC, and AF_KCM, carry
traffic over kernel-owned sockets that the broker never creates and that
are not bound to loopback.

Allow only AF_UNIX, AF_NETLINK (still limited to NETLINK_ROUTE), and the
brokered AF_INET/AF_INET6 families, and restrict socketpair(2) to
AF_UNIX. Both decisions use the scalar domain argument. The broker
independently refuses non-INET families other than AF_UNIX and
AF_NETLINK with EAFNOSUPPORT.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(sandbox): remove legacy read-only mode

After native local accept, only two broker paths still wrote into
workload memory: getpeername on relayed connections and the per-message
lengths of a first sendmmsg to the DNS relay. Legacy mode existed only to
refuse those writes on kernels without WAIT_KILLABLE_RECV, and the
getpeername refusal broke CPython TLS on RHEL 9.

The broker now never writes workload memory:
- getpeername is no longer mediated and reports the kernel peer on every
  kernel; relayed connections report the loopback relay address.
- A first send to the DNS relay pins the broker's socket copy to the
  relay and continues the syscall, so the kernel performs the send and
  writes any per-message results.

Remove the listener mode, the task-memory write path and probe, and the
task_memory_write, cancellation, and task_memory_writes_disabled audit
evidence fields. WAIT_KILLABLE_RECV is still used when available and is
reported for diagnostics, but no mediation decision depends on it.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(sandbox): drop WAIT_KILLABLE_RECV and make handlers restart-safe

The broker no longer writes workload memory, so killable notification
waits only reduced how often a signal restarts a notified syscall. RHEL 9
kernels never had them, so the handlers must tolerate restarts anyway.
Install the listener without the flag on every kernel instead of
special-casing newer ones.

Handlers now check that the notification is still live immediately
before each side effect (bind, connect of the retained socket, listen,
and the first DNS send), and answer a restarted operation the broker
already completed as the kernel would: a repeated TCP connect returns
EISCONN, repeating the same UDP association succeeds, a repeat of a
completed bind succeeds, and a failed relay reports its errno.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): finish teardown only when every workload descendant has exited

Teardown waited only for registered process groups to disappear. A root
is unregistered once reaped, so a descendant that ignored SIGTERM could
outlive it while termination reported success and never sent SIGKILL,
violating the bounded-termination requirement.

The sandbox now becomes a child subreaper when it is not PID 1, so
orphaned descendants stay in its tree and are reaped. Termination waits
until no live descendant remains, and every scanned process is signalled
through a pidfd after confirming its start time, so a reused PID is never
signalled.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): keep a frozen workload from resuming itself

Freezing stops every workload process with SIGSTOP. tgkill and
rt_tgsigqueueinfo were not mediated, and mediated kill, rt_sigqueueinfo,
and tkill passed SIGCONT through, so a workload process that was not yet
stopped could resume the others while the supervisor recovered.

Mediate tgkill and rt_tgsigqueueinfo like tkill: refuse targets in the
sandbox thread group, report a thread outside the named group as
missing, and continue otherwise. The boundary marks the broker frozen
before stopping the workload and clears it after resuming; while frozen,
workload requests to send SIGCONT fail with EPERM.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): send mediated DNS datagrams from the broker socket

The native UDP send path re-ran the workload's own sendmsg, so
per-message ancillary data such as IP_PKTINFO rode along; only the
loopback destination contained a routing override.

The broker now reads the datagram and sends it from its retained,
relay-connected socket, which it builds without ancillary data, so a
per-message override cannot redirect the packet. It writes nothing back
into workload memory. A send carrying control data is refused with
EOPNOTSUPP rather than silently stripped.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): keep sandbox control variables out of the workload environment

The canonical process inherits the sandbox environment so the image's
own ENV reaches the workload, then removed only a denylist of credential
variables. Other variables in the reserved OPENSHELL_ namespace, such as
the serialized user environment and the log level, still reached the
workload.

Remove every inherited OPENSHELL_ variable before applying the declared
environment, and restore OPENSHELL_SANDBOX=1. The image's ordinary ENV
and declared variables are unaffected; declared variables cannot use the
reserved namespace.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): keep /proc read-only for GPU workloads

GPU mode granted read-write access to all of /proc so CUDA's cuInit could
write thread names to /proc/<pid>/task/<tid>/comm.

Keep /proc read-only. Open mediation now serves a writable open of the
caller's own thread comm file: the broker opens it and injects the
descriptor, with no syscall continued. The kernel accepts a comm write
only from the target's own thread group, so a substituted path or reused
thread ID cannot rename another process's thread. Any other /proc write
is left to Landlock, which denies it.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): keep the broker socket for connected UDP DNS sockets

Sending DNS datagrams from the broker's socket required the broker to
keep its copy, but the connect paths to the relay and to loopback still
released it. glibc connects the resolver socket and then sends A and
AAAA together with sendmmsg, which is mediated, so resolution failed.

Release the copy on those connect paths only for TCP. Add a regression
test for connect followed by sendmmsg that fails fast rather than
hanging.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): keep slow loopback connects from stalling mediation

The loopback connect path held the registry lock while polling a TCP
connect for up to five seconds on the single notification dispatcher.
A workload connecting to a busy local listener stalled every other
mediated syscall, including opens and signals.

Connect a duplicate of the retained socket without holding the lock. A
nonblocking socket gets the native EINPROGRESS and the kernel completes
the handshake on the shared socket; a blocking socket waits on a bounded
worker thread.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): contain network broker handler panics

A panic in a notification handler unwound the single broker thread while
the health flag still read healthy, so every later blocked workload
syscall hung until the sandbox was killed.

Run each dispatch under catch_unwind: a panicking handler fails only that
syscall with EIO and the broker keeps mediating. If the broker thread
ever exits, mark it unhealthy so dependent operations fail closed instead
of blocking.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): reject ancillary data on DNS sends instead of broker-sending

Sending mediated DNS datagrams from the broker's own socket required
keeping a broker handle for each DNS socket's whole life, which regressed
real name resolution through reclaim and resource accounting that the
mock-based unit tests did not exercise.

Revert to continuing the kernel send, but refuse a send that carries
ancillary control data (msg_controllen != 0) at read time, so a
per-message routing override such as IP_PKTINFO cannot ride a mediated
DNS send. A loopback destination contains an override that races the
check. Full broker-side UDP mediation is left to a separate change with
deployment e2e.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): address review findings across the hardening changes

Fixes:
- The blocking local-connect worker no longer toggles O_NONBLOCK on the
  open file it shares with the workload; it waits with a plain connect.
- tgkill and rt_tgsigqueueinfo are notified only when they send SIGCONT,
  so ordinary thread signals such as Go preemption stay in the kernel.
  The frozen flag is re-read immediately before delivery.
- A shell redirect to the caller's own thread comm file works again;
  O_CREAT and O_TRUNC are no-ops there and O_EXCL returns EEXIST.
- The teardown scan is a linear walk and fails closed when /proc cannot
  be read.

Simplifications:
- Remove tests that need infrastructure outside the repo or prove
  nothing: the topology harness tests, the static-server benchmark, the
  direct-syscall accept test, the comm rename test without Landlock, a
  serde-default test, and the test-only panic hook in setsockopt.
- Drop the unused Failed connect outcome, the redundant read-back after
  binding to loopback, redundant OPENSHELL_SANDBOX settings, and stale or
  duplicated comments; reuse the reserved environment prefix constant.
- Tighten the support matrix and OpenShift wording.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(cli): wait for input when piped exec stdin is nonblocking

Processes that inherit the same stdin share one open file description,
so any of them can make it nonblocking for all. `sandbox exec` then read
no input yet, got EAGAIN, and failed with "Resource temporarily
unavailable (os error 11)". Parallel e2e tests inherit the runner's
stdin, which made credential_gating fail intermittently.

The stdin reader now blocks in poll() until input or end of file arrives
instead of treating EAGAIN as fatal, and it leaves the shared descriptor's
flags alone. The e2e exec helper also stops inheriting the runner's
stdin, since its commands take no input.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): fix macOS lint and older-glibc test linking

Import HashSet only in the Linux-only process scan, and call gettid
through the raw syscall in a test, since the libc wrapper needs glibc
2.30 or newer.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(sandbox): clarify socket peer address reporting

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-10-06 04:08:04 +00:00
Shiju 8579dfb301 feat(providers): add multiple tokens to one request (#4047)
* feat(providers): compose dynamic credential grants per request

Select the most-specific grant independently for each protected header,
preserving each credential's issuer, audience and cache identity. Resolve
all selected grants before rewriting the request, replace agent-supplied
header copies, and fail closed on collisions or acquisition failures.

Allow distinct-header compositions in gateway validation and redact raw
issuer errors. Cover independent cache entries, atomic TLS relay failure,
header replacement and concurrent request isolation; document the profile
contract.

Fixes #3320

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(providers): correct dynamic grant CI checks

Use if-let for a grant result whose error payload is deliberately ignored,
and name the empty query map type in the regression fixture. Preserve grant
acquisition, failure redaction and request atomicity. Update the token-exchange
failure test to require the sanitized error instead of raw issuer text.

Signed-off-by: Shiju <shiju@nvidia.com>

---------

Signed-off-by: Shiju <shiju@nvidia.com>
2026-10-05 23:15:46 +00:00
Simon Scatton be7af99ff2 fix(e2e): override provider readiness supervisor image via environment (#4207)
Signed-off-by: Simon Scatton <sscatton@nvidia.com>
2026-10-05 17:54:48 +00:00
Matthew GrossmanandEvan Lezar 13c248fd45 fix(podman): create managed workspace volumes owned by the workload identity (#3981)
* fix(podman): create managed workspace volumes owned by the workload identity

Podman now creates the managed /sandbox volume with uid/gid options for the
resolved workload identity, so the workload starts directly as that
identity. This fixes rootful sandboxes whose image USER or policy
run_as_user could not write to a root-owned /sandbox, and removes the
root-then-drop workspace chown start path.

Resource admission accepts the managed workspace volume when its options
match the workload container's final identity, or are empty for volumes
created by older gateways. The channel volume still requires empty options.

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* test(podman): cover managed volume reuse and workspace access

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* refactor(podman): clarify managed volume creation and validation

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(podman): verify workspace access across user namespaces

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>
Co-authored-by: Evan Lezar <elezar@nvidia.com>
2026-10-05 16:42:43 +00:00
Shiju d676e036a4 fix(vm): reject conflicting workload identity selectors (#4036)
* fix(vm): enforce the configured workload identity

Reject conflicting policy users and groups before VM image preparation
and before guest attach or process startup changes state. Validate
supervisor policy updates against the protected VM workload identity.

Preserve the gateway CA transport and capability-free sandbox launcher.

Signed-off-by: Shiju <shiju@nvidia.com>

* test(sandbox): clarify VM identity rejection fixtures

Name invalid user and group fixtures distinctly and move the final
workload identity into its group mismatch test.

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(supervisor): align VM identity startup with current APIs

Pass the optional rejection-log key for VM identity failures and keep
generic startup-write regressions free of VM identity constraints.

Repair the call sites after the branch rebase so the identity and cleanup
proposals compile against the current startup helpers.

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(vm): restore inactive sandbox workload identity

Recover the persisted overlay owner before publishing stopped and terminal
sandboxes. Keep resources manageable when identity metadata is invalid.
Clarify fixed MicroVM ownership in policy-generation guidance.

Signed-off-by: Shiju <shiju@nvidia.com>

* test(vm): flush identity fixture before restart

Persist the canonical identity file before readiness and report the observed
exec, canonical and file-owner identities before comparing them.

Signed-off-by: Shiju <shiju@nvidia.com>

---------

Signed-off-by: Shiju <shiju@nvidia.com>
2026-10-03 20:03:39 +00:00
Drew Newberry 121930be07 fix(sandbox): replace lifetime exec cap with retry deadlines (#4105)
* fix(sandbox): replace lifetime exec cap with retry deadlines

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(sandboxes): remove exec recovery overview change

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(skills): remove exec recovery CLI skill change

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-10-02 21:31:40 +00:00
Drew Newberry 8719fc9f37 fix(sandbox): restrict provider file mode (#4093)
* fix(sandbox): restrict provider file mode (fixes #4091)

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(sandbox): set provider file mode with safe API

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-10-02 06:35:15 +00:00
Drew Newberry 021400be8a refactor(auth): separate sandbox identity from TLS (#3110)
* refactor(auth): separate sandbox identity from TLS

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(auth): clarify gateway mTLS behavior

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(auth): include workspace scope in TLS authorization checks

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(e2e): bound service auth sandbox names for large PIDs

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-10-01 04:33:25 +00:00
John T. Myers 2935e9731b fix(gateway): delete finalized ephemeral sandboxes while connected (#3984)
Start driver cleanup after terminal finalization and retain disconnect fallback. Add detached success and failure e2e coverage across supervisor-based drivers.

Closes #3938

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-10-01 00:09:08 +00:00
Derek Carr 912a077bd6 feat(service): add bearer authorization passthrough (#3796)
* feat(service): add bearer authorization passthrough

Signed-off-by: Derek Carr <decarr@redhat.com>

* docs(sdk): add service authorization migration guide

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(server): remove stale version import

Signed-off-by: Derek Carr <decarr@redhat.com>

* docs(upgrade): remove service authorization SDK guide

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(e2e): relabel provider readiness TLS mount

Signed-off-by: Derek Carr <decarr@redhat.com>

* test(e2e): stabilize exposed service routing

Signed-off-by: Derek Carr <decarr@redhat.com>

* test(e2e): support HTTPS service routing

Signed-off-by: Derek Carr <decarr@redhat.com>

---------

Signed-off-by: Derek Carr <decarr@redhat.com>
2026-09-30 20:22:46 +00:00
Polite_realismandEvan Lezar b8ffe5244c test(podman): move podman_preflight into driver-podman integration tests (#3783)
* test(podman): move podman_preflight into driver-podman integration tests

podman_preflight verifies that openshell-driver-podman fails fast when
its Podman socket is unreachable. It only needs the standalone driver
binary, not a gateway, so it never fit the gateway-backed e2e-podman
harness it lived under and never ran anywhere in CI.

Move it into crates/openshell-driver-podman/tests/ as a plain Cargo
integration test. It now runs via the existing required workspace test
job with no special mise task, workflow step, or coverage exception.

Signed-off-by: politerealism <burdcat17@gmail.com>

* test(podman): make preflight diagnostics portable

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: politerealism <burdcat17@gmail.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>
Co-authored-by: Evan Lezar <elezar@nvidia.com>
2026-09-30 06:05:18 +00:00
Drew Newberry 252882f37f feat(providers): serve sandbox config files on demand (#3832)
* feat(providers): serve sandbox config files on demand

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(providers): defer managed file documentation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(providers): mark managed file api experimental

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* style(go): format provider profile fields

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(providers): preserve legacy environment with managed files

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-30 01:41:32 +00:00
Eric Curtin cfcc3733bd fix(e2e): stop sandbox leaks from async Drop cleanup (#3750)
* fix(e2e): stop sandbox leaks from async Drop cleanup

Closes #2922

SandboxGuard::Drop spawned a detached thread to delete the sandbox.
The thread got killed with the test process before the delete
finished. Switch to a blocking command in Drop, like ManagedCleanup
already does. Also wrap two tests' manual cleanup in RAII guards so
a panic does not leak a sandbox.

Signed-off-by: Eric Curtin <eric.curtin@docker.com>

* test(e2e): arm sandbox guards before create

Address review: install guards with explicit names first.

Signed-off-by: Eric Curtin <eric.curtin@docker.com>

---------

Signed-off-by: Eric Curtin <eric.curtin@docker.com>
2026-09-29 10:10:29 +00:00
Philippe MartinandJohn Myers 9cb72baa2e feat(docker): support corporate proxy CA bundles (#3549)
* feat(docker): support corporate proxy CA bundles

Closes #3545

Validate and stage operator-owned proxy CA bundles for Docker supervisors, add corporate proxy E2E coverage, and document the trust contract.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(docker): validate proxy config on startup

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(docker): use the E2E workload image for proxy tests

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(docker): generate strict corporate proxy certificates

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(docker): surface intercepted TLS fixture errors

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(docker): drain buffered TLS proxy data

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(docker): relay intercepted HTTP deterministically

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: Philippe Martin <phmartin@redhat.com>
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-09-29 05:15:02 +00:00
Drew Newberry acbac9cb79 feat(sandbox): add main restart policy (#2798)
* feat(sandbox): add main restart policy

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): harden policy-driven restarts

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): address restart review feedback

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>

* fix(sandbox): port restart policy to current runtime

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* perf(sandbox): restart promptly after terminal delivery

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>
2026-09-29 02:18:30 +00:00
Shiju 1358941b81 feat(mcp): inspect requests with Tower-selected protocol profiles (#3335)
* fix(sandbox-backend): sort boundary request objects before hashing

Sort boundary request objects recursively before hashing so serde_json's
preserve_order feature cannot change digest identity. Cover canonical
bytes, envelope round trips, and rejection of modified provider values
and operations.

Signed-off-by: Shiju <shiju@nvidia.com>

* feat(mcp): upgrade tower-mcp-types to 0.22.2

Upgrade tower-mcp-types from 0.12.0 to an exact-pinned 0.22.2 and use its
inspection APIs to validate MCP requests against the selected revision.
Carry inspection metadata into policy evaluation and validate requests
after header rewriting, before forwarding.

Add explicit support for the sessionless 2026-07-28 revision while keeping
2025-11-25 as the default. Validate per-request metadata and standard HTTP
header mirrors, and support discovery, tools, and subscription requests.

Delegate batch availability and parameter schemas to Tower. Share typed
request names between policy and HTTP checks, retain the local batch
resource cap, and centralize MCP policy version parsing and ordering.

Keep supported MCP revisions and shared allowlist parsing in the canonical
policy schema; core re-exports those types. Tower owns wire-profile
semantics, and every supported policy revision must map to the matching
inspector profile.

Reject duplicate JSON keys, invalid known-method parameters, unavailable
methods, and unsupported batches. Keep exact extension allow rules and
deny precedence. Document request inspection boundaries and add unit,
forwarding, and sandbox coverage.

Refs #2174.

Signed-off-by: Shiju <shiju@nvidia.com>

* test(mcp): prove authorization at the forwarding boundary

Cover March batch denial in both member orders, valid and malformed
controls, and audit behavior across both relay entry paths. Exercise real
middleware tool rewrites with matching metadata and assert the exact
upstream representation or zero forwarded bytes.

Verify legacy bodyless SSE GET remains usable while GET tool bodies and
unsupported DELETE cleanup are rejected. Clarify request-selected profile
and middleware mutation comments without changing production behavior.

Signed-off-by: Shiju <shiju@nvidia.com>

* test(mcp): exercise permitted profiles through the sandbox proxy

Cover March and June singleton policies and select November and July
separately under one endpoint allowlist. Capture upstream tool receipts
to distinguish proxy policy denial from an upstream rejection.

Extend middleware rewrite coverage to June and multi-version policies,
and preserve the sessionless discovery and subscription checks through
the shared fixture helpers.

Signed-off-by: Shiju <shiju@nvidia.com>

* test(kubernetes): box the admission check future

Keep the admission test future below Clippy's size limit when the
workspace dependency features are unified.

Signed-off-by: Shiju <shiju@nvidia.com>

* test(mcp): reuse the forwarding fixture identity cache

Share the binary identity cache across protocol-profile cases, matching
the proxy lifecycle and avoiding repeated hashes of the test executable.
Keep procfs authorization and all forwarding assertions intact.

Signed-off-by: Shiju <shiju@nvidia.com>

---------

Signed-off-by: Shiju <shiju@nvidia.com>
2026-09-28 20:48:50 +00:00
Evan Lezar eef8bec0c9 test(e2e): run podman suite with tmachine (#3637)
* test(e2e): remove superseded podman userns coverage

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(tmachine): run podman e2e archive

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(tmachine): generate podman e2e archive inventory

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-09-28 11:15:33 +00:00
Evan Lezar 0c29d8e061 test(conformance): migrate file transfer scenarios (#3597)
Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-09-28 09:02:44 +00:00
Evan Lezar c63f8ce564 test(cli): migrate gateway-free smoke coverage (#3641)
* test(cli): migrate gateway-free smoke coverage

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(e2e): remove migrated tests from podman CI

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-09-28 09:00:30 +00:00
Drew Newberry 6f00d5cacc fix(e2e): reserve distinct corporate proxy fixture ports (#3761)
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-28 07:10:23 +00:00
Drew Newberry 7a50c0899f fix(cli): stream piped exec stdin beyond gRPC request limit (#3687)
* fix(cli): stream piped exec stdin across gRPC messages

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(cli): preserve small exec requests and surface stdin errors

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(cli): preserve exec stdin limit across gRPC streaming

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-25 04:45:57 +00:00
Drew Newberry 4688061882 fix(sandbox): deliver complete exec output before success (#3688)
* fix(sandbox): preserve exec output through channel close

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(exec): propagate output delivery failures before exit

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(cli): clarify exec output delivery failures

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-25 04:03:49 +00:00
Polite_realismandDrew Newberry d376c90755 test(podman): close rootful userns, resource-limit, and daemon-failure CI gaps (#3690)
* test(podman): run driver-podman userns suite against rootful Podman too

The driver-specific-integration job only ran the driver-podman testsuite
(default/auto/keep-id/private userns reference checks) against
fedora-podman-rootless, leaving rootful behavior for this scenario
unverified even though the compute driver auto-detects and explicitly
supports rootful Podman.

The default-userns-baseline and userns-profile playbooks hard-asserted a
rootless tmachine gateway user, so pointing them at a rootful environment
would have failed that assertion immediately rather than exercising
anything. They now detect rootful vs. rootless via the existing
tmachine_container_runtime role and branch the reference-capture user
accordingly, while keeping the captured reference file itself owned by
tmachine, since the archived test binary that reads it back always runs
unprivileged as tmachine regardless of daemon mode.

Signed-off-by: politerealism <burdcat17@gmail.com>

* test(podman): add real-daemon coverage for resource limits and daemon failure

Neither the Podman driver's resource-limit enforcement nor its behavior
when the Podman daemon is unreachable had any test coverage against a
real daemon; both were only exercised through unit tests against a
mocked Podman client.

podman_resource_limits.rs creates a sandbox with --cpu/--memory flags and
reads /sys/fs/cgroup/memory.max and cpu.max from inside the sandbox
itself, verifying the limit is actually enforced rather than just echoed
back by the template API. Expected values are cross-checked against the
driver's own parse_cpu_to_microseconds/parse_memory_to_bytes and against
a real local `podman run --cpus/--memory` container.

podman_preflight.rs spawns the standalone openshell-driver-podman binary
against a guaranteed-nonexistent Podman socket and asserts it exits
non-zero within its bounded retry window with an actionable error naming
the socket path, rather than hanging or failing silently.

Signed-off-by: politerealism <burdcat17@gmail.com>

* test(podman): make rootful userns and cgroup checks pass

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(podman): match lifecycle containers by isolation role label

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(podman): accept non-expiring bootstrap tokens

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: politerealism <burdcat17@gmail.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
2026-09-24 20:16:29 -07:00
grs 8369bc11a5 fix(podman): restore host gateway alias mediation (#3606)
* fix(podman): restore host gateway alias mediation

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(podman-e2e-tests): enable broader test podman e2e coverage

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(tests): make test more reliable

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(podman): fix macos linting error

Signed-off-by: Gordon Sim <gsim@redhat.com>

---------

Signed-off-by: Gordon Sim <gsim@redhat.com>
2026-09-24 18:41:59 +00:00
Drew Newberry 0518bd4c83 fix(sandbox): preserve local sessions across host sleep (#3573)
* fix(cli): recover sandbox connect transport

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(auth): propagate non-expiring local sandbox sessions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(cli): distinguish main exit from transport loss

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(cli): bound sandbox connect recovery

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-24 17:22:30 +00:00
krishicks 0b351c4a9b fix(helm)!: reduce gateway Secret privileges (#3616)
* fix(driver-kubernetes-secrets)!: store provider credentials in one namespace

The Kubernetes Secrets credential driver now stores every credential in its
configured namespace in all workspace modes and rejects handles that reference
any other namespace before contacting the Kubernetes API. The gateway reaches
credential Secrets through the Role in that namespace; this allows removing the
Secret rules from the ClusterRole.

- Remove the workspace_mode, gateway_id, and allow_reference_namespace driver
  settings and stop rendering them from Helm. Configurations that set them fail
  at startup. Existing credential state is not migrated.
- Add server.credentialDrivers.kubernetesSecrets.createNamespace to provision a
  dedicated credential namespace. The namespace is kept on uninstall, adopted
  by a reinstall of the same release, and left untouched when something else owns
  it.
- Update the gateway config reference, Kubernetes setup docs, 0.1.0
  upgrade guide, compute-runtime architecture, and cluster debugging skill.

Signed-off-by: Kris Hicks <khicks@nvidia.com>

* fix(helm): reduce gateway Secret permissions

Remove the gateway's Secret list permission in every workspace mode and
grant source Secret reads through a Role in the sandbox namespace.
Bootstrap Secret cleanup deletes Secrets by exact name instead of listing
them.

- Grant get on the copied client TLS and image-pull Secrets through a Role in
  the sandbox namespace. The ClusterRole keeps get and patch on those names
  for the ownership check and server-side apply into workspace namespaces.
- Delete sandbox and supervisor bootstrap Secrets by exact name, derived from
  the runtime generation recorded on the Sandbox and, on restart, the target
  generation, tolerating 404. The generation annotation is cleared only after
  cleanup succeeds, and each bootstrap Secret has a Pod owner reference, so
  garbage collection removes any generation the driver does not name.
- Drop Secret list from the ClusterRole and the shared-mode sandbox Role.
- Extend the managed e2e RBAC checks to Secret list.
- Update the Kubernetes setup and sandbox runtime docs, compute-runtime
  architecture, and cluster debugging skill.

Signed-off-by: Kris Hicks <khicks@nvidia.com>

* fix(driver-kubernetes): stage workspace Secrets per runtime generation

Every Secret the Kubernetes driver writes into a workspace namespace is
now scoped to one sandbox runtime generation, immutable, and created
with create only, so the gateway never reads, patches, or adopts an
existing Secret there. This removes the gateway's cluster-wide get and
patch on the copied client TLS and image-pull Secret names.

- In managed mode, create an immutable copy of each configured
  image-pull Secret per generation, named os-pull-<id>-<generation>-<n>
  and owned by the generation's workload and supervisor Pods. Pods and
  the restarted Sandbox template reference those names, and generation
  cleanup deletes them by name. A Secret already holding a generation
  name fails the create.
- Outside shared mode, stage the gateway client TLS material into the
  supervisor bootstrap Secret instead of copying the client TLS Secret
  into the workspace namespace.
- Remove the fixed-name TLS and image-pull copies, the target ownership
  read, and the ClusterRole get and patch rule on the copied names.
  Source reads stay in the sandbox-namespace Role.
- Update the managed e2e to expect generation image-pull Secrets and
  client TLS material in the supervisor bootstrap Secret, and to check
  that the gateway cannot read the copied names in workspace namespaces.
- Update the gateway config and compute driver references, Kubernetes
  setup and sandbox runtime docs, compute-runtime architecture, driver
  README, and cluster debugging skill.

Signed-off-by: Kris Hicks <khicks@nvidia.com>

* fix(helm)!: grant operator-mode Secret permissions through the workspace chart

The operator-mode gateway ClusterRole grants no Secret permissions.
The openshell-workspace chart Role, installed in each operator-managed
namespace, grants the gateway create and delete on Secrets for sandbox
runtime generations. Operator-managed namespaces require the workspace
chart.

- Fail the chart tests on any ClusterRole rule that includes Secrets in
  operator and shared modes.
- Install the workspace chart when the operator e2e provisions a
  namespace, and assert that the gateway has no Secret permissions in a
  namespace without it.
- Update the Kubernetes setup docs, 0.1.0 upgrade guide, compute-runtime
  architecture, and cluster debugging skill.

Signed-off-by: Kris Hicks <khicks@nvidia.com>

---------

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-09-23 21:35:22 +00:00
Jim Meyer 62df64625b fix(identity): assess leaf and ancestor executable identities (#3633)
* fix(security): pin supplied executable identity chains

Signed-off-by: Jim Meyer <jimeyer@nvidia.com>

* docs(security): document executable identity chain pinning

Signed-off-by: Jim Meyer <jimeyer@nvidia.com>

* fix(binary-identity): compile Linux ancestry hashing

Signed-off-by: Jim Meyer <jimeyer@nvidia.com>

* test(binary-identity): avoid cross-label ancestry fixture

Signed-off-by: Jim Meyer <jimeyer@nvidia.com>

* test(e2e): choose distinct denied TCP port

Fix flaky test due to sequential port assignment on MacOS

Signed-off-by: Jim Meyer <jimeyer@nvidia.com>

* fix(identity): bound executable evidence cache

Reject identity chains atomically when the supervisor cache reaches its hard limit, and preserve existing pins without eviction. Classify malformed or conflicting evidence separately from policy denials at the staged TCP boundary.

Signed-off-by: Jim Meyer <jimeyer@nvidia.com>

---------

Signed-off-by: Jim Meyer <jimeyer@nvidia.com>
2026-09-23 21:07:24 +00:00
Drew Newberry 52aac37866 fix(network): honor HTTP response connection closure (#3581)
* fix(network): honor HTTP response connection closure

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(network): fix EOF fixture socket setup

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(network): keep pipeline probe response reusable

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-23 12:01:46 -07:00
Piotr Mlocek bed9e5eafc fix(supervisor): use better error message when sandbox connect is not available (#3572)
* fix(supervisor): explain unavailable main terminal attachments

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(supervisor): reset cursor after terminal attachment errors

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(supervisor): format read-only warnings for terminal clients

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs: drop terminal attachment documentation additions

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(supervisor): exit read-only viewers on Ctrl-C

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs: simplify read-only viewer Ctrl-C guidance

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

---------

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-09-23 17:31:13 +00:00
Evan Lezar 6cb1140c66 fix(drivers): normalize label namespace (#3609)
Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-09-23 14:30:27 +00:00
krishicks f8002d19ad fix(e2e): repair the credential driver test (#3565)
The credential driver e2e has not passed end to end, and the disabled
kubernetes-credential-drivers CI lane hid three problems.

The test broke when the JSON format of provider list changed.
Continuation-token pagination (#3249) changed
provider list --output json from a bare array of providers to an object
with next_page_token and a providers array. The test still parsed the
output as an array, so it failed before checking either storage
backend. Read the providers array from the new object instead.

Its sandbox name was about 58 characters, but sandbox names are
DNS-routable and limited to 19, so sandbox creation was rejected. Build
a short unique name instead.

The sandbox guard deletes its sandbox from a detached thread on drop, so
the test deleted the provider while the sandbox still existed. The
gateway rejects deleting a provider that is attached to a sandbox, the
test ignored that error, and the credential Secret remained. Delete the
sandbox explicitly before returning from the sandbox check.

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-09-22 20:55:04 +00:00
Piotr Mlocek 4b1c09de28 fix(network): preserve chunked request boundaries (#3530)
* fix(network): preserve chunked request boundaries

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(network): bound chunk framing amplification

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(supervisor): box sandbox runtime future

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(network): buffer chunked relay read-ahead

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(network): flush completed chunks promptly

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

---------

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-09-22 19:04:59 +00:00
Philippe Martin 35e0a68e4a feat(kubernetes): support corporate proxy CA bundle (#3447)
* feat(kubernetes): support corporate proxy CA bundle

The Kubernetes driver had no way to supply a CA bundle for the corporate
egress proxy, so an `https://` proxy with a private CA, or a TLS-intercepting
proxy, could not be used. Podman and VM already expose `proxy_ca_bundle`.

Add `proxy_ca_bundle` to `[openshell.drivers.kubernetes]` as a path the
gateway Pod reads. The gateway stages the PEM into the existing
per-generation supervisor bootstrap Secret and passes
`--upstream-proxy-ca-bundle` on the supervisor argv. That Secret is already
immutable, owner-referenced and garbage-collected, and its volume mounts
every key at /.openshell/supervisor with no items filter, so this needs no
new object kind, volume, mount, or RBAC verb, and works in shared, managed
and operator workspace modes.

The bundle is deliberately read from the gateway's filesystem rather than
referenced as an object in the sandbox namespace. It becomes a trust anchor
for every upstream the sandbox reaches, so it must stay in the gateway's
trust domain; the immutable staging Secret also keeps the anchor from
changing underneath a running sandbox.

Bound the staged bundle at 256 KiB. The shared reader's limit is exactly the
apiserver's own Secret limit and the bootstrap Secret carries four other
keys, so a bundle between the two would pass gateway startup and then fail
every sandbox create with an opaque `data: Too long`.

Delegate the URL, no_proxy, connect_by_hostname and ca_bundle rules to the
shared validate_upstream_proxy_settings, keeping the Secret-specific
credential block local: this driver accepts an explicit
`proxy_auth_allow_insecure = false` without credentials, which the shared
rules reject. This also fixes the acknowledgement being demanded for an
`https://` proxy, where the credential travels inside the verified TLS
session. Add auth_setting_label so the inline-credential diagnostic names
the Secret keys instead of proxy_auth_file, which this driver rejects as an
unknown key.

Document that the bundle should carry only the CA that signs the proxy's
certificate, or that an intercepting proxy re-signs upstream certificates
with. Public roots already reach the sandbox through the supervisor image and
its TLS stack, and the bundle is concatenated with that system store into a
single boundary control frame, so a full merged trust bundle spends the frame
budget on duplicated roots. The frame, not the apiserver Secret limit, is the
tighter of the two ceilings in practice; raising the staging bound requires
checking it.

Closes #3443

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(helm): quote proxy CA ConfigMap references

Signed-off-by: Philippe Martin <phmartin@redhat.com>

---------

Signed-off-by: Philippe Martin <phmartin@redhat.com>
2026-09-22 17:57:40 +00:00
50230616d5 refactor(runtime): retire Community image dependencies (#3386)
* feat(sandbox): default to official Alpine sandbox image

default_sandbox_image() now returns docker.io/library/alpine:3.22, a generic
version-qualified official image, so a fresh install no longer depends on the
community sandbox image catalog. All compute drivers (docker, podman,
kubernetes, vm) inherit this fallback.

Part of #3116.

Signed-off-by: Akram
Signed-off-by: Akram <akram.benaissi@gmail.com>

* feat(deploy): default deployment configs to the official Alpine sandbox image

Update the shared gateway default_image, Helm chart values, the standalone
Kubernetes manifest, and the dev gateway task scripts to use
docker.io/library/alpine:3.22 instead of the community base image, consistent
with default_sandbox_image(). GPU e2e image-build base is left unchanged (CUDA
needs a glibc base).

Part of #3116.

Signed-off-by: Akram
Signed-off-by: Akram <akram.benaissi@gmail.com>

* feat(driver): default to numeric non-root identity for USER-less images

With the default sandbox image now Alpine, images that declare no OCI USER
must start instead of being rejected. When the image declares no USER and
the policy requests none, the Podman and Docker drivers now supply a numeric
non-root identity (DEFAULT_SANDBOX_UID/GID = 1000) instead of rejecting,
matching the numeric-identity behavior of the Kubernetes and VM drivers. The
supervisor's resolved-identity path runs the sandbox as a synthesized
non-root account without the account existing in the image. Images that
declare a USER keep the OCI resolution path unchanged.

Part of #3116.

Signed-off-by: Akram <akram.benaissi@gmail.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(conformance): use Alpine workload image

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* refactor(policy): drop community image /app path from default policy

The restrictive default policy granted read-only access to /app, a directory
that only existed in the community base image. A generic Alpine default has no
/app, so remove it. Landlock best-effort already ignores absent paths; this
just stops advertising a community-specific layout in the default.

Part of #3116.

Signed-off-by: Akram
Signed-off-by: Akram <akram.benaissi@gmail.com>

* docs(config): document Alpine default images

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(podman): report early sandbox termination

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(podman): initialize rootless workspace ownership

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(sandbox): qualify NVIDIA Ubuntu default

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): initialize rootful default workspace

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(sftp): add native sandbox adapter

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sftp): gate runtime helper support to Linux

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sftp): support standard OpenSSH file operations

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sftp): harden rename and special file handling

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(runtime): remove community image dependencies

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(e2e): build provider readiness tool fixture

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(e2e): use a dedicated Noble fixture for Docker tests

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Akram
Signed-off-by: Akram <akram.benaissi@gmail.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Evan Lezar <elezar@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
2026-09-22 14:43:51 +02:00
0a770d9173 feat(kubernetes): support HA gateway rebalancing (#1868)
* feat(kubernetes): support HA gateway rebalancing

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* perf(server): cache peer connections, tokens, and owner lookups

Every forwarded relay rebuilt its setup from scratch: an owner lookup, a
blocking read of the peer token, a TLS connect to the owning replica, and
a TokenReview plus Pod GET on the receiving side. Sandbox service routing
does this per HTTP request, so the apiserver calls scaled with traffic.

Cache all of it on ServerState:

- peer channels pooled per endpoint, so relays multiplex over one
  connection instead of redialing
- peer tokens keyed by SHA-256, expiring at min(ttl, token exp) so a hit
  cannot accept an expired token
- owner records for 3s against a 45s ownership TTL, still freshness
  checked before use

Entries are evicted when a relay fails. Also raise HTTP/2
max_concurrent_streams to 1024, since pooling funnels every relay between
two replicas onto one connection and hyper's default of 200 sits below
the 256 pending-relay budget.

Signed-off-by: divesh <dgude@nvidia.com>

* perf(server): pool upstream connections for sandbox services

Each HTTP request to a sandbox service opened its own supervisor relay,
paying a new TCP connection and HTTP/1 handshake every time. Worse, it
counted against the 32 in-flight relay cap, so a service handling more
than 32 concurrent requests failed outright.

Pool idle upstreams per endpoint and port, up to 8 each for 15s. Reuse is
safe because the pool only returns a connection hyper reports as ready,
and HTTP/1 cannot start a request until the previous body has drained.
Upgrades are never pooled since they take the connection over, and a
failed send evicts that endpoint. Pruning is bounded per key, with the
full sweep limited to once per 30s.

Signed-off-by: divesh <dgude@nvidia.com>

* fix(server): address HA gateway review findings (#3449)

- Let a gateway own supervisor sessions without a peer endpoint. Requiring
  one whenever the store is PostgreSQL broke every single-instance
  PostgreSQL deployment, because no sandbox supervisor could connect.
  A cross-replica request to an owner that advertises no endpoint now fails
  immediately naming the cause, instead of retrying until the wait timeout.
- Close a supervisor session on heartbeat only when another replica owns it,
  or after renewals fail for the ownership TTL. A database error no longer
  drops every session heartbeating during an outage.
- Clamp owner record ages at zero so a skewed or corrupt stored timestamp
  cannot produce a negative age.
- Bound the cross-object advisory lock with a lock timeout, so a stuck holder
  fails instead of blocking every mutation in the fleet.
- Refuse to start when a peer endpoint is configured on a multi-replica
  backend but peer authentication is unavailable, and warn when a
  multi-replica backend has no peer endpoint at all.
- Reject a plaintext peer endpoint when the gateway serves TLS.
- Skip the sandbox watch poller on single-replica backends, where the local
  update bus already sees every write.
- Rate-limit the peer owner cache sweep so an insert no longer scans the
  whole map under the lock.
- Retry GET and HEAD on a pooled upstream the sandbox closed, instead of
  returning 502, and drop an emptied endpoint from the pool right away.
- Document the gateway peer environment variables and the post-rollout
  ownership skew operators should expect.

Signed-off-by: divesh <dgude@nvidia.com>

* fix(server): harden HA supervisor ownership

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Signed-off-by: divesh <dgude@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: divesh <dgude@nvidia.com>
Co-authored-by: Divesh Chowdary <47188680+FrostGod@users.noreply.github.com>
2026-09-21 21:22:12 +00:00
Emilien Macchiands2cube cb6e88acb7 fix(vm): unpack registry images correctly and validate prepared disks (#3524)
* fix(vm): unpack registry images correctly and validate prepared disks

The registry image-prep path expected `umoci raw unpack` to produce a
bundle-style rootfs/ subdirectory, but it extracts the image filesystem
directly into the target. Every registry prep therefore failed after a
successful unpack. Guest init exit codes do not survive the libkrun
boundary, so the failure looked like success and the broken disk was
cached, making every later sandbox for that image fail with "prepared
image disk missing /image-rootfs".

VM E2E started hitting this after the bootstrap image moved to
nvcr.io/nvidia/base/ubuntu:24.04: `--from base` no longer matches the
bootstrap image, so it now goes through registry prep.

- Accept umoci's direct extraction layout in the guest prep script.
- Build the image rootfs under a partial directory and rename it to
  /image-rootfs only after every prep step succeeds.
- Check the prepared disk for /image-rootfs before caching it. On
  failure, leave the cache untouched and report the prep console tail.
- Size the prep disk to hold the payload and the unpacked rootfs at the
  same time. The community base image needs 1.40 GB + 3.32 GB, which
  did not fit in the old payload*3 + 512 MiB.

Fixes #2358

Co-authored-by: s2cube <26961336+s2cube@users.noreply.github.com>
Signed-off-by: Emilien Macchi <emacchi@redhat.com>

* test(e2e): run tool-dependent VM tests from the community base image

The host_gateway_alias and vm_corporate_proxy workloads run curl and
python3. The VM driver now defaults to nvcr.io/nvidia/base/ubuntu:24.04,
which ships neither, so these tests fail in VM E2E with "command not
found". Request the community base image explicitly with `--from base`.

Docker, Podman, and Kubernetes E2E already default to that image, so
their behavior is unchanged.

Signed-off-by: Emilien Macchi <emacchi@redhat.com>

---------

Signed-off-by: Emilien Macchi <emacchi@redhat.com>
Co-authored-by: s2cube <26961336+s2cube@users.noreply.github.com>
2026-09-21 13:25:02 -07:00
Shiju fa8f6d3949 feat(cli): promote profile commands to top level (#3258)
* feat(cli): promote profile commands to top level

Add profile discovery and management commands with shared handlers for
the existing provider entry points. List a flat catalog across scopes
and follow continuation tokens through full and short pages.

Describe metadata, credentials, endpoints, TLS inspection, and MCP access
settings while preserving complete JSON/YAML definitions. Cover parser
equivalence, scope forwarding, pagination, and inspection settings with
focused unit and compiled-CLI integration tests.

Update docs, public skills, examples, and E2E command invocations.

Refs #2588

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(cli): remove redundant workspace selector qualification

Use the imported WorkspaceSelector in the provider integration helper so
the target passes Clippy with warnings denied.

Signed-off-by: Shiju <shiju@nvidia.com>

---------

Signed-off-by: Shiju <shiju@nvidia.com>
2026-09-21 17:18:55 +00:00
Drew Newberry dee4f98dda chore(vm): refresh runtime defaults and hardening (#3446)
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-21 10:19:30 +00:00
Drew Newberry 17ce738bfb fix(ci)!: remove gateway callback listener dependency (#3365)
* fix(ci): repair post-merge release canary

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(packaging): bootstrap canary runtime prerequisites

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci(canary): collect macOS VM diagnostics

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci(canary): pin libkrun-compatible macOS runner

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci(canary): limit macOS smoke test to package startup

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(compute)!: remove gateway callback listeners

Run Docker supervisors on host networking so they use the operator-configured primary gateway endpoint. Remove the unused compute-driver callback listener negotiation and listener-scoped routing machinery.

BREAKING CHANGE: The ComputeDriver API no longer exposes GetGatewayListenerRequirements or GatewayListenerRequirement. External drivers must regenerate bindings and connect supervisors to the configured primary gateway endpoint.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(docker): use sandbox runtime image in launcher

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(podman): exercise production endpoint selection

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(e2e): route supervisors to reachable gateways

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(e2e): align Podman endpoint fixtures

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): preserve host aliases for supervisors

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): align sandbox host gateway pin

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(e2e): address Docker fixtures by bridge IP

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(e2e): serialize sandbox lifecycle cases

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(e2e): host Docker TCP fixture with gateway

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(e2e): use loopback for host-network supervisor

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-18 20:55:55 +00:00
Drew Newberry d91b1999a0 feat(api)!: use sandbox names as canonical RPC references (#3272)
* feat(api)!: use sandbox names as canonical references

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(cli): update forward color fixture for workspace scope

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(supervisor): use sandbox names for settings lookup

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(e2e): use canonical sandbox request fields

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(e2e): use canonical sandbox receipt field

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(server): harden sandbox mutation handling

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(ci): update rebased sandbox references

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(api)!: standardize canonical entity references

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(api): codify protobuf API conventions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(api): preserve workspace selector semantics

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(api): restore workspace selector parity

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(api): preserve descriptive name fields

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(api): update e2e request fixtures

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(core): omit workspace selector during bootstrap

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* chore(api): remove proto convention checker

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* chore(api): refresh schema fingerprints after rebase

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(cli): use canonical provider receipt field

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-18 11:39:55 -07:00
Jorge e2939b919d feat(e2e): run the Kubernetes e2e suite on cargo-nextest with machine- and human-readable reports (#3344)
* feat(e2e): run kubernetes suite on cargo-nextest with JUnit/HTML reports

Switch e2e:kubernetes (and all its variants) from `cargo test` to
`cargo nextest run` for per-test process isolation and output consistent
with the other nextest-based CI runs.

- Add a dedicated `e2e-kubernetes` nextest profile with a JUnit report and a
  generous slow-timeout (60s flag, 5-min terminate) suited to live-cluster
  tests; kept separate from `ci` so its JUnit path and timeouts don't affect
  the workspace run.
- Pin `--target-dir` for the run so the profile's relative JUnit path resolves
  to the repo-root results/ regardless of any inherited CARGO_TARGET_DIR
  (nextest ignores absolute JUnit paths).
- Render the JUnit XML to a standalone HTML report via xsltproc and a committed
  XSLT stylesheet (best-effort; never masks the test exit code).
- Name each report via `OPENSHELL_E2E_REPORT_NAME` (default `e2e-kubernetes`),
  used verbatim for both the `results/<name>.{xml,html}` filenames and the HTML
  heading. Tasks that invoke the script multiple times in one run set a distinct
  name per invocation so the reports no longer clobber the single fixed path:
  the credential-driver runs write results/e2e-kubernetes-secrets.xml and
  -vault.xml, and e2e:kubernetes:agent-sandbox-versions writes
  results/e2e-kubernetes-agent-sandbox-v1beta1.xml and -v1alpha1.xml.
- Declare cargo-nextest in mise [tools] so the task runs without the Nix shell.
- Ignore the results/ output directory.

The results/ reports do not leak information. They are gitignored and no
workflow uploads them as artifacts, so they stay on the ephemeral CI runner
and are discarded when it is torn down. The HTML template renders only test
names, status, timings, and failure messages (no captured stdout/stderr).
Moving from `cargo test -- --nocapture` to nextest's captured, failure-only
output also reduces what lands in the retained, viewable console logs.

Signed-off-by: Jorge Garcia Oncins <jgarciao@redhat.com>

* chore(e2e): revert per-lane report names for agent-sandbox-versions

The agent-sandbox-versions task runs two lanes sequentially against the
same cluster: v0.5.0 (v1beta1 storage version) then v0.4.6 (v1alpha1).
On a reused cluster the second lane fails when kubectl applies the older
CRD, because Kubernetes refuses to drop v1beta1 from spec.versions while
it remains in status.storedVersions (the storage-version downgrade
guardrail). This is a pre-existing issue with the v0.4.6 lane, unrelated
to the nextest reporting work.

The per-lane OPENSHELL_E2E_REPORT_NAME additions do not address that
downgrade failure, so revert them to keep this PR scoped to the nextest
change. Agent Sandbox 0.4.x is also superseded (1.0.0 is published);
dropping or bumping the v1alpha1 lane is left as a follow-up.

Signed-off-by: Jorge Garcia Oncins <jgarciao@redhat.com>

---------

Signed-off-by: Jorge Garcia Oncins <jgarciao@redhat.com>
2026-09-18 15:50:03 +00:00
Evan Lezar 2263685cf3 test(tmachine): migrate Keycloak provider refresh coverage (#3404)
* test(tmachine): add Keycloak provider refresh suite

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* refactor(tmachine): share container runtime detection

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* ci(tmachine): run feature suites in GitHub Actions

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* ci(tmachine): run conformance with Podman tests

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* ci(tmachine): cover provider refresh with Podman

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* ci(integration): split input preparation from runners

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-09-18 13:56:07 +02:00