Commit Graph
215 Commits
Author SHA1 Message Date
Shiju 8aa5846d7a fix(cli): preserve existing directories during single-file upload (#4177)
* fix(cli): preserve existing directories during single-file upload

Inspect remote destination types before choosing archive entry names.
Keep directory and source symlinks intact, recheck detected type changes,
and report the resulting path. Add real tar and live conformance coverage.

Closes #4175

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(cli): box upload future in sandbox_create

Boxing the upload future keeps the sandbox_create future under the clippy::large_futures threshold after sandbox_upload_planned began returning the uploaded path.

Signed-off-by: Shiju <shiju@nvidia.com>

---------

Signed-off-by: Shiju <shiju@nvidia.com>
2026-10-07 22:20:46 +00:00
Divesh 834b79a8c2 feat(server): place supervisor sessions by consistent hash (#3661)
* feat(server): place supervisor sessions by consistent hash and hand them off on shutdown

Signed-off-by: divesh <dgude@nvidia.com>

* feat: route long-lived sandbox connections to the owning gateway replica

Signed-off-by: divesh <dgude@nvidia.com>

* feat(helm): add opt-in per-replica GRPCRoute routing

Signed-off-by: divesh <dgude@nvidia.com>

---------

Signed-off-by: divesh <dgude@nvidia.com>
2026-10-07 05:01:32 +00:00
Jason T. Greene b3a9bd85fb perf(server): drop per-connection session tokens for TCP forwards (#3734)
`openshell forward service` minted an SSH session token before every
forwarded TCP connection and revoked it afterwards: two store commits per
connection. The token added nothing on that path. `ForwardTcp` already
authenticates the caller and authorizes it against the sandbox's workspace
on every stream before it looks at the token, the relay to the supervisor
is opened with the sandbox id and target only, and the token is never
forwarded, audited, or visible to the target service. The mechanism exists
for `openshell sandbox ssh`, where the process that opens the stream is an
ssh ProxyCommand holding nothing but the token. Reusing it per TCP
connection put a store write on the connect path and serialized concurrent
forwards on commit latency: #3494 measured the symptom, and #3543 made the
commits cheaper, but each one still holds SQLite's writer lock for an
fsync, so connection setup under a burst stayed linear in the number of
concurrent connections.

Let `target.tcp` streams omit `authorization_token`. The gateway admits
them on the already-authorized principal, counts them against the same
per-sandbox connection cap, and touches no store. `target.ssh` streams keep
requiring the token. A token supplied with a TCP target is still validated
and counted per token, so an older CLI against a new gateway is unchanged.
The CLI stops minting and revoking a session per forwarded connection;
against a gateway that predates this change it recognizes the
`authorization_token is required` rejection once and falls back to
per-connection tokens for the rest of that forward.

Tests cover token-less TCP admission and slot release, SSH targets still
rejected without a token, a supplied token still validated, and the
per-sandbox cap for token-less forwards. CLI integration tests run
`service_forward_tcp` against a mock gateway: token-less inits echo data
with no CreateSshSession or RevokeSshSession call, and a gateway that
rejects the empty token is detected once, after which every connection in
that forward mints and revokes its own token. Architecture and security
docs describe which targets carry a token, and the per-token connection
limit now reads 3, matching the gateway.

Signed-off-by: Jason T. Greene <jason.greene@redhat.com>
2026-10-06 21:51:35 +00:00
Eric Curtin 3082a9ad7d fix(cli): plan sandbox uploads before provisioning (#4193)
Reject bad uploads before create. Share the upload transfer helper.

Signed-off-by: Eric Curtin <eric.curtin@docker.com>
2026-10-06 21:48:48 +00:00
Mike Nguyen 9fd41e6bfd fix(ssh): persist sandbox host identities (#4094)
* fix(ssh): persist sandbox host identities

Store each sandbox's Ed25519 host key in the gateway credential store and
deliver it only to the supervisor. Preserve identity across restarts,
delete owned credentials with the sandbox, and expose the public SHA256
fingerprint through sandbox and SSH-session APIs and client SDKs.

Cover credential ownership, cancellation, deletion retries, client
compatibility, and pinned SSH connections through lifecycle transitions.

Closes #3835

Signed-off-by: Mike Nguyen <miken@nvidia.com>

* fix(compute): clean up failed sandbox SSH identity creation

Signed-off-by: Mike Nguyen <miken@nvidia.com>

* test(ssh): wait for sandbox deletion before name reuse

Signed-off-by: Mike Nguyen <miken@nvidia.com>

---------

Signed-off-by: Mike Nguyen <miken@nvidia.com>
2026-10-06 20:09:20 +00:00
Drew Newberry 12cec59bf4 fix(sandbox): accept local connections natively on loopback-confined sockets (#4150)
* fix(sandbox): accept local connections natively on loopback-confined sockets

On kernels without SECCOMP_FILTER_FLAG_WAIT_KILLABLE_RECV (RHEL 9 / RHCOS
5.14), the broker cannot safely write an accepted peer address into
workload memory, so accept/accept4 with a peer-address buffer failed with
EOPNOTSUPP. Static binaries and Go servers, which issue the raw syscall,
could not accept connections at all.

Move local acceptance into the kernel and replace per-accept inspection
with standing kernel confinement:

- Bind every broker-created TCP/UDP socket to the loopback device before
  injection and verify the binding. Accepted sockets inherit it, so they
  can neither receive routed ingress nor emit routed egress.
- Stop notifying accept/accept4 and remove the accept workers, the
  SIGUSR2 accept-interrupt monitor, and the 64-accept ceiling.
- Continue getpeername natively for every descriptor except relayed
  connections.
- Deny interface-selection socket options and MSG_FASTOPEN sends from
  scalar syscall arguments, so confinement does not depend on the
  capability state of the user namespace that owns the network namespace
  and covers unregistered descriptors.
- Drop loopback-interface ingress on the non-loopback TCP control
  listener before it listens, so a workload cannot reach it by
  reconnecting a natively accepted socket.
- Mark descriptors above stdio close-on-exec in every workload pre_exec
  path.
- Require an active confinement probe at qualification and carry it as
  required authenticated audit evidence.

Document OpenShift 4.19 as the minimum release: RHCOS kernels for 4.16
through 4.18 are built without Landlock.

Closes #4058

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): close review gaps in native local accept confinement

- Reject a TCP boundary control listener on a loopback address,
  including IPv4-mapped loopback. Production drivers bind the
  unspecified address; a loopback listener would be reachable from
  workload sockets the broker does not track.
- Extend the confinement probe to IPv6 and prove that an accepted
  socket keeps its loopback binding after an AF_UNSPEC disconnect.
- Pin the accepted-peer contract with tests: a loopback-bound listener
  admits only clients in the sandbox network namespace, and workload
  sockets cannot bind a non-loopback source address.
- Correct the support matrix: accepted connections are bounded by
  per-process descriptor limits, the runtime PID limit, and the sandbox
  memory limit, not by a broker limit.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(sandbox): replace unsafe socket calls with socket2 and rustix

Socket confinement now uses socket2 for device binding and the ingress
filter, and rustix for interface lookup and AF_UNSPEC disconnect, so the
module contains no unsafe code. The control-listener filter is no longer
locked: no safe API exposes SO_LOCK_FILTER, and the listener descriptor
never leaves the trusted sandbox process, which marks every descriptor
above stdio close-on-exec before running workload code.

Tests added for native accept use rustix and socket2 instead of raw libc
calls. rustix issues accept4 and getpeername as raw syscalls, so the
direct-syscall test keeps its meaning. The close-on-exec sweep test runs
in a re-executed test process instead of a forked child.

The close_range(CLOSE_RANGE_CLOEXEC) syscall in the pre_exec hook remains
the only unsafe added by this branch; neither rustix nor nix wraps it.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): allowlist workload socket families

The workload seccomp filter denied only AF_PACKET, AF_BLUETOOTH, and
AF_VSOCK, and the broker continued socket() for every non-INET family.
Several protocol families, including AF_RXRPC, AF_SMC, and AF_KCM, carry
traffic over kernel-owned sockets that the broker never creates and that
are not bound to loopback.

Allow only AF_UNIX, AF_NETLINK (still limited to NETLINK_ROUTE), and the
brokered AF_INET/AF_INET6 families, and restrict socketpair(2) to
AF_UNIX. Both decisions use the scalar domain argument. The broker
independently refuses non-INET families other than AF_UNIX and
AF_NETLINK with EAFNOSUPPORT.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(sandbox): remove legacy read-only mode

After native local accept, only two broker paths still wrote into
workload memory: getpeername on relayed connections and the per-message
lengths of a first sendmmsg to the DNS relay. Legacy mode existed only to
refuse those writes on kernels without WAIT_KILLABLE_RECV, and the
getpeername refusal broke CPython TLS on RHEL 9.

The broker now never writes workload memory:
- getpeername is no longer mediated and reports the kernel peer on every
  kernel; relayed connections report the loopback relay address.
- A first send to the DNS relay pins the broker's socket copy to the
  relay and continues the syscall, so the kernel performs the send and
  writes any per-message results.

Remove the listener mode, the task-memory write path and probe, and the
task_memory_write, cancellation, and task_memory_writes_disabled audit
evidence fields. WAIT_KILLABLE_RECV is still used when available and is
reported for diagnostics, but no mediation decision depends on it.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(sandbox): drop WAIT_KILLABLE_RECV and make handlers restart-safe

The broker no longer writes workload memory, so killable notification
waits only reduced how often a signal restarts a notified syscall. RHEL 9
kernels never had them, so the handlers must tolerate restarts anyway.
Install the listener without the flag on every kernel instead of
special-casing newer ones.

Handlers now check that the notification is still live immediately
before each side effect (bind, connect of the retained socket, listen,
and the first DNS send), and answer a restarted operation the broker
already completed as the kernel would: a repeated TCP connect returns
EISCONN, repeating the same UDP association succeeds, a repeat of a
completed bind succeeds, and a failed relay reports its errno.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): finish teardown only when every workload descendant has exited

Teardown waited only for registered process groups to disappear. A root
is unregistered once reaped, so a descendant that ignored SIGTERM could
outlive it while termination reported success and never sent SIGKILL,
violating the bounded-termination requirement.

The sandbox now becomes a child subreaper when it is not PID 1, so
orphaned descendants stay in its tree and are reaped. Termination waits
until no live descendant remains, and every scanned process is signalled
through a pidfd after confirming its start time, so a reused PID is never
signalled.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): keep a frozen workload from resuming itself

Freezing stops every workload process with SIGSTOP. tgkill and
rt_tgsigqueueinfo were not mediated, and mediated kill, rt_sigqueueinfo,
and tkill passed SIGCONT through, so a workload process that was not yet
stopped could resume the others while the supervisor recovered.

Mediate tgkill and rt_tgsigqueueinfo like tkill: refuse targets in the
sandbox thread group, report a thread outside the named group as
missing, and continue otherwise. The boundary marks the broker frozen
before stopping the workload and clears it after resuming; while frozen,
workload requests to send SIGCONT fail with EPERM.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): send mediated DNS datagrams from the broker socket

The native UDP send path re-ran the workload's own sendmsg, so
per-message ancillary data such as IP_PKTINFO rode along; only the
loopback destination contained a routing override.

The broker now reads the datagram and sends it from its retained,
relay-connected socket, which it builds without ancillary data, so a
per-message override cannot redirect the packet. It writes nothing back
into workload memory. A send carrying control data is refused with
EOPNOTSUPP rather than silently stripped.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): keep sandbox control variables out of the workload environment

The canonical process inherits the sandbox environment so the image's
own ENV reaches the workload, then removed only a denylist of credential
variables. Other variables in the reserved OPENSHELL_ namespace, such as
the serialized user environment and the log level, still reached the
workload.

Remove every inherited OPENSHELL_ variable before applying the declared
environment, and restore OPENSHELL_SANDBOX=1. The image's ordinary ENV
and declared variables are unaffected; declared variables cannot use the
reserved namespace.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): keep /proc read-only for GPU workloads

GPU mode granted read-write access to all of /proc so CUDA's cuInit could
write thread names to /proc/<pid>/task/<tid>/comm.

Keep /proc read-only. Open mediation now serves a writable open of the
caller's own thread comm file: the broker opens it and injects the
descriptor, with no syscall continued. The kernel accepts a comm write
only from the target's own thread group, so a substituted path or reused
thread ID cannot rename another process's thread. Any other /proc write
is left to Landlock, which denies it.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): keep the broker socket for connected UDP DNS sockets

Sending DNS datagrams from the broker's socket required the broker to
keep its copy, but the connect paths to the relay and to loopback still
released it. glibc connects the resolver socket and then sends A and
AAAA together with sendmmsg, which is mediated, so resolution failed.

Release the copy on those connect paths only for TCP. Add a regression
test for connect followed by sendmmsg that fails fast rather than
hanging.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): keep slow loopback connects from stalling mediation

The loopback connect path held the registry lock while polling a TCP
connect for up to five seconds on the single notification dispatcher.
A workload connecting to a busy local listener stalled every other
mediated syscall, including opens and signals.

Connect a duplicate of the retained socket without holding the lock. A
nonblocking socket gets the native EINPROGRESS and the kernel completes
the handshake on the shared socket; a blocking socket waits on a bounded
worker thread.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): contain network broker handler panics

A panic in a notification handler unwound the single broker thread while
the health flag still read healthy, so every later blocked workload
syscall hung until the sandbox was killed.

Run each dispatch under catch_unwind: a panicking handler fails only that
syscall with EIO and the broker keeps mediating. If the broker thread
ever exits, mark it unhealthy so dependent operations fail closed instead
of blocking.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): reject ancillary data on DNS sends instead of broker-sending

Sending mediated DNS datagrams from the broker's own socket required
keeping a broker handle for each DNS socket's whole life, which regressed
real name resolution through reclaim and resource accounting that the
mock-based unit tests did not exercise.

Revert to continuing the kernel send, but refuse a send that carries
ancillary control data (msg_controllen != 0) at read time, so a
per-message routing override such as IP_PKTINFO cannot ride a mediated
DNS send. A loopback destination contains an override that races the
check. Full broker-side UDP mediation is left to a separate change with
deployment e2e.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): address review findings across the hardening changes

Fixes:
- The blocking local-connect worker no longer toggles O_NONBLOCK on the
  open file it shares with the workload; it waits with a plain connect.
- tgkill and rt_tgsigqueueinfo are notified only when they send SIGCONT,
  so ordinary thread signals such as Go preemption stay in the kernel.
  The frozen flag is re-read immediately before delivery.
- A shell redirect to the caller's own thread comm file works again;
  O_CREAT and O_TRUNC are no-ops there and O_EXCL returns EEXIST.
- The teardown scan is a linear walk and fails closed when /proc cannot
  be read.

Simplifications:
- Remove tests that need infrastructure outside the repo or prove
  nothing: the topology harness tests, the static-server benchmark, the
  direct-syscall accept test, the comm rename test without Landlock, a
  serde-default test, and the test-only panic hook in setsockopt.
- Drop the unused Failed connect outcome, the redundant read-back after
  binding to loopback, redundant OPENSHELL_SANDBOX settings, and stale or
  duplicated comments; reuse the reserved environment prefix constant.
- Tighten the support matrix and OpenShift wording.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(cli): wait for input when piped exec stdin is nonblocking

Processes that inherit the same stdin share one open file description,
so any of them can make it nonblocking for all. `sandbox exec` then read
no input yet, got EAGAIN, and failed with "Resource temporarily
unavailable (os error 11)". Parallel e2e tests inherit the runner's
stdin, which made credential_gating fail intermittently.

The stdin reader now blocks in poll() until input or end of file arrives
instead of treating EAGAIN as fatal, and it leaves the shared descriptor's
flags alone. The e2e exec helper also stops inheriting the runner's
stdin, since its commands take no input.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): fix macOS lint and older-glibc test linking

Import HashSet only in the Linux-only process scan, and call gettid
through the raw syscall in a test, since the libc wrapper needs glibc
2.30 or newer.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(sandbox): clarify socket peer address reporting

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-10-06 04:08:04 +00:00
Shiju 8983642e28 fix(gateway): give image preparation its own deadline (#4038)
* feat(server): separate image preparation and admission deadlines

Give sandbox image preparation its own deadline and start the admission
deadline after preparation finishes. Persist both phases across gateway
restarts and show recovery guidance for the phase that expired.

Preserve the current service authorization schema and regenerate the Go
bindings with the preparation timestamps.

Fixes #3952
Related to #3955

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(cli): simplify preparation timeout fallback selection

Use lazy Option fallbacks while preserving timeout messages and retained
sandbox behavior.

Signed-off-by: Shiju <shiju@nvidia.com>

* docs(server): clarify admission timer prerequisites

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(compute): enforce deadlines during initial sandbox create

Release stalled create operations after preparation expires and preserve
the timeout diagnosis across late driver results. Keep failed-create cleanup
bound to its original attempt so another replica can retry safely.

Signed-off-by: Shiju <shiju@nvidia.com>

* test(server): satisfy deadline regression lints

Drop the create-error mutex guard before matching its cloned value and use
idiomatic iteration and timeout matching in the deadline fixtures.

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(server): retain ownership of pending provisioning operations

Keep submitted create and start requests alive after caller cancellation,
monitor failure, or preparation timeout. Persist request ownership before
dispatch and retain staged uploads while the driver response is pending.

Require timeout cleanup to stop compute after driver settlement without
discarding an active cleanup claim. Fence late result handling against newer
operations, preserve failed-start recovery, and defer automatic restart while
another request owns the sandbox. Expose pending ownership in CLI JSON.

Add ordered multi-replica and cancellation regressions for the review findings.

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(openshell): preserve tracing and accept ready create responses

Carry the request span into the detached provisioning worker so compute
driver calls remain attached to their parent trace after task handoff.

Update the compensation regression to require no backend DELETE when the
durable cleanup claim fails, matching the operation ownership requirement.

Accept the gateway's current Ready snapshot when a sandbox becomes ready
before CREATE returns. Do not require the client to observe an earlier
provisioning phase. Cover the Ready-only watch and command attachment.

Signed-off-by: Shiju <shiju@nvidia.com>

---------

Signed-off-by: Shiju <shiju@nvidia.com>
2026-10-03 20:48:18 +00:00
Eric Curtin e7d14edd88 fix(relay): close outbound relay stream when target closes first (#3772)
* fix(relay): close outbound relay stream when target closes first

Closes #3724

out_tx was cloned into the target-reading task, so the function's
own copy kept the outbound relay stream open until the client side
also ended. A target that closed a keep-alive connection never
reached the client as EOF, so reused connections hung forever.

Move out_tx into the target-reading task instead of cloning it, so
dropping it on target EOF ends the outbound stream right away.

Signed-off-by: Eric Curtin <eric.curtin@docker.com>

* fix(relay): keep client uploads after target half-close

Signed-off-by: Eric Curtin <eric.curtin@docker.com>

---------

Signed-off-by: Eric Curtin <eric.curtin@docker.com>
2026-10-03 20:30:24 +00:00
Shiju 48d9ab3d0d fix(cli): check final exec status and warn on partial input (#3803)
* feat(cli): stream non-TTY exec input before EOF

Add --stream-stdin using the existing interactive exec RPC without a PTY. Preserve separate output streams and enforce the existing 4 MiB cumulative input cap while forwarding input.

Require explicit clean stdin EOF and drain the response through its final gRPC status. Cover held-open input, limits, cancellation, trailers, and default finite-input behavior with subprocess and live sandbox regressions.

Signed-off-by: Shiju <shiju@nvidia.com>

* test(cli): distinguish exec cancellation from stdin EOF

Treat transport termination and response cancellation as separate test observations. Verify explicit stdin EOF through the shared frame writer and cover cancellation in the pinned Tonic decoder.

Signed-off-by: Shiju <shiju@nvidia.com>

* style(cli): use lazy optional stdin dispatch

Signed-off-by: Shiju <shiju@nvidia.com>

* test(cli): use imported duration in stdin EOF regression

Signed-off-by: Shiju <shiju@nvidia.com>

---------

Signed-off-by: Shiju <shiju@nvidia.com>
2026-10-02 22:10:01 +00:00
alangou 36819f476d fix(cli): stop uploads when Git filtering fails or selects no files (#3957)
Signed-off-by: Adrien Langou <alangou@nvidia.com>
2026-10-02 14:24:32 +00:00
Fede Kamelhar 71440b28f4 fix(policy): refresh pending proposals when the sandbox policy changes (#3923)
* fix(policy): refresh pending proposals when the sandbox policy changes

Approving, removing, or undoing a rule, or updating the sandbox policy,
changes the inputs every other pending proposal was evaluated against.
Only proposals the new policy covered were reconciled; the rest kept
their old prover result and review token. The review surface
(GetDraftPolicy) therefore showed a stale evaluation, and the first
approval of the next proposal refreshed it and failed with
FAILED_PRECONDITION, so approving proposals one after another always
failed once.

Re-evaluate the remaining pending proposals at each policy change,
reusing the cached prover result unless the proposal's inputs changed.
Approval still rejects a review token that does not match the stored
evaluation, so a reviewer holding a pre-refresh evaluation must still
refetch it.

When a refresh does happen at approval time (inputs changed between
fetch and approve), the CLI now explains that the rule was re-evaluated
and how to review it, instead of printing the raw gRPC status.

Closes #3884

Signed-off-by: fede-kamel <fkamelhar@gmail.com>

* fix(policy): make pending proposal refresh race-safe and bounded

Store refreshed evaluations with a compare-and-swap: the store re-reads
the proposal, refuses when its rule name, proposed rule, or review token
changed since the evaluation read it, copies only the evaluation fields
onto the stored record, and updates only if the payload is still the one
it read. A refresh can no longer revert a concurrent edit or observation,
and the edit path uses the same guard against a concurrent refresh.

Bound each refresh to the 32 newest pending proposals; the rest keep the
approval-time recheck, which still refuses a stale review token. Operator
decisions (approve, approve-all, remove, undo) refresh before responding.
UpdateConfig, which holds the gateway-wide sandbox sync guard, and
agent-driven auto-approval refresh in a background task instead, one per
sandbox with later changes coalesced into a single rerun.

Refs #3884

Signed-off-by: fede-kamel <fkamelhar@gmail.com>

* docs(policy): describe proposal rechecks after approvals and approve-all

Explain that approving, removing, or undoing a rule rechecks the other
pending proposals so they can be approved one after another, when the
recheck is deferred or bounded, and what rule approve reports when a
proposal changed after it was listed. Show rule approve-all in Run Your
First Agent with its security-flag behavior.

Refs #3884

Signed-off-by: fede-kamel <fkamelhar@gmail.com>

* fix(policy): refresh pending proposals after a full policy replacement

A full policy UpdateConfig (openshell policy set) re-reads the latest
revision after its atomic write, finds the revision it just committed,
and returns before reaching the pending-proposal refresh at the end of
the handler. Pending proposals kept their stale evaluation, so rule get
showed the old candidate and the next approval failed with the refresh
precondition. Schedule the background refresh right after the commit.

Refs #3884

Signed-off-by: fede-kamel <fkamelhar@gmail.com>

---------

Signed-off-by: fede-kamel <fkamelhar@gmail.com>
2026-10-01 16:24:51 +00:00
Fede Kamelhar ffcbe6280c fix(cli): start sandbox exec without waiting for piped stdin EOF (#4006)
With a non-terminal stdin, sandbox exec read stdin to EOF before it sent
the exec request. A pipe that never closes (CI runners, supervisors, agent
harnesses) blocked the CLI forever in read(2) without the gateway ever
seeing the request, and a slow producer delayed the command until EOF.

Collect piped stdin on a detached reader thread for at most 200 ms. Input
that reaches EOF within that window still travels in the single request
that older gateways need. If the pipe is still open, start the command
through the streaming RPC and forward the collected prefix plus the rest of
stdin as it arrives, closing remote stdin at EOF. The 4 MiB cap covers the
prefix and the streamed remainder together.

Closes #3993

Signed-off-by: Federico Kamelhar <federico.kamelhar@oracle.com>
2026-10-01 16:11:49 +00:00
Drew Newberry fe38637533 fix(runtime): recover SSH relays and bound startup diagnostics (#4011)
* fix(runtime): recover SSH relays and bound startup diagnostics

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(server): deliver pending relays once per supervisor session

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(server): satisfy relay delivery clippy diagnostics

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(server): bound relay setup with one absolute deadline

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-10-01 12:47:20 +00:00
Derek Carr 912a077bd6 feat(service): add bearer authorization passthrough (#3796)
* feat(service): add bearer authorization passthrough

Signed-off-by: Derek Carr <decarr@redhat.com>

* docs(sdk): add service authorization migration guide

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(server): remove stale version import

Signed-off-by: Derek Carr <decarr@redhat.com>

* docs(upgrade): remove service authorization SDK guide

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(e2e): relabel provider readiness TLS mount

Signed-off-by: Derek Carr <decarr@redhat.com>

* test(e2e): stabilize exposed service routing

Signed-off-by: Derek Carr <decarr@redhat.com>

* test(e2e): support HTTPS service routing

Signed-off-by: Derek Carr <decarr@redhat.com>

---------

Signed-off-by: Derek Carr <decarr@redhat.com>
2026-09-30 20:22:46 +00:00
Eric CurtinandDrew Newberry 07a486d751 fix(cli): accept sandbox name before -- in exec (#3901)
* fix(cli): accept sandbox name before -- in exec

Closes #3882

Signed-off-by: Eric Curtin <eric.curtin@docker.com>

* fix(cli): define exec grammar in clap

Signed-off-by: Eric Curtin <eric.curtin@docker.com>

* docs(sandboxes): remove exec overview change from PR

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Eric Curtin <eric.curtin@docker.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
2026-09-30 18:50:04 +00:00
Shiju ba16b9f2c7 fix(mcp): explain revision-scoped policy and rejections (#3850)
* fix(mcp): explain revision-scoped policy and rejections

Explain the selected-revision method set in profile output and policy docs.
Distinguish protocol and policy rejection causes and give a next step while
preserving authorization, response statuses, error codes and YAML keys.

Cover CLI serialization, revision selection, exact extension rules, deny
precedence and rejection before forwarding with focused regressions.

Signed-off-by: Shiju <shiju@nvidia.com>

* docs(mcp): correct HTTP cancellation revision support

Limit notifications/cancelled to the three 2025 revisions in the core
method matrix. State that MCP 2026-07-28 HTTP cancellation closes the
response stream, matching the runtime rejection and sessionless docs.

Signed-off-by: Shiju <shiju@nvidia.com>

---------

Signed-off-by: Shiju <shiju@nvidia.com>
2026-09-29 21:50:35 +00:00
Shiju 12ef86c285 fix(cli): keep policy and provider diagnostics readable (#3444)
Reuse character-safe truncation for policy history errors so a multibyte character cannot panic the table renderer.

Distinguish unavailable provider-profile YAML from an absent profile, display a bounded diagnostic, and preserve strict serialization and redacted object navigation.

Cover the actual CLI renderer and TUI display/navigation paths, including invalid and absent profiles, Unicode input, and redacted errors.

Signed-off-by: Shiju <shiju@nvidia.com>
2026-09-29 15:35:56 +00:00
Drew Newberry acbac9cb79 feat(sandbox): add main restart policy (#2798)
* feat(sandbox): add main restart policy

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): harden policy-driven restarts

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): address restart review feedback

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>

* fix(sandbox): port restart policy to current runtime

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* perf(sandbox): restart promptly after terminal delivery

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>
2026-09-29 02:18:30 +00:00
Drew Newberry e63cfa1182 fix(cli): keep SSH forwards owned by spawned process (#3759)
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-28 19:34:13 +00:00
Drew Newberry 36b0386c92 feat(cli): detach sandbox sessions with Ctrl-D (#3744)
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-28 09:36:14 -07:00
Drew Newberry 9244868056 docs: refresh architecture and agent guides (#3705)
* docs: refresh architecture and agent guides

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs: describe updated security architecture neutrally

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs: highlight new isolation primitives

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(sandboxes): clarify how to disconnect

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs: align architecture and guides with current navigation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(extensibility): streamline extension authentication guidance

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-25 02:38:00 -07:00
Drew Newberry c93fd94a5f feat(cli): import provider profiles from HTTP URLs (#3706)
* feat(cli): import provider profiles from HTTP URLs

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(cli): initialize crypto provider for remote profiles

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-25 06:12:18 +00:00
Drew Newberry 7a50c0899f fix(cli): stream piped exec stdin beyond gRPC request limit (#3687)
* fix(cli): stream piped exec stdin across gRPC messages

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(cli): preserve small exec requests and surface stdin errors

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(cli): preserve exec stdin limit across gRPC streaming

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-25 04:45:57 +00:00
Drew Newberry 0518bd4c83 fix(sandbox): preserve local sessions across host sleep (#3573)
* fix(cli): recover sandbox connect transport

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(auth): propagate non-expiring local sandbox sessions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(cli): distinguish main exit from transport loss

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(cli): bound sandbox connect recovery

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-24 17:22:30 +00:00
Gaizka Menendez 123d95e2ed fix(pagination): document list contract and harden SDK pagers (#3279)
* fix(pagination): document list contract and harden SDK pagers

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* fix(pagination): bound pager token history

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* fix(pagination): preflight token history limits

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* fix(pagination): paginate sandbox providers

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* fix(pagination): harden TypeScript pager

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* fix(pagination): expose provider pagers in SDKs

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* docs(pagination): describe a uniform list contract

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* fix(pagination): use stable provider cursors

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* docs(pagination): clarify mutation semantics

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* fix(cli): expose sandbox provider pagination

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* fix(proto): refresh pagination schema inventory

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* fix(proto): refresh rebased schema inventory

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>
Signed-off-by: Kris Hicks <khicks@nvidia.com>

---------

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>
Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-09-23 18:51:59 +00:00
Mrunal Patel bdffa102c3 feat(api): add durable exec launch admission (#3324)
* feat(api): add durable exec launch admission

Fence duplicate exec launches with keyed durable admission and producer-owned terminal completion. Keep uncertain launches unresolved and never replay output or interactive input.

Part of #3051 (phase 4a).

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(api): fence exec identity across authorization lookups

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

---------

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
2026-09-22 21:52:19 +00:00
Drew Newberry 1e34e8c576 fix(drivers): require admission labels for external resources (#3538)
* fix(drivers): require admission labels for external resources

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(drivers): address resource admission review findings

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(core): reserve driver-owned admission labels

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(core): clarify workspace admission label

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(drivers): clarify resource admission failures

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(e2e): configure resource admission fixtures

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): retry forbidden admission lookups

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(e2e): preserve external driver admission defaults

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-22 21:22:30 +00:00
Artem LytvynandJohn Myers 470a34635d fix(api): make WatchSandbox loss-aware and resumable (#3209)
* fix(api): emit warning on WatchSandbox broadcast lag instead of terminating

Broadcast lag on the status, log, and platform receivers was converted to a RESOURCE_EXHAUSTED status that terminated the whole watch stream. Lag is recoverable: the receiver resumes at the oldest surviving message. Emit a SandboxStreamWarning and continue streaming instead; keep terminating on Closed. Add helpers and unit tests covering the warning payload and receiver recovery after lag.

Partially addresses #3055 (cursor/resume follow up separately).

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* refactor(server): group per-sandbox log bus state and stamp sequence numbers

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* feat(proto): add resume cursor fields to sandbox watch API

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* feat(server): stamp watch cursors from a shared per-sandbox sequence

Allocate cursors from a single SeqAllocator shared by the log and
platform event buses, so a sandbox's merged watch stream carries
unique, strictly increasing cursors. A single resume_after_cursor can
then unambiguously locate a client's position across both sources.

Rewrite both publish paths to allocate the sequence, stamp
event.cursor, send, and append to the tail under one lock. This
removes the previous get_mut().expect() TOCTOU race where a concurrent
remove() between the two lock sections could panic.

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* feat(server): serve WatchSandbox resume from cursor with gap detection

Add tail_after() to the log and platform event buses, returning every
buffered event newer than a client's resume cursor. Each PerSandbox now
tracks last_trimmed_seq (the highest seq it has evicted) so a resume is
reported as an unrecoverable ResumeGap only when this bus dropped an
event the client still needs.

Judging gaps by evictions, not by the tail's oldest seq, is required
under the shared cursor space: each bus's tail is non-contiguous in the
global sequence because the other bus owns the missing seqs, so
comparing against tail.front() would flag false gaps.

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* feat(server): resume WatchSandbox from cursor across log and platform buses

Wire resume_after_cursor into the watch producer. On a non-zero cursor,
replay events strictly after it from both the log and platform buses,
merge by shared cursor, and emit in order before entering the live loop.
A trimmed range on either bus is an unrecoverable gap and terminates the
stream with OUT_OF_RANGE carrying the requested and earliest-available
cursors, distinct from recoverable lag which warns and continues.

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* test(server): cover WatchSandbox cursor resume paths

Add handler-level tests for the resumable watch stream: replay strictly
after the client cursor, merge log and platform events in shared-cursor
order, suppress duplicates when resuming at the latest cursor, and
terminate with OUT_OF_RANGE when the requested cursor has been trimmed.

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* docs(api): document WatchSandbox loss-awareness and resume

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* fix(server): deliver watch events once and harden cursor teardown

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* feat(sdk): add loss-aware resumable watch_logs to Rust SDK client

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* fix(server): keep watch cursors monotonic across teardown and restart

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* fix(server): merge live watch sources by cursor before emission

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* fix(api): bind watch cursors to a cursor space and merge tail sources

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* fix(server): revalidate the watch cursor space after collecting replay

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* test(server): update the public RPC schema fingerprint for the string cursor

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* fix(server): hold watch events above the publication watermark and emit the watch lag warning before its batch

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* test(server): synchronize the watch live-order test with the end of initialization

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* fix(sdk): use canonical sandbox name in watch_logs

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* test(sdk): guard canonical-name addressing in watch_logs

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* fix(server): fix public rpc schema

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* fix(api): reconcile watch resume rebase

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(server): bound interactive relay cleanup

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-09-22 19:45:55 +00:00
Yuedong Wu 718dba3430 fix(policy)!: reject removed tls endpoint values (#3414)
Signed-off-by: Yuedong Wu <dwcn22@outlook.com>
2026-09-22 17:58:19 +00:00
50230616d5 refactor(runtime): retire Community image dependencies (#3386)
* feat(sandbox): default to official Alpine sandbox image

default_sandbox_image() now returns docker.io/library/alpine:3.22, a generic
version-qualified official image, so a fresh install no longer depends on the
community sandbox image catalog. All compute drivers (docker, podman,
kubernetes, vm) inherit this fallback.

Part of #3116.

Signed-off-by: Akram
Signed-off-by: Akram <akram.benaissi@gmail.com>

* feat(deploy): default deployment configs to the official Alpine sandbox image

Update the shared gateway default_image, Helm chart values, the standalone
Kubernetes manifest, and the dev gateway task scripts to use
docker.io/library/alpine:3.22 instead of the community base image, consistent
with default_sandbox_image(). GPU e2e image-build base is left unchanged (CUDA
needs a glibc base).

Part of #3116.

Signed-off-by: Akram
Signed-off-by: Akram <akram.benaissi@gmail.com>

* feat(driver): default to numeric non-root identity for USER-less images

With the default sandbox image now Alpine, images that declare no OCI USER
must start instead of being rejected. When the image declares no USER and
the policy requests none, the Podman and Docker drivers now supply a numeric
non-root identity (DEFAULT_SANDBOX_UID/GID = 1000) instead of rejecting,
matching the numeric-identity behavior of the Kubernetes and VM drivers. The
supervisor's resolved-identity path runs the sandbox as a synthesized
non-root account without the account existing in the image. Images that
declare a USER keep the OCI resolution path unchanged.

Part of #3116.

Signed-off-by: Akram <akram.benaissi@gmail.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(conformance): use Alpine workload image

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* refactor(policy): drop community image /app path from default policy

The restrictive default policy granted read-only access to /app, a directory
that only existed in the community base image. A generic Alpine default has no
/app, so remove it. Landlock best-effort already ignores absent paths; this
just stops advertising a community-specific layout in the default.

Part of #3116.

Signed-off-by: Akram
Signed-off-by: Akram <akram.benaissi@gmail.com>

* docs(config): document Alpine default images

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(podman): report early sandbox termination

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(podman): initialize rootless workspace ownership

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(sandbox): qualify NVIDIA Ubuntu default

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): initialize rootful default workspace

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(sftp): add native sandbox adapter

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sftp): gate runtime helper support to Linux

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sftp): support standard OpenSSH file operations

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sftp): harden rename and special file handling

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(runtime): remove community image dependencies

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(e2e): build provider readiness tool fixture

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(e2e): use a dedicated Noble fixture for Docker tests

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Akram
Signed-off-by: Akram <akram.benaissi@gmail.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Evan Lezar <elezar@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
2026-09-22 14:43:51 +02:00
0a770d9173 feat(kubernetes): support HA gateway rebalancing (#1868)
* feat(kubernetes): support HA gateway rebalancing

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* perf(server): cache peer connections, tokens, and owner lookups

Every forwarded relay rebuilt its setup from scratch: an owner lookup, a
blocking read of the peer token, a TLS connect to the owning replica, and
a TokenReview plus Pod GET on the receiving side. Sandbox service routing
does this per HTTP request, so the apiserver calls scaled with traffic.

Cache all of it on ServerState:

- peer channels pooled per endpoint, so relays multiplex over one
  connection instead of redialing
- peer tokens keyed by SHA-256, expiring at min(ttl, token exp) so a hit
  cannot accept an expired token
- owner records for 3s against a 45s ownership TTL, still freshness
  checked before use

Entries are evicted when a relay fails. Also raise HTTP/2
max_concurrent_streams to 1024, since pooling funnels every relay between
two replicas onto one connection and hyper's default of 200 sits below
the 256 pending-relay budget.

Signed-off-by: divesh <dgude@nvidia.com>

* perf(server): pool upstream connections for sandbox services

Each HTTP request to a sandbox service opened its own supervisor relay,
paying a new TCP connection and HTTP/1 handshake every time. Worse, it
counted against the 32 in-flight relay cap, so a service handling more
than 32 concurrent requests failed outright.

Pool idle upstreams per endpoint and port, up to 8 each for 15s. Reuse is
safe because the pool only returns a connection hyper reports as ready,
and HTTP/1 cannot start a request until the previous body has drained.
Upgrades are never pooled since they take the connection over, and a
failed send evicts that endpoint. Pruning is bounded per key, with the
full sweep limited to once per 30s.

Signed-off-by: divesh <dgude@nvidia.com>

* fix(server): address HA gateway review findings (#3449)

- Let a gateway own supervisor sessions without a peer endpoint. Requiring
  one whenever the store is PostgreSQL broke every single-instance
  PostgreSQL deployment, because no sandbox supervisor could connect.
  A cross-replica request to an owner that advertises no endpoint now fails
  immediately naming the cause, instead of retrying until the wait timeout.
- Close a supervisor session on heartbeat only when another replica owns it,
  or after renewals fail for the ownership TTL. A database error no longer
  drops every session heartbeating during an outage.
- Clamp owner record ages at zero so a skewed or corrupt stored timestamp
  cannot produce a negative age.
- Bound the cross-object advisory lock with a lock timeout, so a stuck holder
  fails instead of blocking every mutation in the fleet.
- Refuse to start when a peer endpoint is configured on a multi-replica
  backend but peer authentication is unavailable, and warn when a
  multi-replica backend has no peer endpoint at all.
- Reject a plaintext peer endpoint when the gateway serves TLS.
- Skip the sandbox watch poller on single-replica backends, where the local
  update bus already sees every write.
- Rate-limit the peer owner cache sweep so an insert no longer scans the
  whole map under the lock.
- Retry GET and HEAD on a pooled upstream the sandbox closed, instead of
  returning 502, and drop an emptied endpoint from the pool right away.
- Document the gateway peer environment variables and the post-rollout
  ownership skew operators should expect.

Signed-off-by: divesh <dgude@nvidia.com>

* fix(server): harden HA supervisor ownership

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Signed-off-by: divesh <dgude@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: divesh <dgude@nvidia.com>
Co-authored-by: Divesh Chowdary <47188680+FrostGod@users.noreply.github.com>
2026-09-21 21:22:12 +00:00
Seth Jennings 2493d415c2 feat(extensions)!: normalize protocol negotiation (#3352)
* feat(extensions)!: normalize protocol negotiation

Closes #3057

Introduce a shared extension handshake, enforce protocol and capability compatibility across extension families, and expose immutable negotiated snapshots through gateway info and the Go SDK.

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(credentials): fail fast on negotiation errors

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(extensions): validate gateway handshake metadata

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(go-sdk): re-export extension kind constants

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(extensions): fail fast on credential handshake rejection

Signed-off-by: Seth Jennings <sjenning@redhat.com>

---------

Signed-off-by: Seth Jennings <sjenning@redhat.com>
2026-09-21 13:19:39 -07:00
Shiju fa8f6d3949 feat(cli): promote profile commands to top level (#3258)
* feat(cli): promote profile commands to top level

Add profile discovery and management commands with shared handlers for
the existing provider entry points. List a flat catalog across scopes
and follow continuation tokens through full and short pages.

Describe metadata, credentials, endpoints, TLS inspection, and MCP access
settings while preserving complete JSON/YAML definitions. Cover parser
equivalence, scope forwarding, pagination, and inspection settings with
focused unit and compiled-CLI integration tests.

Update docs, public skills, examples, and E2E command invocations.

Refs #2588

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(cli): remove redundant workspace selector qualification

Use the imported WorkspaceSelector in the provider integration helper so
the target passes Clippy with warnings denied.

Signed-off-by: Shiju <shiju@nvidia.com>

---------

Signed-off-by: Shiju <shiju@nvidia.com>
2026-09-21 17:18:55 +00:00
Drew Newberry 1905069948 feat(sandbox): expose services during creation (#3439)
* feat(sandbox): expose services during creation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(providers): refresh Codex credentials in gateway

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(cli): normalize create-time service URLs

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(server): roll back failed service exposure

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(example): simplify Codex provider setup

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(example): separate provider setup commands

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): harden create-time service exposure

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(example): bundle Codex provider profile

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(example): allow npm-installed Codex binary

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-20 23:19:25 -07:00
Drew Newberry d91b1999a0 feat(api)!: use sandbox names as canonical RPC references (#3272)
* feat(api)!: use sandbox names as canonical references

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(cli): update forward color fixture for workspace scope

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(supervisor): use sandbox names for settings lookup

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(e2e): use canonical sandbox request fields

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(e2e): use canonical sandbox receipt field

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(server): harden sandbox mutation handling

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(ci): update rebased sandbox references

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(api)!: standardize canonical entity references

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(api): codify protobuf API conventions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(api): preserve workspace selector semantics

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(api): restore workspace selector parity

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(api): preserve descriptive name fields

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(api): update e2e request fixtures

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(core): omit workspace selector during bootstrap

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* chore(api): remove proto convention checker

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* chore(api): refresh schema fingerprints after rebase

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(cli): use canonical provider receipt field

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-18 11:39:55 -07:00
John T. Myers 903d9a0e7a chore(license): align repository compliance text (#3467)
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-09-18 17:05:18 +00:00
Krzysztof Malczuk e38d7254e6 fix(policy): reject unknown endpoint security modes (#3187)
* fix(policy): reject unknown endpoint security modes

Closes #3046

Validate TLS, enforcement, and access values across policy and provider profile ingress, and prevent runtime parsing from falling back to audit for unknown enforcement values.

Signed-off-by: Krzysztof Malczuk <kmalczuk@redhat.com>

* fix(policy)!: use enums for endpoint security modes

Replace the public TLS, enforcement, and access strings with protobuf enums and carry the typed values through policy composition, provider profiles, drivers, and runtime conversion.

Preserve the documented YAML spellings, reject unknown and invalid numeric enum values consistently, and update generated Go bindings, SDK conversions, tests, and policy documentation.

Signed-off-by: Krzysztof Malczuk <kmalczuk@redhat.com>

---------

Signed-off-by: Krzysztof Malczuk <kmalczuk@redhat.com>
2026-09-18 16:48:33 +00:00
John T. Myers 1d010f4187 feat(sandbox): validate configuration before workload activation (#3259)
* feat(sandbox): validate configuration before workload activation

Closes #3145

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(sandbox): bound startup failures and preserve activation history

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(sandbox): enforce deadlines on startup RPC attempts

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(sandbox): isolate provider auto-create policy fixture

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(sandbox): remove unnecessary fixture string delimiters

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(sandbox): capture startup logs before ephemeral cleanup

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(sandbox): preserve credential revocation after rebase

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(server): retain provider revision helper for endpoint reports

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(server): import provider object trait in production

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(supervisor): preserve admission across boundary extraction

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(sandbox): use existing boundary discovery imports

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(sandbox): distinguish workload and supervisor containers

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(server): reconcile admission schema with timestamp migration

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(server): reconcile admission with typed deletion schema

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(server): reconcile admission with mutation request IDs

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(supervisor): adapt local startup fixture to admission state

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(supervisor): deduplicate startup quarantine diagnostics

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(tui): show configuration blockers in sandbox notes

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* feat(sandbox): expire provisioning repair attempts after five minutes

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(supervisor): reconcile admission with provider readiness

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(tui): separate configuration summaries from full diagnostics

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(tui): shorten invalid configuration note

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(server): refresh schema inventory after rebase

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-09-18 00:16:18 +00:00
Philippe MartinandJohn Myers 04146692d9 refactor(providers)!: make provider profiles import-only (#3383)
* chore(providers): remove dead provider plugin modules

Twelve modules under crates/openshell-providers/src/providers/ were never
declared in providers/mod.rs, so they have not been compiled since the plugin
registry was narrowed to the two adapters it still registers. Four of them
(generic, gitlab, opencode, outlook) key off provider type identifiers that
normalize_provider_type already retired.

Keep google_cloud and vertex, which are the only plugins ProviderRegistry::new
registers.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* test(providers): load example profiles from providers/ at test time

Add an example_profiles module that reads the YAML under providers/ from the
source checkout at run time and parses it with the existing profile loader. It
is gated behind a new non-default example-profiles feature so it is available to
this crate's own tests and, once wired into dev-dependencies, to gateway and CLI
tests, while never reaching a release binary.

Nothing consumes it yet; later changes move the test fixtures off the compiled
catalog and onto these files.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* test: source profile fixtures from providers/ instead of the compiled catalog

Unit and integration fixtures reached into builtin_profiles() to get a profile
to work with, which ties the tests to the compiled catalog rather than to the
files an operator would import. Point them at the example_profiles loader
instead.

The profiles.rs tests keep their coverage unchanged and become explicit golden
tests over providers/*.yaml: they are what keeps those files valid once nothing
compiles them. builtin_profiles_are_sorted_by_id widens into a test that the
whole example set parses, sorts, and lints clean as one catalog.

The CLI fake gateways serve the example profiles the way a real gateway serves
what an operator imported, through a shared helper.

No behavior change: the gateway still loads the built-in source by default and
serves the same profiles.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* docs(providers): document the example provider profiles

Each file under providers/ now opens with a header naming its expected client
binary identities, the image layout those paths assume, the credential scope,
the endpoint access it grants, and a smoke test. Add a README covering the
import commands and why a profile should be copied and edited rather than
imported unchanged.

Several of these profiles bind network access to paths that only exist in the
OpenShell Community image — /sandbox/.venv, /app/.venv, /sandbox/.cursor-server,
/usr/lib/node_modules. Imported unchanged into another image the profile matches
nothing: the catalog still advertises it, but the credential is never injected
and the traffic is denied. The headers say so where it applies.

Comments only; the profile schema has no documentation fields and all fifteen
files still lint clean.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* test(server): seed unit-test state with the example provider profiles

Server unit tests inherited the builtin + user source default from ServerState,
so around a hundred and forty assertions about github, openai and the rest
resolved against the compiled catalog. Point the test state at the user source
alone and import the example profiles from providers/ into its store first, the
way an operator would. Every one of those assertions keeps passing unchanged,
which is the point: it proves the gateway behaves identically with an imported
catalog before the default moves.

Four tests asserted builtin-source semantics specifically. Under import-only the
only profiles a gateway cannot edit are the ones a non-user source vends, so
they now exercise a source-managed profile composed with the user source; the
read-only guard they cover is the one that still applies to interceptor
catalogs. The list test asserts the imported profile's user/platform identity
instead of builtin with an empty scope.

test_server_state_with_user_only_github_profile is gone: the default test state
now is a user-only gateway with github imported.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* refactor(cli): resolve credential suggestions from the gateway catalog

The --env credential warning scanned a profile table compiled into the CLI, so
its suggestions described the binary rather than the gateway the user is talking
to: a profile the gateway does not serve was suggested anyway, and an imported
custom profile never was.

Fetch the catalog from the connected gateway instead. The warning moves out of
argument parsing and into sandbox create and sandbox template create, where a
client already exists. A catalog fetch failure is not fatal — the warning
degrades to its generic form rather than blocking sandbox creation.

Extract the ListProviderProfiles paging loop from provider list-profiles into a
shared fetch_provider_profile_catalog, which the profile-driven paths now share.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* refactor(cli)!: infer providers from the gateway catalog

Command-to-provider inference went through a hardcoded alias table in
openshell-providers, a second copy of the built-in catalog's identifiers. A
custom profile could never be inferred no matter what binaries it declared, and
the table drifted from the profiles it mirrored.

Infer from the connected gateway's catalog instead: match the command's basename
against each profile's ID and against the basenames of the binaries the profile
authorizes. A profile that names /usr/bin/claude is the profile for running
claude. The match must be unique — where several profiles claim a command, the
user names one with --provider — and an empty catalog infers nothing.

No command is special-cased. `binaries` is the operator's authorization
statement, so a profile that declares a binary claims the command that runs it,
whatever that binary is; narrowing that belongs in the profile rather than in a
list compiled into the CLI, which could never cover an unbounded catalog anyway.

Breaking: the retired aliases stop resolving, and commands the old table never
listed can now infer. git, pip and uv are declared by the github and pypi
example profiles, so they infer where they previously did not.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* refactor(providers,server)!: resolve provider profiles by exact ID

normalize_provider_type was a hardcoded alias map — gh to github, claude to
claude-code, vertex to google-vertex-ai — and a second copy of the built-in
catalog's identifiers, independent of the YAML it mirrored. It let a provider
type resolve to a profile the operator never named, and it made the built-in IDs
behave as a reserved namespace.

Remove it, along with detect_provider_from_command and the alias-normalizing
ProviderRegistry::inject_env. A provider type now names a profile exactly:

- the effective catalog resolves an ID or reports it absent, with no alias retry
- plugins activate only for the ID of a profile the gateway resolved; a provider
  with no resolvable profile gets no plugin projection, instead of falling back
  to an alias guess
- the CLI surfaces the gateway's not-found instead of retrying under an alias
- telemetry buckets by profile ID, and the gitlab, opencode and outlook buckets
  go with the aliases that were their only source

Breaking: `--type gh`, `--type claude` and the other aliases no longer resolve.
Use the profile's own ID.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* feat(server): fail closed when a provider's profile is absent

A sandbox composed from a provider whose profile the gateway cannot resolve
started anyway: the credential and policy builders warned and skipped, so the
sandbox came up carrying none of that provider's credentials or network policy.
The operator learned about it later, as a denied connection or a missing
environment variable, rather than as the configuration error it is.

Check at the two composition boundaries — CreateSandbox and
AttachSandboxProvider — right beside the catalog snapshot already taken there,
and reject with a bounded diagnostic naming the provider, the profile it refers
to, and the import command that supplies it. Scope-aware: a platform-scoped
provider is told to import with --global.

Read paths are untouched. ListProviders, GetProvider and profile export keep
working so an operator can see and recover an affected provider, and the shared
policy builders keep their warn-and-skip for the diagnostic paths that also
reach them.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* refactor(server): scope the vendor base-URL pin to declared endpoints

provider_profile_endpoints_are_active withheld a credential and the provider's
policy layer when an openai or anthropic provider pointed its client somewhere
other than the public vendor endpoint. The guard was keyed on
profile.source == "builtin" and on those two profile IDs, so it protected only
profiles OpenShell shipped. Once profiles are import-only no profile is ever
builtin, and the guard would silently stop applying — including to an operator
who imported providers/openai.yaml verbatim.

Key it on the profile instead of on where the profile came from. A profile's
endpoints are the boundary its credential is bound to, so if the provider
configures a *_BASE_URL pointing at a host the profile does not declare, the
profile no longer describes where that credential goes and is treated as
endpointless. Host matching reuses the DNS-label-aware matcher in
openshell-core, so wildcard endpoints such as Vertex's
*-aiplatform.googleapis.com resolve correctly.

A profile with no declared endpoints has no boundary to contradict, and a config
value that names no host is not a redirect this can reason about. Both keep the
profile active. The control now covers every endpoint-bearing profile, including
an operator's own, and the last "builtin" string leaves the gateway.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* test(e2e): import example provider profiles during gateway bring-up

The e2e suites create providers from github, openai, nvidia, claude-code and
google-cloud, and the GitHub lane runs a real git clone through the profile's
binary attribution. A gateway serves only the profiles an operator imported, so
the lanes have to import them.

Add e2e_import_example_provider_profiles to the shared bring-up helpers and call
it from the Docker, Podman, Kubernetes and VM wrappers once the gateway is
healthy and registered. It runs the same command the upgrade notes give
operators, against the repository's own providers/ directory, so the lanes
exercise the documented path rather than a test-only shortcut.

Lands before the default changes: a same-ID user profile already shadows a
built-in, so importing works today.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* refactor(providers)!: make provider profiles import-only

OpenShell compiled fifteen provider profile YAML files into every release binary
and selected them by default, so a fresh gateway published a catalog it was
never configured with. Those profiles are not image-neutral: their binary
selectors name paths that exist in the OpenShell Community image, so changing
the sandbox image could make a profile inert while the catalog still advertised
it. Every endpoint, credential name and binary path in providers/ was also
effectively part of the 0.1.0 public contract.

A gateway's catalog is now exactly what an operator imported:

- providers/*.yaml is no longer include_str!'d, and builtin_profiles() is gone
  from the openshell-providers API
- the default provider_profile_sources is [{ type = "user" }], and a gateway
  with nothing imported reaches ready and serves an empty catalog — an empty
  catalog is a valid state, not a startup failure
- the builtin source type is removed from the configuration schema, and a
  gateway.toml that still names it is rejected at parse time with the import
  command rather than an unknown-variant error
- nothing reserves the canonical identifiers any more, so github, pypi,
  anthropic and the rest import at their own IDs and the imported profile is the
  only definition for that ID; the static_fallback that kept a shipped
  definition resident behind an imported one is gone with the source that
  produced it

A collision between an interceptor-vended profile and an imported one still
fails closed, which is the behavior that shadowing quietly bypassed for
built-ins.

Operators upgrading should export the profiles their deployment relies on
before upgrading, or copy them from the providers/ directory of the matching
release tag, then import them at the scope their providers use.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* docs(providers): document import-only provider profiles

Rewrite the provider profile documentation around a catalog the operator builds
rather than one the gateway ships.

- gateway-config: the default is [{ type = "user" }], the builtin source type is
  gone and rejected at startup, and an empty catalog is a valid ready state
- profiles: replace the built-in profile table and its shadowing semantics with
  the import workflow, and say plainly that a profile whose binary paths do not
  match the image is inert
- manage-providers: replace the two fixed provider-type tables with
  `provider list-profiles`, and describe catalog-driven command inference
- inference-routing: drop the "still loads built-in profiles for compatibility"
  transition text; its migration walkthrough is now the normal path
- quickstart and the Docker Compose, GitHub, AWS, Google Cloud and Vertex
  tutorials: import the profile before creating the provider, since nothing
  resolves without it
- release notes: a 0.1.0 migration section covering export-before-upgrade,
  importing from the release tag's providers/ directory, what happens to a
  provider whose profile is missing, and the removal of the legacy type aliases
- README, architecture, the openshell-cli and debug-openshell-cluster skills,
  and the governance interceptor example follow the same change; the example's
  smoke assertion now imports a profile to prove an authoritative interceptor
  hides it

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(cli,e2e): distinguish catalog failures and authenticate OIDC seeding

Addresses two findings from review (GATOR-c1bd9867-02 and -03).

The provider profile catalog lookup collapsed its error into an empty
catalog, so an authorization, availability or transport failure on
ListProviderProfiles was indistinguishable from a gateway that genuinely has
no profiles. Command inference then resolved nothing and the sandbox was
created without the provider it needed, deferring the failure to the workload.

Keep the lookup's outcome instead of discarding it. The credential warning is
advisory and still degrades to its generic form, but inference now consults the
catalog only when there is a command to resolve and surfaces the lookup failure
when there is, naming the gateway and pointing at explicit --provider selection.
Two regression tests cover it through a fake gateway whose ListProviderProfiles
returns UNAVAILABLE while every other RPC succeeds: sandbox creation fails
without sending a provider-less create request, and a sandbox with no trailing
command still succeeds because it needs no catalog.

The e2e profile seeding also ran unauthenticated in the OIDC lanes. Those lanes
deliberately skip gateway registration and start the gateway without a TLS
client CA, so no mTLS identity exists and no token has been acquired when the
import runs; the wrapper exited during setup. Skipping the import is not
sufficient because the provider tests now require the claude-code profile.

Add e2e_register_oidc_admin_session, which mints an administrator token with
Keycloak's password grant — the same grant the OIDC test helpers use — and
writes the gateway metadata and token bundle that an interactive login would
have stored, so the import runs as an authenticated administrator. The mTLS
lanes keep the direct import unchanged. Both affected wrappers are covered:
with-podman-gateway.sh had the same defect as with-docker-gateway.sh.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* refactor(cli)!: remove command-derived provider attachment

Addresses GATOR-c1bd9867-01.

A profile's `binaries` list authorizes a binary to reach that profile's
endpoints. It is not a statement that running the binary asks for the provider,
and reading it as attachment intent let a command silently gain provider
authority: the aws-s3 example declares /bin/bash, so `sandbox create -- bash`
resolved an existing aws-s3 provider and attached its credential-backed
capability and network policy to the shell and its descendants. Attachment of
an already-created provider never prompted, so the confirmation flow did not
guard it.

The distinction that would make inference sound — whether a declared binary is
a profile's client or merely a permitted runtime — cannot be expressed:
NetworkBinary carries only a path, and the field that encoded it was removed in
0.1.0. Any substitute is a guess. Restricting the guess to a unique claimant
does not help, because uniqueness measures how sparse the catalog is rather
than what the user intended, and a compiled list of "generic" commands could
never cover an unbounded operator catalog.

Remove trailing-command inference rather than approximate it. A provider is
attached only when named with --provider, which still creates a missing
provider from local discovery when the name matches an imported profile ID.
Sandboxes with no providers remain a normal, fully supported state.

Removing inference also settles GATOR-c1bd9867-02: with no consumer deriving
authority from the catalog, its only remaining use is the advisory credential
warning, so a failed lookup degrades that warning instead of blocking creation.
The regression test now asserts that an unreachable catalog still creates the
sandbox and attaches nothing.

While repurposing the deduplication test, a pre-existing defect surfaced:
repeating a name in --provider auto-created it twice and the second attempt
failed with "provider already exists", because the explicit pass never
consulted the set of names it had already handled. Guard it.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(e2e): load the OIDC token and trust the gateway when seeding profiles

Addresses the carried finding GATOR-c1bd9867-03.

e2e_register_oidc_admin_session established a session the CLI never used. The
CLI decides whether to load a stored bearer token by matching on the auth_mode
field of the gateway metadata alone; the helper omitted that field, so the
metadata fell through to the default arm and oidc_token.json stayed on disk
unread. The profile import went out unauthenticated exactly as it had before
the helper existed. The helper also installed no trust anchor, leaving the CLI
unable to verify the gateway's self-signed serving certificate.

Write auth_mode = "oidc" so the stored token is loaded, and install the CA at
<gateway>/mtls/ca.crt. Only the CA is installed: with no client certificate or
key on disk the CLI falls back to CA-only server verification and authenticates
with the bearer token, which is what these lanes need because they start the
gateway without --tls-client-ca. Certificate verification stays on; no insecure
transport override is introduced.

Assert the session before anything depends on it. ListProviderProfiles is
annotated auth_mode: "bearer", so it cannot succeed unless the token was loaded
and accepted. The helper now fails at that point, naming the gateway config
directory and echoing the CLI output, rather than letting the defect surface
later as an opaque profile import error.

Both affected wrappers pass the PKI directory and CLI binary the helper needs;
the mTLS lanes are untouched.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(e2e): request the OpenShell scopes when minting the admin token

The OIDC lanes failed at the session assertion added in b25dd0298 with
PERMISSION_DENIED and "scope 'provider:read' required". Authentication was
working -- the gateway logged ListProviderProfiles at gRPC status 7, which it
can only reach once the bearer token has been loaded and accepted. The token
simply carried no OpenShell scope.

sandbox:*, provider:*, config:*, workspace:* and openshell:all are optional
client scopes on the openshell-cli client in scripts/keycloak-realm.json, so
Keycloak mints them only when the request asks for them. The password grant
here asked for nothing, leaving the realm defaults (openid, profile, email,
roles, web-origins, acr) and an access token that authorizes no RPC.

Request "openid openshell:all", as e2e/python/oidc/oidc_auth_test.py already
does for its administrator tokens. openshell:all is SCOPE_ALL in
crates/openshell-server/src/auth/authz.rs, so one scope covers the setup calls
without enumerating them. Verified against the realm: the token goes from
"email profile openid" to "email profile openshell:all openid".

Record the same scopes in metadata.json. The helper writes the bundle
openshell gateway login would have stored, and oidc_scopes is the field that
login path reads back, so leaving it out would misdescribe the stored token.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(e2e): seed provider profiles once per VM lane gateway

The VM lanes failed on the second test target with "custom provider profile
'anthropic' already exists" for all fifteen example profiles, and the import
exited non-zero.

e2e_import_example_provider_profiles was called from run_e2e_test, so it ran
once per target -- four times against one long-lived gateway. That was
harmless while import overwrote silently, but profiles are now import-only:
ImportProviderProfiles is create-only and reports an existing id as an
error-severity diagnostic, with no overwrite flag on the request. The first
import therefore succeeds and every later one fails.

Hoist the call to just after the conformance run, which is where the docker,
podman and kube lanes already seed their catalogs. The profiles persist for
the gateway's lifetime, so every target still finds them, and the
E2E_TEST_OVERRIDE path is covered by the same single call.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* test(e2e): name the provider explicitly in the OIDC workspace-user test

user_can_create_sandbox_with_inferred_provider_command reached the gateway
once the admin session was fixed, and then failed: it asserted a
missing-provider error, but the sandbox was created and died provisioning
with "failed to spawn sandbox entrypoint process 'claude-code'".

The test drove provider resolution by passing claude-code as the trailing
command and relying on the CLI to infer the provider type from it. That
inference is what this branch removed, so the trailing word is now nothing
but an entrypoint, and the image has no such binary.

The regression the test guards is not inference itself: it is that a
workspace user resolving a provider is not gated behind Platform Admin.
Name the provider with --provider, the only remaining way to attach one.
The CLI still has to fetch the claude-code profile before it can auto-create
the provider, so the lookup a workspace user must be allowed to make still
happens, and auto-creation still stops at the non-interactive branch with
"missing required provider". Both assertions therefore keep their meaning.

Rename the test and rework its comments to describe what it now exercises;
the old name would otherwise outlive the behavior it was named for.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(providers): reconcile import-only profiles with main

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: Philippe Martin <phmartin@redhat.com>
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-09-17 22:59:27 +00:00
Shiju a72351d370 fix(policy)!: require explicit L7 append targets and scope (#3380)
Make allow and deny appends identify a rule and endpoint and declare every
affected binary and port. Reject incomplete, stale, ambiguous, and provider
targets atomically so a small append cannot silently change a broader scope.

Update CLI previews, wire requests, Go types and generated bindings, SDK
regressions, operator documentation, and live policy-update coverage.

BREAKING CHANGE: AddAllowRules and AddDenyRules require L7RuleTarget instead
of host and port. CLI appends require a rule name and explicit binary scope.

Signed-off-by: Shiju <shiju@nvidia.com>
2026-09-17 22:53:07 +00:00
Shiju 9708ba9999 feat(providers): report applied sandbox provider changes (#3391)
* feat(providers): report applied sandbox provider changes

Record exact provider mutation targets in shared configuration operations.
Require authenticated evidence that credentials, effective policy, and the
workload launch environment have been installed before reporting readiness.

Add bounded CLI and Rust SDK status and wait support, preserving ordinary
revision-scoped references for existing processes. Verify new-client rotation
and acknowledged detach revocation without external-stable resolver changes.

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(providers): align readiness times with protobuf contracts

Represent readiness receipts, status, and operation times with Timestamp
and report intervals with Duration. Reserve the scalar field tags, update
all consumers and generated bindings, and preserve timestamp presence
and nanosecond identity through storage and client validation.

Qualify both empty-map constructors in the Linux boundary test so its
module compiles while retaining the explicit default required by Clippy.

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(cli): preserve provider mutation storage uncertainty

Recognize the gateway's exact structured storage-uncertainty reason for
provider attach, detach, and update. Explain that the change may already
be saved and must be reconciled before retrying, without exposing server
messages or metadata. Preserve uncertainty ahead of generic retry hints.

Exercise saved mutations through the CLI and verify single submission,
redaction, missing receipt handling, and untrusted error-detail rejection.
Document the recovery guidance for users and the public CLI skill.

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(cli): explain denied provider profile lookups

Report exact and alias profile lookup denials with fixed permission and workspace guidance. Keep backend details redacted and stop before provider mutations.

Cover denied create and update calls through the CLI. Verify the complete provider list independently in the cross-workspace OIDC regression, extracting its JSON object from surrounding startup diagnostics.

Signed-off-by: Shiju <shiju@nvidia.com>

---------

Signed-off-by: Shiju <shiju@nvidia.com>
2026-09-17 19:10:43 +00:00
Mrunal Patel a316fd7832 feat(api): add durable workspace mutation admission and replay (#3321)
* feat(api): add durable workspace mutation admission and replay

Part of #3051 (phase 3a). Preserve unresolved admissions, reauthorize replay, and protect same-name replacements.

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* feat(api): extend mutation replay through gateway interceptors (#3323)

Add typed durable replay receipts for 24 ordinary unary mutations, protect sensitive payload fingerprints, and revalidate intercepted retries without repeating post-commit observation.

Part of #3051 (phase 3b).

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(api): scope workspace request IDs by target

Include the requested workspace name in create/delete admission keys while leaving workspace UUID guards unset. Cover cross-target UUID reuse, replay, and missing targets with server and live gateway regressions.

Merge the latest phase-two SDK fixes and preserve the approved interceptor replay changes.

Refs #3051.

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

---------

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
2026-09-17 14:51:56 +00:00
Mrunal Patel c502be9fd7 feat(api): return typed deletion outcomes with explicit missing-target semantics (#3317)
* feat(api): return typed deletion outcomes with explicit missing-target semantics

Implement phase 2 of #3051 across the public gateway API and first-party SDKs. Preserve asynchronous sandbox deletion and observed resource identities, reserve legacy wire fields, and document the coordinated migration.

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(sdk): bind deletion waits to sandbox identity

Track accepted deletions by original sandbox identity in Rust and TypeScript. Preserve name-only waits and cover replacement races and lookup failures.

Merge the phase-one cleanup fix and adapt its interceptor regressions to typed deletion outcomes.

Refs #3051.

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* test(e2e): accept asynchronous stopped sandbox deletion

Allow either completed or accepted deletion output, then continue polling for actual sandbox absence. This preserves the lifecycle assertion under the phase-two deletion contract.

Refs #3051.

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

---------

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
2026-09-16 22:11:19 +00:00
Derek Carr d68b7069c3 refactor(proto)!: use well-known time types (#3113)
* refactor(proto)!: use well-known time types

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(proto): preserve time migration behavior

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(proto): preserve timestamp boundary semantics

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(proto): convert sandbox token expiry to timestamp

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(proto): preserve time compatibility semantics

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(proto): preserve exact endpoint and profile times

Signed-off-by: Derek Carr <decarr@redhat.com>

* test(e2e): use duration for interactive exec timeout

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(sdk-go)!: remove legacy profile duration fields

BREAKING CHANGE: Go provider profile callers must use RefreshBefore, MaxLifetime, and CacheTTL with ProfileDuration instead of the whole-second fields.

Signed-off-by: Derek Carr <decarr@redhat.com>

---------

Signed-off-by: Derek Carr <decarr@redhat.com>
2026-09-16 19:51:45 +00:00
Johnny Greco 2ccef97769 feat(policy): establish one canonical authored policy representation (#3334)
* chore(policy): restart schema implementation

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* feat(policy-schema): add canonical authored policy model

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* refactor(policy): use canonical authored schema

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* refactor(prover): project canonical policy documents

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): describe shared schema boundary

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(policy): close reviewed parser gaps

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(policy-schema): fail closed on unsupported fields

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(policy): preserve partial process identities

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(policy-schema): harden authored policy inspection

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* refactor(policy): rename policy schema crate

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* revert(policy): restore policy schema crate

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* test(e2e): serialize OIDC PKCE scenarios

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

---------

Signed-off-by: Johnny Greco <jogreco@nvidia.com>
2026-09-16 17:36:31 +00:00
John T. MyersandJohn Myers dfd5238d0d fix(gator): make supervised lifecycle sandbox-native (#3343)
* fix(gator): run supervisor as sandbox main process

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* feat(gator): persist supervised state history

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* refactor(gator): remove obsolete background launch mode

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <johntmyers@users.noreply.github.com>
2026-09-15 19:04:13 +00:00
Shiju fd3fd9cf74 feat(sandbox): explain failed calls to external tool servers (#3207)
Show configured tool server addresses and their last observed connection
results together in sandbox status. Keep sandbox lifecycle readiness
separate so an external connection failure does not mark the sandbox unready.

Expose direct endpoint records through the CLI and SDKs, with plain-language
failure explanations and gateway acceptance times. Keep observation tracking,
runtime reporting, and gateway validation in dedicated endpoint status modules.

Preserve bounded reporting, request attribution, retry ordering, and
configuration and supervisor authority checks. Clear obsolete observations
while retaining the configured addresses, and document the distinction
between an observed HTTP response, current availability, and tool success.

Signed-off-by: Shiju <shiju@nvidia.com>
2026-09-15 13:46:38 +00:00
VarshaandJohn Myers 0803c4aa4c refactor(policy)!: remove NetworkBinary harness field (#3222)
* refactor(policy)!: remove NetworkBinary harness field

Closes #3054

Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>

* test(policy): preserve unknown fields during migration

Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>

* fix(policy): preserve advisor provenance during merge

Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>

* fix(policy): remove redundant network binary default

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(policy): record harness schema migration

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-09-11 16:49:14 +00:00
Drew Newberry 33bbda3d33 refactor(persistence): adopt continuation-token pagination (#3249)
* refactor(persistence): adopt continuation-token pagination

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(pagination): address continuation review findings

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(tui): recover completed list refreshes

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(pagination): address review scalability findings

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(pagination): repair branch validation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(go): use page size in template example

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-11 00:02:07 +00:00
Mrunal PatelandJohn Myers 90dbe5454b feat(api): add typed workspace selectors (#3245)
* feat(api)!: add typed workspace selectors

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(cli): preserve template workspace metadata

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* test(e2e): migrate workspace request selectors

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(api): update public schema inventory

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-09-10 20:38:50 +00:00