152 Commits
Author SHA1 Message Date
Prekshi Vyas b031dc037a fix(mxc): reject unsupported live policy updates (#3480)
* fix(mxc): reject unsupported live policy updates (NVBug 6782891)

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* fix(mxc): gate all live policy mutations

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* fix(mxc): gate composed policy mutations

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* fix(ci): satisfy provider update lint

* fix(ci): order provider validation branches

* test(mxc): make policy synchronization deterministic

* fix(server): scope MXC policy synchronization

* fix(server): serialize provider-backed sandbox creation

* test(server): use valid provider create fixture name

---------

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>
2026-09-22 15:44:44 -07:00
Drew Newberry 49b4f0eb7f feat(mxc): add UI policy, credentials, relay lifecycle, and proxy auth
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-17 12:06:06 -07:00
John T. MyersandJohn Myers dfd5238d0d fix(gator): make supervised lifecycle sandbox-native (#3343)
* fix(gator): run supervisor as sandbox main process

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* feat(gator): persist supervised state history

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* refactor(gator): remove obsolete background launch mode

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <johntmyers@users.noreply.github.com>
2026-09-15 19:04:13 +00:00
Shiju fd3fd9cf74 feat(sandbox): explain failed calls to external tool servers (#3207)
Show configured tool server addresses and their last observed connection
results together in sandbox status. Keep sandbox lifecycle readiness
separate so an external connection failure does not mark the sandbox unready.

Expose direct endpoint records through the CLI and SDKs, with plain-language
failure explanations and gateway acceptance times. Keep observation tracking,
runtime reporting, and gateway validation in dedicated endpoint status modules.

Preserve bounded reporting, request attribution, retry ordering, and
configuration and supervisor authority checks. Clear obsolete observations
while retaining the configured addresses, and document the distinction
between an observed HTTP response, current availability, and tool success.

Signed-off-by: Shiju <shiju@nvidia.com>
2026-09-15 13:46:38 +00:00
Drew Newberry 1860010850 feat(sdk): add lazy pagination pagers (#3256)
* feat(sdk): add lazy pagination pagers

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(sdk): harden pager edge cases

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(sdk): cover initial resume token

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(go): fix all-workspaces pager examples

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-11 00:02:08 +00:00
Drew Newberry 33bbda3d33 refactor(persistence): adopt continuation-token pagination (#3249)
* refactor(persistence): adopt continuation-token pagination

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(pagination): address continuation review findings

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(tui): recover completed list refreshes

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(pagination): address review scalability findings

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(pagination): repair branch validation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(go): use page size in template example

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-11 00:02:07 +00:00
John T. MyersandJohn Myers 226a83b323 fix(supervisor): classify credential placeholders in request bodies (#3246)
* fix(supervisor): classify credential placeholders in request bodies

Closes #2904

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* fix(supervisor): preserve same-provider placeholders in request bodies

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <johntmyers@users.noreply.github.com>
2026-09-10 23:10:01 +00:00
Mrunal PatelandJohn Myers 90dbe5454b feat(api): add typed workspace selectors (#3245)
* feat(api)!: add typed workspace selectors

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(cli): preserve template workspace metadata

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* test(e2e): migrate workspace request selectors

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(api): update public schema inventory

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-09-10 20:38:50 +00:00
Shiju ea8eda6d5b feat(supervisor): enforce MCP request protocol versions (#3241)
* feat(supervisor): enforce MCP request protocol versions

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(supervisor): enforce MCP versions across HTTP forwarding

Apply shared request-version guards before authorization and after forward-request rewriting. Require version metadata to survive HTTP header cleanup, and cover valid initialization and selected-revision forwarding through middleware.

Signed-off-by: Shiju <shiju@nvidia.com>

---------

Signed-off-by: Shiju <shiju@nvidia.com>
2026-09-09 20:28:55 +00:00
John T. Myers f4dc6be4b2 refactor(inference): remove managed inference routes (#3195)
* refactor(inference): remove managed inference routes

Closes #3172

Remove the inference route control plane, inference.local data path, built-in router crate, and SDK surface. Move inference workloads to explicitly imported provider profiles and native endpoints, with migration cleanup and updated tests and documentation.

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(policy): preserve alternate upstream isolation

Restore the provider policy activation guard so legacy OpenAI and Anthropic providers configured for alternate base URLs do not grant egress to the built-in public vendor endpoints.

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-09-09 18:47:22 +00:00
Evie Howard 6e6b3c8905 refactor(cli): remove local Dockerfile image builds (#3214)
Signed-off-by: Evie Howard <evhoward@redhat.com>
2026-09-09 13:28:17 +00:00
Philippe Martin 118b250f01 feat(sandbox): support rootfs tar as --from source for VM driver (#2863)
* feat(sandbox): support rootfs tar as --from source for VM driver

Accept flat rootfs tar archives (.tar, .tar.gz, .tgz) via the --from
flag for VM-backed gateways. The CLI detects the archive extension,
validates that the gateway uses the VM compute driver, and passes the
tar path through driver_config. The VM driver copies the tar into its
staging area and feeds it into the existing rootfs extraction and ext4
disk creation pipeline, skipping the container image pull/export steps.

Closes #2175

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(sandbox): validate rootfs tar path at the VM driver boundary

The rootfs_tar_path field in driver_config was passed from the API
caller directly to tokio::fs::copy without validation. An authenticated
user bypassing the CLI could supply arbitrary host paths (e.g.
/dev/zero for disk exhaustion, or readable host files for data
exfiltration).

Introduce a trusted staging directory that the VM driver creates on
startup and advertises via GetCapabilities. The CLI now copies the tar
into the staging directory before creating the sandbox, and the driver
validates that the received path is a regular file inside the staging
root and within a configurable size limit (default 10 GiB) before any
I/O.

New VmDriverConfig options:
- rootfs_tar_staging_dir: override the staging directory
  (default: <state_dir>/rootfs-tar-staging)
- rootfs_tar_max_bytes: override the size limit (default: 10 GiB)

Addresses GATOR-28b5152e-01.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(sandbox): request-scoped staging, size pre-check, and cleanup for rootfs tar

Tighten the rootfs tar staging flow to address the remaining GATOR-01
obligations:

- Request-scoped staging: the CLI creates a unique per-request
  subdirectory (req-<pid>) under the staging root instead of placing
  files directly in the shared directory. The driver enforces that the
  tar path is at depth 2 (staging_root/<subdir>/<file>), preventing
  cross-request path selection.

- Size pre-check: the driver advertises rootfs_tar_max_bytes via
  GetCapabilities. The CLI reads this limit and rejects oversized files
  before copying, avoiding disk exhaustion in the staging directory.

- Cleanup: the driver removes the request staging subdirectory after
  consuming the tar (on cache hit, copy success, or copy failure),
  ensuring staged data does not persist beyond the request.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(vm): restore rootfs-tar sandboxes from persisted image identity

On restore or restart, the one-shot staged tar archive has already been
cleaned up. Reading the persisted image identity from the sandbox state
directory and resolving the cached disk path directly avoids re-accessing
the deleted staging path.

Addresses GATOR-168b9210-01.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(cli): use random staging dirs and enforce byte limit during rootfs tar copy

Replace PID-based request staging directories with tempfile-generated
random names to prevent collisions and make paths unpredictable.
Replace bare tokio::fs::copy with a streaming copy loop that enforces
the advertised max_bytes limit during transfer, closing the TOCTOU gap
between the pre-copy size check and the actual copy.

Signed-off-by: Philippe Martin <phmartin@nvidia.com>
Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(sandbox): issue rootfs tar staging slots from the gateway

A caller could name any host path in `driver_config.vm.rootfs_tar_path`,
which the privileged VM driver then read. The CLI-side locality check did
not apply to direct API requests.

The gateway now owns staging. `BeginRootfsTarStaging` allocates a
request-scoped directory under the driver-advertised staging root and
returns an opaque single-use token; `CreateSandbox` carries the token, and
the gateway substitutes the path it allocated before dispatching to the
driver. `template.driver_config.<driver>.rootfs_tar_path` is rejected
outright in request validation, so a caller-supplied path never reaches
privileged I/O.

Tokens are bound to the issuing workspace and subject, consumed once, and
expire after 30 minutes. Outstanding slots are capped per caller and
overall, so one caller can neither exhaust the staging filesystem nor
starve others. An RAII guard reclaims the directory on every failure path
after consumption, and an age-gated sweep runs at startup and on each
reconcile pass for directories whose driver died before its own cleanup.

The token is stripped from the public sandbox before persistence: the
stored copy is returned verbatim by GetSandbox, ListSandboxes and
WatchSandbox to every member of the workspace.

Also fixes two defects this exposed:

- The CLI wrote `rootfs_tar_path` at the top level of `driver_config`, but
  the gateway forwards only `driver_config.<driver_name>`, silently
  dropping unmatched keys. The archive never reached the VM driver, so the
  documented `--from ./rootfs.tar` flow did not work at all. Config is now
  nested under `vm` and deep-merged, so a caller's existing VM settings
  survive instead of being clobbered by a shallow extend.

- Staging previously required `GetGatewayInfo`, which is restricted to
  `platform_admin`, making the feature unusable for ordinary users on any
  RBAC-enabled gateway. The new RPC matches CreateSandbox at
  `sandbox:write` / `workspace_role: user`.

`compute_driver.proto` is unchanged; the gateway reads the staging root
from the capabilities it already stores.

Refs #2175

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(vm): derive rootfs tar cache identity from archive contents

The prepared-disk cache key combined the archive's full path with an mtime
truncated to seconds, then mapped punctuation to `-`. Distinct paths such as
`/tmp/a/b.tar` and `/tmp/a-b.tar` collapsed onto the same key and reused each
other's disk, a rewrite within the same second kept stale contents, and a long
path could exceed filesystem component limits.

Identity is now a SHA-256 of the archive contents. This is also what makes the
cache work at all now that the gateway allocates a fresh staging directory per
request: a path-derived key would miss on every create.

The archive is hashed, the cache checked, and only on a miss copied — so a hit
skips writing a multi-gigabyte file. The copy is hashed as it is written and
rejected if the digest differs from the first pass, which closes the window
where the source changes during staging rather than approximating it with a
re-stat.

Refs #2175

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(vm): decompress gzip rootfs tar archives during staging

`--from` accepts `.tar.gz` and `.tgz`, but the driver staged whatever bytes
it was given and the guest image-prep VM extracts the staged file with a
plain `tar -xpf`. Compressed sources therefore depended on the guest tar
auto-detecting gzip, and the prepared disk was sized from the compressed
length, which is far too small for the expanded rootfs.

Staging now detects gzip from the archive's magic bytes -- the driver only
ever sees a gateway-issued staging path, never the caller's file name -- and
writes an uncompressed tar. The digest still covers the source bytes, so the
"archive changed while staging" check is unaffected, and expansion is bounded
by `rootfs_tar_max_bytes` so a compression bomb cannot fill the host disk.

`extract_rootfs_archive_to` sniffs gzip as well, so the host-side extraction
path matches.

Adds unit coverage for gzip staging, bounded expansion, and gzip extraction,
plus an e2e sandbox created from a gzip-compressed export.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

---------

Signed-off-by: Philippe Martin <phmartin@redhat.com>
Signed-off-by: Philippe Martin <phmartin@nvidia.com>
2026-09-08 22:36:40 +00:00
Shiju 592df3e014 feat(policy): preserve exact MCP revision allowlists (#3027)
* feat(mcp): add version-aware wire profile metadata

Signed-off-by: Shiju <shiju@nvidia.com>

* feat(policy): canonicalize MCP version allowlists

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(policy): align MCP policy tests with current main

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(policy): canonicalize supervisor protobuf ingress

Materialize defaultable MCP revisions before ambiguity checks, OPA construction, and sidecar delivery. Reject invalid sidecar policies with bounded errors.

Signed-off-by: Shiju <shiju@nvidia.com>

---------

Signed-off-by: Shiju <shiju@nvidia.com>
2026-09-05 04:24:49 +00:00
John T. MyersandJohn Myers fc0929749c fix(policy): harden advisor transport proposals (#3136)
Signed-off-by: John Myers <johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <johntmyers@users.noreply.github.com>
2026-09-04 19:38:00 +00:00
Akram Ben Aissi 08eac8c46d fix(sandbox): detect an available login shell instead of hardcoding /bin/bash (#3147)
* fix(sandbox): detect an available login shell instead of hardcoding /bin/bash

The built-in default sandbox command and the interactive SSH session
hardcoded /bin/bash. Minimal images such as Alpine ship only /bin/sh
(BusyBox ash), so sandbox startup failed with an opaque "No such file or
directory (os error 2)" that never named the missing binary.

Add openshell-core::shell with shell-path constants and a runtime
detect_login_shell() that resolves a shell present in the sandbox image
($SHELL if executable, then bash, then /bin/sh). Use it for:

- the built-in default command (only the default is remapped; explicit
  user commands are never rewritten), resolved in the supervisor so it
  inspects the sandbox filesystem rather than the gateway's
- the SSH interactive shell
- the SHELL environment variable

Also name the program in the spawn error so a missing shell/binary is
diagnosable instead of a bare ENOENT.

Refs #3146

Signed-off-by: Akram <akram.benaissi@gmail.com>

* fix(sandbox): drop $SHELL preference in shell detection

$SHELL is image/user-controlled and the detected shell is later invoked
with `-lc`, so an executable that is not a compatible shell (e.g.
SHELL=/bin/false) would pass the executable check and then break command
execution even when /bin/sh is available. Resolve only from known shell
paths instead.

Also add a USR_BASH constant for /usr/bin/bash rather than a string
literal in SHELL_CANDIDATES.

Refs #3146

Signed-off-by: Akram <akram.benaissi@gmail.com>

* fix(sandbox): resolve the default login shell in the supervisor (empty command = default)

Addresses review: interactive PTY SSH now uses the detected shell, the shell
tests are portable across the Windows lane, and default-shell provenance is
carried without a new spec field.

An omitted command is left empty end to end and resolved in the supervisor,
which is the only place that sees the sandbox image:
- The CLI forwards the command as-is; the gateway persists an omitted command
  as empty (no baked /bin/bash -l) and requests a TTY.
- MainProcessConfig carries the command empty (the transport now allows it);
  the supervisor resolves a login shell that exists in the sandbox image (bash
  when present, otherwise /bin/sh on minimal images like Alpine) and logs the
  resolved shell.
- Interactive PTY SSH (spawn_pty_shell) uses the detected shell; a shared
  build_ssh_shell_command helper covers the PTY and non-PTY paths, with a
  deterministic sh-only regression test.
- Unix-only shell tests are gated with cfg(unix).

An explicit command is always run verbatim.

Refs #3146

Signed-off-by: Akram <akram.benaissi@gmail.com>

---------

Signed-off-by: Akram <akram.benaissi@gmail.com>
2026-09-03 22:25:29 +00:00
Mrunal PatelandEvan Lezar 03003cd017 fix(cli): require ANSI-capable terminal before colorizing (#3121)
* fix(cli): require ANSI-capable terminal before colorizing

Follow-up to #3026, raised in review.

`auto` treated any terminal as styleable, so `TERM=dumb openshell ...`
still emitted escapes into a terminal that renders them literally. An
unset TERM had the same problem.

This is partly a regression that #3026 introduced. `console`, which
drives indicatif and dialoguer, already refused to colorize when TERM is
`dumb` or unset, and miette applies the same check through
supports-color. #3026 overrides both with its own switch, so it replaced
two working checks rather than only failing to add one. tracing and the
owo-colors wrapper never had detection, so those two are a gap rather
than a regression.

Add the capability check to the `auto` branch only, matching console's
unix rule: `dumb` is not capable, and an unset TERM is not capable
because nothing identifies a capable terminal. Empty is treated as unset,
which diverges from console — it reads `TERM=""` as capable since the
value is not `dumb` — because an empty value names no terminal type and
every other variable here already treats empty as unset.

Because the check sits after the explicit branches, `--color always` and
FORCE_COLOR still force styling on a dumb terminal, and `--color never`
and NO_COLOR still suppress it on a capable one. TERM is a unix signal;
Windows consoles enable virtual terminal processing and do not set it, so
the check does not apply there.

The existing pty test now pins TERM. It previously inherited the ambient
value, which would make its outcome depend on the environment now that
capability is consulted — CI runners frequently leave TERM unset.

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* refactor(cli): combine stream and terminal capability checks

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* docs(cli): clarify table color behavior

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(cli): cover redirected status table colors

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>
Co-authored-by: Evan Lezar <elezar@nvidia.com>
2026-09-02 15:03:10 +00:00
grs 74960ebfae feat(server): add sandbox templates (#2833)
* feat(server): add sandbox workload templates

Signed-off-by: Gordon Sim <gsim@redhat.com>

* feat(go-sdk): add sandbox workload template support

Signed-off-by: Gordon Sim <gsim@redhat.com>

* feat(rust-sdk): add sandbox workload template support

Signed-off-by: Gordon Sim <gsim@redhat.com>

* feat(python-sdk): add sandbox workload template support

Signed-off-by: Gordon Sim <gsim@redhat.com>

* feat(typescript-sdk): add sandbox workload template support

Signed-off-by: Gordon Sim <gsim@redhat.com>

* docs(agents): document sandbox workload templates

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(cli): support default GPU requests in sandbox templates

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(server): cap sandbox templates per workspace

Signed-off-by: Gordon Sim <gsim@redhat.com>

* docs(architecture): document sandbox workload template boundaries

Signed-off-by: Gordon Sim <gsim@redhat.com>

* feat(cli+sdk): expose sandbox workload template provenance

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(cli): include sandbox template annotations in output

Signed-off-by: Gordon Sim <gsim@redhat.com>

* feat(sandbox): add label selectors to template listing

Signed-off-by: Gordon Sim <gsim@redhat.com>

* test(e2e): cover sandbox template failure paths

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(server): preserve command and ttl when creating sandbox from template

Signed-off-by: Gordon Sim <gsim@redhat.com>

* test(server): add field coverage test for template merge

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(go-sdk): add pagination support to fake client

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(cli): warn on env vars that looks like secrets

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(sdk-ts): propagate sandbox workspace through lifecycle calls

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(docs): update workspace management docs

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(ts-sdk): support command and tty when creating from template

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(go-sdk): support command and tty when creating from template

Signed-off-by: Gordon Sim <gsim@redhat.com>

* test(python-sdk): verify command and tty handling when creating from template

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(go-sdk): update docs and ClientInterface

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(server): validate sandbox create specs before I/O

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(go-sdk): guard empty DNS-1123 label validation

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(python-sdk): allow empty template builder mappings

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(cli): align template GPU JSON default output

Signed-off-by: Gordon Sim <gsim@redhat.com>

---------

Signed-off-by: Gordon Sim <gsim@redhat.com>
2026-09-02 00:55:29 +00:00
grs 7b64c5c88e fix(cli): continue multi-item deletes after failures (#3111)
Signed-off-by: Gordon Sim <gsim@redhat.com>
2026-09-01 23:01:32 +00:00
Mrunal Patel b4afcd8a43 fix(cli): suppress ANSI color when stdout is not a terminal (#3026)
* fix(cli): suppress ANSI color when stdout is not a terminal

The CLI colorized output unconditionally. owo-colors is built without
its `supports-colors` feature, so `.green()` and friends emitted escape
sequences regardless of destination, and nothing in the CLI read
NO_COLOR. Piping any command through grep or awk matched against bytes
the caller could not see; `forward list` was the case that surfaced it,
where an escape sits immediately before the STATUS word and defeats a
pattern anchored on whitespace.

Add a `color` module holding a process-wide switch resolved once in
run_async, before any output. Command modules import its `Colorize`
trait in place of `OwoColorize`; the method names match, so the ~450
call sites are unchanged, but each consults the switch when it renders
and delegates to owo-colors so the escape bytes stay identical. The two
traits collide by design: importing both in one module is an ambiguity
error, which keeps unconditional coloring from returning.

owo-colors is not the only styled path, and the rest each carry their
own default, so the switch governs them too:

  - tracing_subscriber formats with ANSI on, does no terminal detection,
    and writes to stdout, so `openshell -v ... | ...` leaked escapes the
    same way the tables did. It now takes the setting via with_ansi.
  - indicatif and dialoguer both style through console, which has its
    own detection but cannot learn about --color. Overriding console's
    global switch covers every progress bar and prompt rather than the
    specific ones constructed today. Both the stdout and stderr switches
    are set, since prompts and progress bars draw to stderr.
  - miette renders errors through its own handler, likewise unaware of
    --color, so init installs one built from the setting.

Resolution order: `--color always|never`, then NO_COLOR, then
CLICOLOR_FORCE, then whether stdout is a terminal. The decision is made
against stdout even for stderr text, since stdout is what gets parsed;
`--color always` restores styling when redirecting.

Padding is unaffected — the format spec is forwarded to the inner
Display, so widths measure text rather than text plus escapes.

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(cli): resolve color per output stream

Review feedback on #3026.

Resolving one answer from stdout and handing it to every library meant a
redirected stream inherited the other stream's terminal check. Running
`openshell ... 2> build.log` from a terminal wrote escapes into the log,
because console's stderr switch and miette's handler were both given
stdout's answer. That is worse than the behavior before this branch,
where both libraries did their own per-stream detection.

Resolve `auto` separately for stdout and stderr and hand each library
the answer for the stream it writes to: tracing and console's stdout
switch get stdout, miette and console's stderr switch get stderr. The
owo-colors wrapper is the exception, since its call sites are split
across println! and eprintln! and a Painted value cannot tell which
macro will consume it; it styles only when both streams accept escapes,
erring toward plain text rather than risking a redirected stream.

Existing tests could not catch this: Command::output gives both streams
pipes, so a per-stream decision and a single stdout-derived one look
identical. Add a test that puts stdout on a pty and stderr on a pipe,
which fails when stderr is handed stdout's answer.

Replace CLICOLOR_FORCE with FORCE_COLOR. The clicolors spec does not say
how to treat `0`, and implementations that special-case it disagree with
force-color.org, which keys on presence and non-emptiness only. Using
FORCE_COLOR gives it the same rule as NO_COLOR: set and non-empty means
yes, whatever the value. Nothing depended on CLICOLOR_FORCE, which was
introduced earlier on this branch and never released.

Carry the whole style in an owo_colors::Style rather than dispatching a
local enum through a six-arm match, and merge styles when chaining so
`x.green().bold()` emits one `\x1b[32;1m...\x1b[0m` instead of nesting
two wrappers. No call site styles already-styled text, so merging is
safe; the emitted bytes are shorter and there is a single reset.

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

---------

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
2026-09-01 21:11:31 +00:00
John T. MyersandJohn Myers 07df822090 feat(providers): make profiles authoritative (#2962)
* feat(providers): make profiles authoritative

Closes #1988

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* docs(providers): move profiles into provider navigation

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* docs(providers): clarify provider attachment lifecycle

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* fix(tui): scroll provider profile picker

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* fix(providers): honor profile credential semantics

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* fix(providers): prefer exact profile IDs

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* fix(providers): harden authoritative profile adoption

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* test(oidc): align provider fixtures with profiles

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* fix(providers): preserve authoritative profile lifecycle

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <johntmyers@users.noreply.github.com>
2026-09-01 19:22:07 +00:00
Mrunal Patel d5742e01a3 feat(cli): add structured output to list commands (#3067)
Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
2026-09-01 16:16:01 +00:00
John T. Myers 8ffc6c2a13 fix(policy): compose advisor proposals with provider endpoints (#2935)
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-09-01 01:04:13 +00:00
Artem Lytvyn eb15e1a4c9 feat(sandbox): add --no-login-shell to skip shell startup files on exec (#2852)
* feat(sandbox): add --no-login-shell to skip shell startup files on exec

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* chore(sdk/go): regenerate proto bindings for no_login_shell

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* test(supervisor-process): cover login-shell flag selection

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* feat(sandbox): gate --no-login-shell on supervisor SSH banner

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

---------

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>
2026-08-31 16:17:58 +00:00
Drew Newberry 69a05ebb3b fix(sandbox): complete successful main processes (#2884) 2026-08-28 18:21:20 -07:00
Vyncint Ng 0e79653a7f feat(tui): show persisted guidance on rejected policy chunks (#2908)
* feat(tui): show persisted guidance on rejected policy chunks

The reviewer's rejection note is stored in PolicyChunk.rejection_reason and
already reaches the TUI in GetDraftPolicyResponse.chunks, but openshell-tui
never read the field. The note was dropped at the last step, so a reviewer
had no way to recall why a chunk had been rejected.

Render it in two places, following the truncate-in-list / full-in-popup
convention in the TUI development guide: a truncated, dimmed suffix on the
list row, and a "Guidance:" line in the detail popup.

Gate the accessor on status == "rejected" rather than on the field alone.
Approving a chunk passes None for the reason and the gateway writes the
field only when Some, so a chunk that was rejected and later approved still
carries the old note. Reading the field unconditionally would surface a
stale rejection on an approved rule.

Part of #1098, Definition of Done item "Rejected chunks show persisted
guidance". The other items on that issue are untouched.

Signed-off-by: Vyncint Ng <115854244+vyncint@users.noreply.github.com>

* fix(tui): scroll the draft detail popup so long guidance stays reachable

A rejection reason has no server-side length cap, so a paragraph-length one
overflowed the fixed 22-row detail popup: the tail was clipped and the later
fields and both action hints were pushed off screen. The same overflow already
affected the unbounded rationale and security_notes fields, so fix the popup
rather than special-case the guidance line.

Split the popup's inner area into a scrolling body and a pinned hint row,
following the pattern in ui/create_provider.rs, and drive it with j/k, the
arrow keys, PageUp, PageDown, g and G. The approve and close controls now stay
on screen at every scroll position, and the bottom border carries the scroll
position.

Wrap free-form values explicitly instead of relying on Paragraph's Wrap, so the
rendered row count is exactly lines.len() and the scroll clamp cannot under-run
the content. Paragraph::line_count would answer the same question but sits
behind ratatui's unstable-rendered-line-info feature.

Add deterministic TestBackend coverage at 80x24 with a 2,000-character reason,
covering the head and tail, the pinned hints, over-scroll clamping, and the
pre-existing long-rationale case. Document the reviewer-facing guidance in the
policy advisor page.

Part of #1098.

Signed-off-by: Vyncint Ng <115854244+vyncint@users.noreply.github.com>

* fix(tui): wrap the denial row and hint footer on a narrow popup

Making the scroll clamp exact meant dropping Paragraph's Wrap, which also
removed wrapping from the lines that do not go through push_wrapped. At 70
columns the denial row clipped mid-value and lost the last-seen timestamp, and
at 60 a pending chunk with scrollable content pushed [Esc] Close past the right
edge of the single hint row. The key still worked, but that is the same
controls-not-visible failure one axis over.

Keep the denial row on one line while it fits and wrap it onto the label indent
when it does not, so the common 80-column layout is unchanged. Pack the footer
hints into as many rows as they need without splitting a hint, and derive the
footer height from that: adding a hint row shrinks the body and can itself
change whether the content scrolls, so the two settle together.

The earlier tests were all 80x24, which is why they missed this. Cover the
denial timestamps and the approve, reject and close hints at 60, 70 and 80
columns, plus the packing helper directly.

Part of #1098.

Signed-off-by: Vyncint Ng <115854244+vyncint@users.noreply.github.com>

* fix(tui): measure popup wrapping in display columns, not chars

wrap_value packed rows by chars().count(), and a CJK glyph is one char but two
terminal columns. A double-width rejection reason therefore produced rows about
twice the popup's width: the right half of every row was clipped, and unlike
the vertical case there is no horizontal scroll to recover it.

Measure width with Span::width(), the same measurement ratatui applies when it
lays cells out, so the wrap agrees with the renderer by construction and no new
dependency is needed. This covers word packing, the hard-break path for a word
wider than the line, the label indent, the footer hint packing, and the denial
row's fits-on-one-line check.

The list row's shortened copy had the same defect through truncate_str, which
counts chars. Leave truncate_str alone for its two existing callers and add
truncate_display for the guidance row.

Cover it with an 80x24 render regression that interleaves markers through the
CJK text: a trailing marker lands on its own short row in both the broken and
fixed layouts, so it would not detect this. The two unit tests measure with an
independent column oracle rather than the helper under test, which would make
them tautological.

Part of #1098.

Signed-off-by: Vyncint Ng <115854244+vyncint@users.noreply.github.com>

---------

Signed-off-by: Vyncint Ng <115854244+vyncint@users.noreply.github.com>
2026-08-25 18:37:36 +00:00
Mrunal Patel 18ce13b9b1 feat(providers): expose actionable OAuth refresh failures (#2887)
* fix(providers): classify OAuth refresh failures

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* test(providers): add Keycloak refresh e2e lane

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(providers): harden OAuth refresh recovery

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(providers): classify post-mint refresh failures

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(providers): tolerate malformed OAuth subtypes

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

---------

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
2026-08-25 18:00:00 +00:00
Seth Jennings 5206bc51b2 feat(sdk): add OAuth Client Credentials support to SDKs (#2907)
* feat(sdk): add renewable client credentials auth

Implement lazy OAuth client-credentials acquisition and renewal for the Python, TypeScript, and Go SDK clients, with shared security conformance coverage and service-account documentation.\n\nCloses #2803

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(sdk): address client credentials review

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(sdk): address client credentials review

Signed-off-by: Seth Jennings <sjenning@redhat.com>

---------

Signed-off-by: Seth Jennings <sjenning@redhat.com>
2026-08-25 17:53:47 +00:00
grs 40d1b48666 feat(provider): support for SPIFFE backed token exchange (#1970)
* feat(provider): add ability to request token exchange instead of client credentials as OAuth grant_type

Signed-off-by: Gordon Sim <gsim@redhat.com>

* test(proxy): add further tests for token exchange

Signed-off-by: Gordon Sim <gsim@redhat.com>

* test(provider): add runnable example for token exchange

Signed-off-by: Gordon Sim <gsim@redhat.com>

* test(e2e): cover Podman token exchange grants

Signed-off-by: Gordon Sim <gsim@redhat.com>

* refactor(oauth): extract duplicated functionality from server and supervisor

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(provider): evict nearest-to-expiry entry from intermediate token cache

Signed-off-by: Gordon Sim <gsim@redhat.com>

* doc(supervisor): add podman example for token exchange

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(provider): withhold token-exchange subject credentials

Signed-off-by: Gordon Sim <gsim@redhat.com>

---------

Signed-off-by: Gordon Sim <gsim@redhat.com>
2026-08-24 05:42:30 +00:00
John T. MyersandJohn Myers 679fe4c334 fix(policy): validate the applicable advisor candidate (#2850)
* fix(policy): bind reviews to applicable candidates

Build and validate the exact effective-policy candidate before approval, bind review to live policy/provider/credential inputs, and preserve inspected endpoint contracts during mechanistic expansion.

Closes #2821

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* fix(policy): canonicalize advisor review inputs

Serialize nested protobuf maps in stable key order for proposal review tokens and effective-policy hashes. Narrow reused multi-port endpoint contracts to the denied port so advisor proposals cannot widen binary access. Add regressions for both cases.

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* test(e2e): keep advisor sandbox running

Create the issue 2821 regression sandbox detached with a durable canonical main process so policy denial, approval, and hot-reload checks run before lifecycle exit.

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* fix(policy): apply reviewed draft batches atomically

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <johntmyers@users.noreply.github.com>
2026-08-21 19:05:53 +00:00
Drew Newberry ef296806f5 feat(sandbox): add canonical main process (#2726)
* feat(sandbox): add canonical main process

Closes #2710

Persist and supervise one canonical workload per sandbox, attach sandbox connect to its retained session, and make every unexpected main-process exit terminal.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(sandbox): simplify canonical main process contract

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): preserve legacy VM main compatibility

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): preserve main status across driver updates

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): satisfy macOS process lint

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): gate Linux exit acknowledgement publisher

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(sandbox): simplify main process plumbing

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(supervisor): make controlling tty ioctl portable

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(supervisor): initialize canonical process environment

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* perf(supervisor): optimize retained main session

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(sandbox): detach main session on ctrl-c

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): use explicit main detach keys

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(dev): atomically stage Docker supervisor

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(test): align Docker main environment assertion

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(sdk): expose canonical main process fields

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-08-20 22:18:23 +00:00
John T. Myers b2ea81822b feat(network): enable Docker and Podman policy DNS and transparent TCP (#2723)
* feat(network): enable Docker transparent TCP egress

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(e2e): cover Docker transparent TCP egress

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* feat(network): correlate transparent TCP audit events

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(examples): add transparent TCP Redis demo

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(examples): demonstrate blocked TCP connections

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(examples): focus Redis demo audit output

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): close transparent TCP policy bypasses

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(sandbox): reject unsupported TCP policy reloads

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(ci): satisfy Linux transparent TCP lints

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* feat(podman): enable transparent TCP egress

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(podman): permit policy DNS port binding

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(e2e): use qualified transparent TCP hostname

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(podman): preserve exact policy DNS names

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(podman): route policy DNS over TCP

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(dns): serve multiple TCP queries per connection

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(network): explain native DNS and TCP egress

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(sandbox): reconcile runtime reload with upstream

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): harden transparent DNS capture

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(podman): preserve resolver behavior for native tcp

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(network): clarify native tcp runtime constraints

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): remove unused transparent tcp pin

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): admit redirected transparent tcp

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): restore podman transparent networking

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(podman): permit alpine busybox binaries

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(podman): use portable alpine keepalive

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(podman): build musl networking fixture

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(podman): isolate musl DNS probe

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(podman): keep privileged port capability dropped

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): preserve transparent TCP port 53

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): report synthetic pool pressure by family

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(podman): bind tcp fixtures before readiness

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(podman): grant fixture low-port bind

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-08-20 19:26:53 +00:00
John T. Myers 4d7f402ce2 feat(policy): establish direct TCP egress foundation (#2711)
* feat(policy): accept explicit tcp endpoint protocol

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* refactor(network): snapshot authoritative egress decisions

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(policy): document explicit tcp protocol

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(policy): defer transparent TCP release guidance

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* chore(go): regenerate sandbox protobuf bindings

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): complete tcp egress foundation

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(policy): document explicit tcp contract

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): fail closed on authorization errors

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(podman): fence delayed exit events before restart

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(policy): validate network endpoint destinations

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(providers): opt in tcp credential fixture

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(policy): require dns host for transparent tcp

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-08-20 16:43:09 +00:00
Mrunal Patel 6e90f3d5a0 feat(providers): store refresh credentials in credential drivers (#2801)
* feat(providers): store refresh credentials in credential drivers

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(providers): harden refresh credential lifecycle

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(providers): migrate legacy refresh secrets before skip

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* refactor(providers): defer credential migration

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(providers): make refresh configuration atomic

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

---------

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
2026-08-19 21:37:24 +00:00
Drew Newberry 998db04780 feat(policy): allow non-root sandbox identities (#2785)
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-08-19 21:20:14 +00:00
alangou 0d708d6d51 fix(policy): gate uninspected credentialed endpoints (#2493)
* fix(policy): gate uninspected credentialed endpoints

Signed-off-by: Adrien Langou <alangou@nvidia.com>

* refactor(cli): extract allowed-ip option parsing

Signed-off-by: Adrien Langou <alangou@nvidia.com>

* fix(policy): gate endpointless credential bindings

Signed-off-by: Adrien Langou <alangou@nvidia.com>

---------

Signed-off-by: Adrien Langou <alangou@nvidia.com>
2026-08-19 15:18:02 +00:00
Mrunal Patel 8d67250a5d fix(providers): keep refresh credential handles stable (#2780)
* fix(providers): keep refresh credential handles stable

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(providers): protect refresh-owned credentials

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

---------

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
2026-08-18 20:45:36 +00:00
Piotr Mlocek 44bf0df485 feat(middleware): inspect WebSocket text messages (#2477)
* feat(middleware): inspect websocket text messages

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): address websocket review feedback

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(network): bound websocket message assembly

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(network): harden websocket upgrade lifecycle

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* refactor(middleware): unify in-process and remote transports

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* feat(middleware): support regex websocket redaction

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): bound persistent streaming sessions

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): accept websocket sequence gaps

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* refactor(middleware): refine websocket introspection contract

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): clarify websocket preflight lifecycle

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): clarify websocket coverage semantics

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(network): type websocket frame failures

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(network): return 503 when middleware admission is exhausted

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): align streaming API contract

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): clarify WebSocket event result scope

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(rfc): simplify middleware revision history

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* feat(examples): add WebSocket content guard support

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): unify binding payload limits

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* refactor(middleware): align payload limit terminology

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): address websocket review feedback

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(network): address websocket review findings

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* test(network): allow Linux handler setup in preflight regression

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(network): harden websocket relay finalization

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(network): inspect compressed websocket messages

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* test(network): stabilize compressed websocket regressions

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(go-sdk): regenerate middleware protobuf binding

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): clarify websocket skip lifecycle

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): address WebSocket review feedback

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

---------

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-08-14 21:51:42 +00:00
Seth Jennings 0f8fad23c4 feat(sandbox): add stop and start operations (#2653)
* feat(sandbox): add suspend and resume operations

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(server): preserve lifecycle work after cancellation

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(server): reconcile ambiguous lifecycle outcomes

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(server): complete suspended session cleanup

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(vm): preserve suspension state on resume failure

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(server): retry retained lifecycle transitions

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(server): clean sessions after suspend reconciliation

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* test(sandbox): cover deleting suspended sandbox

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(kubernetes): preserve progressing sandbox suspension

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(kubernetes): bound suspend status polling

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(kubernetes): detect legacy sandbox suspension

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(tui): render suspended sandbox phases

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* refactor(sandbox): rename suspend and resume lifecycle

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* perf(server): clean stopped sessions on transition

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(kubernetes): fail fast on rejected stop

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(compute): fence stale restart lifecycle events

Signed-off-by: Seth Jennings <sjenning@redhat.com>

---------

Signed-off-by: Seth Jennings <sjenning@redhat.com>
2026-08-13 06:13:54 +00:00
Artem Lytvyn dd2b4e3bc0 feat(cli): warn when --env values look like credentials (#2655)
* feat(cli): add credential env match validation

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* feat(cli): warn when --env values look like credentials

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* docs(sandbox): add flag --no-credential-warnings details + polishing

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* fix(cli): match credential keywords on underscore segments

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

---------

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>
2026-08-11 21:40:55 +00:00
John T. MyersandJohn Myers 0120535efc feat(proxy): bind static credentials to provider endpoints (#2510)
* feat(proxy): bind static credentials to provider endpoints

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* test(e2e): verify static credential endpoint isolation

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* docs(provider): explain static credential endpoint binding

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* fix(e2e): use valid endpoint isolation fixtures

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* docs(provider): explain static credential endpoint binding

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* fix(credentials): preserve binding identity across rotations

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(proxy): enforce bindings across request lifecycle

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(proxy): close credential relay gaps

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(credentials): clarify binding failure behavior

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(credentials): hash selected provider profile scope

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(proxy): resolve credentials after request admission

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(credentials): clarify binding failure diagnostics

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(proxy): align single-route credential denials

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(credentials): harden endpoint-bound rotation

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(credentials): enforce identity and authority binding

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(credentials): snapshot provider environment atomically

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(e2e): include authority port in query proxy requests

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(credentials): close credential revocation gaps

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(proxy): explain authority mismatch diagnostics

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(credentials): enforce binding lifecycle invariants

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(provider): reject credential config collisions

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): capture credential scope atomically

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): distinguish origin and absolute targets

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(provider): isolate endpointless profile credentials

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): normalize IPv6 request authorities

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(credentials): clarify endpointless profile isolation

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* feat(policy): bind endpointless provider credentials

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(credentials): use current GCP placeholder revision

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(providers): explain policy credential bindings

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(credentials): cover endpointless fail-closed invariant

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(policy): expect ambiguity rejection at creation

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(server): authenticate rebased policy requests

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* refactor(proxy): share credential mismatch finding builder

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(credentials): cover malformed binding metadata

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(credentials): verify multi-key endpoint isolation

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(e2e): cover same-host credential path denial

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(credentials): document serialized refresh contract

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* refactor(proxy): consolidate L7 log formatting

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* perf(credentials): precompile endpoint binding patterns

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* perf(credentials): share identity epoch revisions

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(proxy): require explicit request default ports

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(policy): validate SigV4 credential sources

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(credentials): preserve endpoint bindings for credential handles

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* feat(go-sdk): expose network credential bindings

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <johntmyers@users.noreply.github.com>
2026-08-10 19:19:45 +00:00
Shiju 5e2f0d1b37 fix(policy): prevent implicit authorization inheritance (#2499)
* fix(policy): prevent implicit authorization inheritance

A network rule authorizes every listed binary to reach every listed
endpoint, so unioning an AddRule operation's binaries and endpoints
independently grants binary-by-endpoint pairs the operation never
declared. Require AddRule to declare the complete product before
merging, and reject an operation that would give one host and port two
different MCP inspection contracts. The rejection names the binaries the
operation still has to declare.

Fix proposal coverage on the same surface. An any-binary proposal was
vacuously covered by a binary-restricted loaded rule, and a complete
product split across several loaded rules was reported as uncovered.
Coverage compares merge-widened endpoint fields by containment and
exact-matches only the fields the merge never widens, so a policy the
gateway just merged always reads back as covered and the sandbox
policy.local /wait long-poll cannot spin to its deadline.

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(policy): close authorization-inheritance gaps in merge and coverage

Coverage treated an unset proposal value for a field the merge retains as a
request for the default, so a proposal that merged cleanly into an endpoint
carrying enforcement, protocol, or tls read back as uncovered and left the
policy.local /wait long poll spinning. Unset now means unspecified.

Ports are a set on the wire but each port is an independent authorization.
Coverage and the inheritance check both resolve one binary, host, path, and
port at a time, so ports spread across loaded rules resolve and a complete
declaration split across incoming endpoints is accepted.

An incoming empty binary list means any binary. It now has to declare every
merged endpoint like a new concrete path does, and once declared the promotion
is applied instead of appending an empty list and leaving the restricted scope
in place.

The endpoint-overlap fallback folded a new rule name into an existing rule
where inheritance validation then rejected it, leaving no way to grant a binary
part of a rule. Folding now keeps the requested rule name when it would widen,
and reports that it did. An MCP contract conflict still propagates because one
host and port carry a single inspection contract.

AddAllowRules and AddDenyRules select an endpoint by host and port alone and
now reject a target that resolves more than once, including two paths on one
rule. RemoveBinary rejects an any-binary rule rather than reporting a success
that leaves the binary authorized.

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(policy): gate port and MCP-contract widening independently of binary scope

The changed-port check only rejected when the operation also left an existing
binary undeclared. An operation listing every existing binary could therefore
declare one port of a multi-port endpoint and have its widened fields land on
the endpoint the merge shares across all of them, authorizing L7 rules on a
port it never named. Each changed port must now be named by the operation on
its own, independently of binary-scope coverage.

MCP contract compatibility was checked inside the endpoint fold, which only
compares endpoints agreeing on host, path, and a shared port. The sandbox
resolves one extended configuration per host and port and never consults the
path, so a second MCP endpoint under another path, or in another rule, left the
effective strict-tool-name, method-profile, and body-limit contract decided by
match order. Contract agreement is now enforced across the whole merged policy,
including provider-composed rules. A conflict already present in the baseline
is left alone so unrelated updates still apply.

An empty binary list authorizes any binary. Appending an incoming named list
made it non-empty and revoked every process the operation did not name, turning
an additive update into a silent mass revocation. An already-empty scope is now
kept and reported; only a restricted scope is replaced by an incoming
any-binary scope.

Warnings raised during a fold now name the rule that was actually modified
rather than the rule name the operation requested, which differ when the
endpoint-overlap fallback redirects the operation.

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(policy): route undeclared-port conflicts through the separate-rule fallback

The fold-only classifier decides which merge errors disappear when the incoming
authorization stays on its own rule. UndeclaredPortWouldChange was added ahead
of the existing-binary conflict but never classified, so a differently named
narrow update against a multi-port endpoint failed outright instead of landing
separately. A same-key update still returns the error, because there the
operation chose the target.

The classifier is now an exhaustive match rather than a matches! with an
implicit false. A new variant defaulting to "not fold-only" is what withdrew the
separate-rule remedy here, so adding one has to be an explicit decision.

Inspection-contract agreement now covers protocol, not only MCP options. The
sandbox resolves one extended configuration per host and port and never
consults the path, so an MCP endpoint and a REST endpoint on the same host and
port left the effective inspection protocol decided by match order. Endpoints
with no protocol carry no contract and are skipped.

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(policy): compare only MCP contracts when detecting endpoint conflicts

The post-merge conflict scan was broadened to compare inspection protocol as
well as MCP options on one host and port. The supervisor selects among matching
endpoint configs by most-specific path, so a broad REST endpoint and a narrower
GraphQL endpoint on the same host and port are unambiguous and supported. The
broader comparison rejected those updates even though the equivalent full policy
loads and serves correctly.

MCP options are not selected that way, so the scan keeps comparing them: two MCP
endpoints on one host and port still have to agree on strict tool names, method
profile, and body limit, whatever paths or rules hold them.

The policy page returns to describing the MCP-specific rule and the path-aware
selection it sits alongside.

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(policy): keep MCP off a host and port shared with other inspection

Narrowing the conflict scan back to MCP options let an MCP endpoint sit under a
path already covered by a broader REST endpoint. The supervisor picks the parser
by most-specific path, so the MCP endpoint parses the request, but
_policy_allows_l7 is existential over every endpoint matching it. A plain REST
rule on the overlapping path can therefore make allow_request true for a
JSON-RPC tool call the MCP endpoint never allowed, and the relay forwards it.

The scan now records every inspected protocol on a host and port. MCP may not
share one with a differently inspected endpoint, and two MCP endpoints there
still have to agree on one contract. Endpoints that are not inspected carry no
contract and never compete, and two non-MCP endpoints stay supported because
they share one method-and-path rule vocabulary.

Signed-off-by: Shiju <shiju@nvidia.com>

---------

Signed-off-by: Shiju <shiju@nvidia.com>
2026-08-09 23:54:04 +00:00
5548405fcb feat(credentials): add provider credential storage drivers (#2437)
* feat(credentials): add provider credential storage drivers

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(credentials): harden credential update handling

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(credentials): harden credential driver security, correctness, and performance

Address review findings from the credential storage drivers PR:

- Route additional_credentials through the driver on refresh to prevent
  silent data loss for multi-credential providers (e.g. AWS STS)
- Clean up stored credential handles on CAS failure during refresh to
  prevent orphaned secrets in external backends
- Enforce namespace validation in the Kubernetes Secrets driver to
  prevent cross-namespace credential access when allow_reference_namespace
  is not enabled
- Cache Vault Kubernetes auth tokens with 80% TTL to avoid re-authenticating
  on every credential operation
- Parallelize resolve_credentials in all three drivers using try_join_all
  for faster sandbox startup
- Add existingSecret support for the KEK Secret to fix helm template/GitOps
  workflows where lookup returns empty and regenerates the key
- Document RBAC blast radius for the Kubernetes Secrets credential driver
  and recommend a dedicated namespace

Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com>
Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>

* fix(credentials): add optimistic concurrency, fix thundering herd, parallelize operations

Use resourceVersion optimistic concurrency with retry loop for K8s
Secret ownership checks to prevent TOCTOU races. Switch Vault token
cache from RwLock to Mutex with double-check pattern to prevent
thundering herd on cache miss. Parallelize credential store and delete
operations across independent keys using try_join_all.

Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com>
Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>

* fix(credentials): handle partial failures, add delete retry, consolidate cleanup

Replace try_join_all with join_all in credential store/delete operations
to handle partial failures — successfully-stored handles are cleaned up
when another key fails. Add retry loop with conflict detection to
db-credstore delete_credential, matching the K8s driver pattern.
Consolidate 4 manual cleanup_pre_stored_provider_credentials call sites
into a single error handler using an async block. Remove inconsistent
.trim() from db-credstore validate_handle_owner.

Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com>
Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>

* fix(credentials): fix retry loop guard and remove unprotected validation

Remove attempt-count guard from 409/Aborted match arms in retry loops
so the post-loop Status::aborted error is reachable after exhausting
retries. Previously, last-attempt conflicts fell through to the
catch-all error arm, producing misleading Status::unavailable errors.

Remove duplicate validation calls that ran after
prepare_provider_credential_update but outside the cleanup-protected
async block, which would leak pre-stored handles on failure.

Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com>
Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>

* fix(credentials): add workspace/provider UUID to credential backend paths

Include workspace and provider ID in credential backend object paths to ensure
cross-workspace uniqueness and prevent credential collision (GATOR-1806c9be-01).

- Updated credential driver proto to include workspace and provider_id fields
- Modified Vault driver to include workspace/provider_id in managed_secret_path
- Modified Kubernetes Secrets driver to include workspace/provider_id in
  credential_owner_id and managed_secret_name
- Updated all credential runtime calls to pass workspace/provider_id
- Updated tests to use the new signatures

This prevents two workspaces sharing the same external credential store from
colliding on provider names, which was a critical security issue (CWE-639).

* fix(credentials): preserve provider-level expiration for handle-backed credentials

Compute effective expiration from both provider and driver values using the
earliest non-zero timestamp and skip expired values before insertion
(GATOR-1806c9be-02).

- Modified resolve_provider_handles to check provider credential_expires_at_ms
- Skip expired credentials during resolution instead of returning them
- Use effective expiration (min of provider and driver) in resolution results
- Fix inference.rs to preserve earliest expiration when merging

This ensures handle-backed credentials respect the same expiration semantics
as inline credentials.

* fix(credentials): stage refresh changes under new handles before validation

Stage credential replacements under new immutable handles instead of reusing
existing handles to prevent overwriting committed values before validation/CAS
(GATOR-1806c9be-03).

- Stage credentials with empty existing_handles map to force new handle creation
- Validate and CAS before the new values are committed to backend storage
- Delete old handles only after successful CAS
- On CAS failure, delete only the newly staged handles
- This prevents CWE-362/CWE-367 race conditions where failed refreshes could
  still modify or delete the active credential

The fix ensures that a rejected refresh cannot modify the backend object still
referenced by the committed provider record.

* fix(credentials): add timeouts to credential driver RPCs

Apply configured timeouts to both startup capability negotiation and runtime
RPCs to prevent indefinite hangs (GATOR-1806c9be-05).

- Add DEFAULT_CREDENTIAL_DRIVER_RPC_TIMEOUT_SECS constant (30s)
- Apply timeout to GetCapabilities during startup connection
- Apply timeout to all runtime RPCs (store, delete, resolve)
- Use tokio::time::timeout to bound the entire GetCapabilities operation
  during startup, not just the socket connection
- Return contextual deadline errors on timeout

This prevents a faulty or overloaded driver from hanging gateway operations
indefinitely.

* fix(credentials): fix test to use consistent workspace/provider identity

The Kubernetes auth Vault resolve test was constructing a managed path with
test-workspace/test-provider-id but sending default/prov-123 in the request,
causing validation to reject the request (GATOR-18e32351-01).

- Update test to use test-workspace and test-provider-id in the request to
  match the logical_path construction
- This ensures the test exercises the intended code path and validates
  Kubernetes auth resolution properly

The test now passes and correctly validates identity enforcement.

* fix(credentials): use unique staging ID for refresh to avoid overwrites

Stage refresh replacements under genuinely distinct immutable handles using
a unique staging ID to prevent overwriting committed values (GATOR-1806c9be-03).

- Generate a unique staging ID using UUID for each refresh operation
- Use this staging ID when storing credentials instead of the real provider ID
- Pass the same staging ID during cleanup on failure to delete only staged objects
- This ensures deterministic paths (Vault) and object names (K8s) don't collide
  with the committed provider's credentials

The fix prevents failed refreshes from silently replacing active credentials
or breaking providers by deleting still-referenced backend objects.

* fix(credentials): wrap credential driver RPCs in local timeouts

Add local tokio::time::timeout wrappers around credential driver RPCs to
bound non-compliant or stalled UDS peers (GATOR-1806c9be-05).

- Wrap StoreCredential, DeleteCredential, and ResolveCredentials in local timeouts
- Return contextual deadline_exceeded errors when timeouts occur
- Keep existing gRPC timeout metadata for compliant implementations
- GetCapabilities during startup was already wrapped in previous commit

This ensures a faulty local driver cannot hang gateway operations indefinitely,
even if it accepts the connection but never responds to the RPC.

* fix(credentials): preserve ownership for staged refreshes

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(credentials): bound startup capability probe

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* test(provider): authenticate credential handler requests

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(ci): grant actions read to credential driver e2e

Signed-off-by: Seth Jennings <sjenning@redhat.com>

---------

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com>
Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>
Signed-off-by: Seth Jennings <sjenning@redhat.com>
Co-authored-by: Taylor Mutch <taylormutch@gmail.com>
Co-authored-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>
2026-08-05 16:29:41 +00:00
Matthew Grossman 537805568d feat(sandbox): honor OCI image working directories (#2530)
* feat(sandbox): honor Docker OCI working directories

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* fix(sandbox): honor effective workspace access

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* test(sandbox): cover enforced workspace denial

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* docs(docker): explain effective workdir checks

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* fix(sandbox): validate effective workspace writes

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* fix(sandbox): reserve supervisor control roots

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* refactor(sandbox): centralize control paths

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* fix(sandbox): reserve OCI runtime mount roots

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

---------

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>
2026-08-04 17:38:51 +00:00
John T. MyersandJohn Myers 905b554c7c refactor(network): consolidate proxy egress pipeline (#2373)
* refactor(network): introduce shared egress pipeline

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(network): cover shared proxy egress paths

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* refactor(network): make destination authorization explicit

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* refactor(network): pin proxy relay policy context

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(network): lock relay generation contracts

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(network): establish phase zero compatibility baseline

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* feat(policy): detect ambiguous network endpoints

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* feat(sandbox): fail closed on invalid policy updates

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* refactor(network): invalidate relays on policy changes

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(policy): document validation failure posture

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(network): cover validation and middleware egress

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(config): move policy failure mode to gateway toml

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(network): name proxy contracts by behavior

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): align overlap validation with endpoint selection

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(network): expect hard loopback denial

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(network): match declared endpoint denial

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(policy): preserve path-specific endpoint overrides

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(network): respect hard-blocked host gateways

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): reconcile proxy refactor with main

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): preserve CONNECT policy generation

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(policy): cover runtime endpoint glob semantics

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(server): reject ambiguous policies before persistence

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(policy): explain ambiguity preflight behavior

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(policy): compare body limits within protocol

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(proxy): avoid global tracing capture race

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* chore(server): format rebased provider tests

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(server): authenticate rebased policy requests

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(sandbox): retain runtime on middleware outage

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(server): preflight provider composition activation

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(sandbox): distinguish runtime failure transitions

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
Signed-off-by: John Myers <johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <johntmyers@users.noreply.github.com>
2026-07-31 20:59:00 +00:00
Derek Carr 9c019a93f5 Wire authorization into workspace model (#2445)
* feat(auth): implement RFC 0011 Phase 2 workspace authorization

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(auth): address PR review feedback on workspace authorization

- Docker e2e: add --health-port and switch readiness probe from
  `openshell status` to `curl /healthz`, fixing a false-positive
  readiness check in OIDC mode where the CLI exited 0 without
  actually contacting the gateway

- ListWorkspaces: move membership filtering from post-query N+1
  lookups into a SQL EXISTS subquery so pagination applies to the
  visible set, not the global ordering. Add generic
  list_with_membership to the persistence layer.

- Descriptor validator: reject role/scope fields on unauthenticated
  and sandbox auth modes, and allow-list workspace_role as
  user/admin and global_role as platform_admin to catch typos at
  startup

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(server): use authed request in delete telemetry test

The workspace authorization added by the Phase 2 auth changes requires
a Principal on every delete request. The delete-telemetry test was still
using a bare Request::new, so extract_principal failed before the handler
could acquire the delete gate, causing a 5-second timeout flake.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(auth): address gator review findings for workspace authorization

- Inject unauthenticated-local-dev principal in no-auth gateway mode so
  handlers that call extract_principal() always find one.
- Cap label-selector membership query at MAX_PAGE_SIZE instead of
  u32::MAX to bound the in-memory read.
- Authorize workspace membership before resolving workspace existence in
  all sandbox RPCs to prevent workspace-name enumeration by non-members.
- Remove dead_code allow on AuthorizedWorkspace.workspace now that
  callers use the normalized name from the authz result.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(auth): close workspace-name oracle and label-selector truncation

Swap authorize-before-resolve ordering in 27 handlers across
provider.rs, service.rs, policy.rs, and workspace.rs to prevent
CWE-203 workspace-name enumeration by non-members.

Add combined membership+label SQL query (list_with_membership_and_selector)
to both persistence backends so ListWorkspaces with label selectors no
longer silently drops results beyond the first page of membership matches.

Signed-off-by: Derek Carr <decarr@redhat.com>

* test(auth): add non-member rejection and membership+label persistence tests

Add comprehensive test coverage for workspace authorization changes:
- Non-member rejection tests across all 44 workspace-scoped handlers
  (sandbox, provider, service, policy, workspace, inference) verifying
  PERMISSION_DENIED is returned instead of NOT_FOUND to prevent
  CWE-203 workspace-name oracle
- Persistence test for list_with_membership_and_selector verifying
  SQL-level membership EXISTS + label filtering, multiple predicates,
  no-match cases, and pagination

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(auth): format merged import line in sandbox tests

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(auth): address gator re-review findings on workspace authorization

- Fix TUI unconditionally setting providers_v2_enabled after provider
  refresh; read the actual gateway setting via GetGatewayConfig at
  startup instead
- Fix SQLite json_extract with dotted label keys (e.g. example.com/env)
  by quoting the key in the JSON path
- Add authed_request wrappers to upstream OCI identity tests that were
  missing a principal after rebase
- Add test proving GetGatewayConfig is accessible without Platform Admin
- Add test for dotted/prefixed Kubernetes-style label key filtering

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(auth): address second gator re-review findings

- Loosen GetGatewayConfig from platform_admin to scope-only so workspace
  users can discover providers_v2_enabled during sandbox creation with
  inferred-provider commands; update proto descriptor, descriptor
  validation, and RFC 0011 access table
- Add validate_label_selector to handle_list_workspaces and escape
  single quotes in SQLite json_extract interpolation (CWE-89
  defense-in-depth)
- Re-fetch providers_v2_enabled after TUI gateway switch so the new
  gateway's capability is reflected
- Add e2e test for workspace user with inferred-provider command
- Add persistence test for adversarial label keys with SQL injection
  attempts
- Add handler test for invalid label selector rejection in
  ListWorkspaces

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(auth): address third gator review findings

- Cap label selector pairs at 64 (CWE-400) to bound SQLite dynamic SQL
- Add SCOPE_ONLY_METHODS allowlist for scope-without-role RPCs (CWE-863)
- Normalize ID-based data-plane handlers to return NOT_FOUND for
  unauthorized sandboxes, closing the cross-workspace oracle (CWE-203)
- Fix TUI provider profile cache lookup key mismatch for legacy
  providers with empty profile_workspace
- Add whoami to CLI skill reference command tree
- Update TUI skill doc with workspace, provider, and settings coverage
- Document scope/workspace orthogonality on GetGatewayConfig proto

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(auth): extend CWE-203 normalization to policy.rs sandbox handlers

GetSandboxConfig and GetSandboxLogs in policy.rs had the same
fetch-before-authorize pattern that leaked cross-workspace sandbox
existence. Promote fetch_and_authorize_sandbox to pub(super) and
use it from both sandbox.rs and policy.rs handlers.

Signed-off-by: Derek Carr <decarr@redhat.com>

* test(auth): update assertions for CWE-203 sandbox ID normalization

Cross-workspace sandbox access via ID-based handlers now returns
NOT_FOUND instead of PERMISSION_DENIED to prevent existence inference.
Update the unit test and OIDC e2e assertion to match.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(auth): narrow CWE-203 error mapping and correct whoami output formats

Only remap PERMISSION_DENIED to NOT_FOUND in fetch_and_authorize_sandbox
and RevokeSshSession, letting INTERNAL and UNAUTHENTICATED propagate
as-is. Fix whoami --output format values in cli-reference.md to match
the actual CLI (table/json/yaml, not text/json).

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(ci): share network namespace with Keycloak in containerized CI

In GitHub Actions job containers, Docker port publishing lands on the
host, not inside the job container. Detect this environment and attach
Keycloak to the job container's network namespace instead, with
hardened defaults (cap-drop ALL, no-new-privileges, loopback-only
listener).

Signed-off-by: Derek Carr <decarr@redhat.com>

---------

Signed-off-by: Derek Carr <decarr@redhat.com>
2026-07-30 00:31:32 +00:00
Evan LezarandDrew Newberry 662dee68eb refactor(compute): make sandbox readiness gateway-owned across all drivers (#2153)
* refactor(compute): extract create_sandbox_record and update_sandbox_record helpers

Split apply_sandbox_update_locked into named helpers to make the two
distinct paths explicit: create_sandbox_record for first-observation events
and update_sandbox_record for subsequent driver snapshots on existing sandboxes.
The dispatcher now uses a match on the existing record rather than an
early-return guard.

No behavior change.

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* refactor(compute): make sandbox readiness gateway-owned across all drivers

Introduce compute_phase_components and apply_readiness_conditions to
centralise the gateway's phase composition logic. The public SandboxPhase
is now determined by combining the backend phase reported by the driver with
supervisor session presence, independent of the driver implementation.

Remove SupervisorReadiness from the driver contract. Running containers
always report BackendReady; the gateway owns the Ready decision. Rename
the dispatcher match to three arms so that status-less events for existing
sandboxes are a documented no-op rather than a silent pass-through.

Drop backend_ready_no_session and the SupervisorNotConnected condition.
The BackendReady driver condition plus the Provisioning phase already
communicates that the backend is up but the supervisor has not connected.
The redundant condition added noise without new information.

Closes #1951

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(compute): expose disconnected supervisor readiness

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(sandboxes): clarify supervisor readiness lifecycle

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Evan Lezar <elezar@nvidia.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
2026-07-29 10:10:53 +00:00
Matthew Grossman bc14018cad feat(sandbox): use policy-first OCI image identity (#2509)
* feat(sandbox): use policy-first OCI image identity

Closes #2331

Preserve per-field policy omission, derive Docker and Podman fallbacks from the inspected immutable image, and resolve the final numeric identity before starting agent children.

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* fix(sandbox): preserve declared process identities

Keep explicit policy values and OCI-declared names intact, defer passwd lookup until a primary GID is required, and refresh stale policy examples.

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* fix(supervisor): reuse resolved OCI identity

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* fix(supervisor): allow Linux pre-exec arguments

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* fix(kubernetes): protect resolved sandbox identity

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* fix(sandbox): prepare workspace for OCI identity

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* refactor(sandbox): own only workspace root

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* fix(sandbox): harden partial identity drops

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* test(sandbox): scope OCI image e2e to Docker

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* fix(sandbox): narrow OCI identity fallback scope

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* test(podman): cover OCI identity launch

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* fix(podman): exercise OCI fallback in E2E

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

---------

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>
2026-07-29 05:27:21 +00:00
Grace Smith deced8716b refactor(policy): extract shared L7 endpoint validation (#2389)
Move L7 endpoint semantic checks into the openshell-policy crate so both
profile lint and the runtime validator share one implementation. This
eliminates drift between the two validation paths.

The shared validator covers 9 checks: unknown protocol, rules/access
mutual exclusivity, JSON-RPC family access rejection, json-rpc requires
rules, non-JSON-RPC protocol requires rules or access, MCP requires
rules when allow_all is false, rules-would-deny-all detection,
deny_rules require protocol, and deny_rules require base allow set.

Changes rules/deny_rules fields to Option<Vec<...>> so absent vs empty
is distinguishable at lint time. Adds is_effectively_empty() to
L7AllowProfile for deny-all detection of allow: {} objects. Makes
rules_would_deny_all MCP-aware by checking tool/params.name selectors
before classifying a rule as deny-all. Adds params field to
L7AllowProfile so MCP tool selectors survive proto round-trip.

Signed-off-by: Grace Smith <gsmith@redhat.com>
Signed-off-by: Grace Smith <grasmith@redhat.com>
2026-07-24 17:48:24 +00:00
Roland HussandEvan Lezar b422b6783a feat(cli): add --output json/yaml to sandbox get, status, and sandbox create (#1989)
* feat(cli): add --output json/yaml to sandbox get, status, and sandbox create

Add structured output support (JSON and YAML) to sandbox lifecycle
commands: sandbox get, status, and sandbox create.

Introduce ProgressOutput enum to cleanly separate interactive,
plain, and silent display modes during sandbox provisioning.

Signed-off-by: Roland Huß <rhuss@redhat.com>

* fix(cli): reject structured create output with side effects

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(cli): suppress upload ssh stdout

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Roland Huß <rhuss@redhat.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>
Co-authored-by: Evan Lezar <elezar@nvidia.com>
2026-07-24 16:10:41 +00:00
Drew Newberry 59f7839f6b fix(auth): report gateway authentication status (#2435)
* fix(auth): report gateway authentication status

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(auth): reuse gateway info for status probe

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-07-23 20:01:28 +00:00