Commit Graph
52 Commits
Author SHA1 Message Date
pkhodade-NV ab64e84bfc fix(core): enforce owner-only Windows ACLs on sensitive files and dirs (#3495)
* fix(core): enforce owner-only Windows ACLs on sensitive files and dirs

set_dir_owner_only/set_file_owner_only were unconditional no-ops on
Windows, so the CLI's mTLS client private key, OIDC/edge tokens, cached
SSH keys, and the gateway's key-encryption key relied entirely on
inherited NTFS ACLs with no OpenShell-applied restriction. Apply an
owner-only DACL via SetEntriesInAclW/SetNamedSecurityInfoW with
PROTECTED_DACL_SECURITY_INFORMATION to strip inherited ACEs, matching
the 0700/0600 guarantee already provided on Unix. is_file_permissions_too_open
now also works on Windows instead of being Unix-only, closing the
detection gap alongside the prevention gap.

Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
(cherry picked from commit 71560e947f85819efbcddf70ddda94befab62b0b)

* fix(core): treat a NULL DACL as too open in is_file_permissions_too_open

has_foreign_trustee conflated a NULL DACL with an unreadable/invalid
ACL and returned Some(false) (not too open) for both. Per the Win32
contract, a NULL DACL means the object grants full access to everyone
-- the most permissive state possible -- so it must be flagged as too
open. Split the null and invalid-ACL branches: null now returns
Some(true), invalid ACL keeps the existing unreadable-ACL fallback
(None, which the caller maps to false via unwrap_or). Adds a
regression test that constructs a real NULL DACL via a
SetNamedSecurityInfoW helper confined to the windows_acl module,
consistent with the existing unsafe-FFI confinement in that module.

Found by CodeRabbit review on MR !113.

Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
(cherry picked from commit 46e635a4ef1d6937cdb088f46aa85baee3d6ad28)

* fix(core): close three false-negative gaps in the Windows ACL audit

restrict_to_current_user() updated only the DACL, leaving a foreign
owner's implicit WRITE_DAC right intact -- they could later replace
the DACL we just set. Query OWNER_SECURITY_INFORMATION and take
ownership in the same SetNamedSecurityInfoW call; if the caller can't
(a genuinely foreign-owned object), the call now fails instead of
silently leaving the object insecure.

is_file_permissions_too_open() mapped every Win32 inspection failure
(missing READ_CONTROL, an invalid ACL, a token-query failure) to
"not too open" via unwrap_or(false). Fail closed instead: an
inspection failure is a security false-negative risk, not a green
light.

has_foreign_trustee()'s ACE loop only recognized plain
ACCESS_ALLOWED_ACE_TYPE and treated every other type as non-granting.
Windows also defines access-allowed object, callback, and
callback-object ACE variants that can grant rights to a foreign
trustee; this audit doesn't parse their wider layouts, so their mere
presence is now conservatively flagged as too open instead of
silently skipped.

Also updates architecture/gateway.md, which still described the
SQLite file-tightening behavior only in terms of Unix mode 0o600, to
distinguish it from the owner-only DACL behavior on Windows.

Addresses review comments on PR #3495.

Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>

* fix(core): conditional owner claim and audit owner in Windows ACL helpers

restrict_to_current_user: query the current owner before calling
SetNamedSecurityInfoW. Include OWNER_SECURITY_INFORMATION only when the
path has a foreign owner -- requesting it unconditionally fails with
ACCESS_DENIED (0x80070005) on standard credentials even when the current
user is already the owner, because WRITE_OWNER is not implied by object
ownership. A foreign-owned path still triggers an ownership claim and
fails hard if the claim is denied, preserving the security contract.

has_foreign_trustee: request OWNER_SECURITY_INFORMATION alongside
DACL_SECURITY_INFORMATION and reject paths with a foreign owner
immediately, before inspecting the DACL. A foreign owner has implicit
WRITE_DAC rights and can replace any DACL we set, so a clean DACL is not
sufficient evidence of safety on a foreign-owned object.

architecture/gateway.md: clarify that the Windows path-hardening behavior
sets mode 0o600 on Unix and applies a protected owner-only DACL on
Windows, with conditional ownership claim and fail-hard semantics for
foreign-owned objects.

Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>

---------

Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
2026-09-23 10:13:32 -07:00
Prekshi Vyas 2d5c9bfaab fix(network): scope Windows egress by socket owner (#3491)
* fix(ocsf): attribute MXC proxy events to sandboxes

NVBug 6783086

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* fix(network): scope Windows egress by socket owner

Resolve each accepted Windows proxy connection to its unique owning PID and executable before evaluating binary-scoped network policy. Fail closed when ownership or process identity cannot be established, and cover allowed and undeclared child processes with a real MXC regression.

* fix(network): preserve sandbox context with socket owners

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>
(cherry picked from commit 32ea316deb6c83c157fdf2de4a2cdc3a7d2da1e3)

* fix(network): harden Windows process identity

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

---------

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>
2026-09-22 20:22:40 -07:00
Prekshi Vyas b2d6280794 fix(windows): repair OpenClaw ProcessContainer startup (NVBug 6782898) (#3475) 2026-09-21 16:17:48 -07:00
Mrunal Patel 39cf4823f7 feat(api): add structured gateway errors and SDK decoding (#3313)
* feat(api): expose structured gateway errors across SDKs

Refs #3051. Add standard validation, conflict, and retry details; preserve raw transport status in Rust, Go, TypeScript, and Python; document status and recovery guidance.

This is the structured-error foundation only. Mutation result shapes, allow_missing, durable request deduplication, and exec retry semantics remain follow-up work.

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(python): preserve wrapped RPC cleanup handling

Inspect the original gRPC call when handling missing sandboxes during deletion waits and managed cleanup. Add intercepted cleanup regressions and clarify the error-wrapper migration contract.

Addresses the cleanup review on #3313; part of #3051.

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

---------

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
2026-09-15 17:51:11 +00:00
ae57979b03 feat(mxc): add Windows ETW-to-OCSF audit trail (#3015)
* feat(mxc): ETW->OCSF audit consumer + Windows OCSF JSONL parity (cp6 P1)

Add a Windows MXC ETW->OCSF audit trail in openshell-driver-mxc: a real-time
Sandboxing-provider ETW consumer that decodes events (TDH), attributes each to
an OpenShell sandbox_id, and maps them to OCSF (lifecycle 6002, config 5019,
process 1007, finding 2004).

cp6 Phase 1 - durable OCSF JSONL audit-file parity with Linux:
- openshell-ocsf: add emit_ocsf_event_routed (populates the event-bridge
  thread-local AND stamps sandbox_id+message in one dispatch) plus public
  set/clear_current_event; OS-aware device (Device::windows/for_current_os) so
  device.os.name reflects the host instead of a hardcoded Linux stub.
- etw_consumer: emit via the routed emit (previously fired a bare info! that
  never populated the bridge, so the structured event was dropped).
- openshell-server: install OcsfJsonlLayer over a synchronous daily-rotated
  appender (durable under force-kill), gated by OPENSHELL_OCSF_JSON, path via
  %PROGRAMDATA%\OpenShell\logs (override OPENSHELL_OCSF_LOG_DIR).
- device.hostname now resolves to the real gateway machine name.

Box-proven on 7F203-MXC-001: JSONL lines == shorthand OCSF rows, all valid
OCSF JSON, per-sandbox attribution intact, disabled state writes nothing.

Signed-off-by: Akber Raza <akberr@nvidia.com>

* feat(mxc): map remaining Sandboxing ETW events to OCSF

Close the last three ETW->OCSF gaps so the audit trail covers the full
set of events the Sandboxing provider emits (12/12):

- ProcessLaunched -> Process Activity [1007] "Launch" (confirmed start;
  carries the real processId/threadId, the twin of CreateProcessInSandbox
  which only has the request + command line).
- SandboxProxyConfigured -> Device Config State Change [5019] (the one
  network-plane setup event; surfaces proxyPort, "no proxy" when 0).
- SandboxConsoleReferencePlumbed -> Device Config State Change [5019]
  (console-handle plumbing).

map_config_state now handles the full config/hardening/setup family and
carries proxyPort/hasConsoleReference/creationFlags as unmapped fields.
Verified on 7F203-MXC-001: 11/12 event types emit OCSF without a proxy
(SandboxProxyConfigured requires proxy config to fire).

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc): seed ETW attribution under registry lock + Device tests

Address CodeRabbit review on !31:

- Prevent stale ETW attribution on a delete/launch race: register the
  wxc-exec pid while holding the registry lock, and bail if the sandbox
  entry is already gone. Previously the attribution key could be seeded
  after `delete` had removed the sandbox, leaving a stale key that could
  misroute later Sandboxing ETW events to a dead sandbox_id. Lock order
  (registry -> attribution) matches the delete path, so no deadlock.
- Add unit tests for the new Device::windows and Device::for_current_os
  constructors to harden Windows/Linux OCSF device parity.

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc-etw): buffer+replay racing events and harden attribution keys

Addresses two ETW->OCSF attribution review items (Shailendra #1, #2).

#2 early-event loss: ETW delivers the sandbox create/config burst the instant wxc-exec starts, which can beat the driver's register_launch (now under the registry lock post-Ready). process_event previously dropped anything unresolved, losing the racing burst. Add a bounded, time-bounded pending buffer (PENDING_MAX=4096, PENDING_TTL=5s): unresolved events are held and replayed once attribution lands, aged-out ones dropped. Consumer switched to a timed recv_timeout(200ms) so the buffer is re-driven after each event and on a tick. Emit path factored into shared emit_resolved().

#1 attribution collisions: a Windows PID is recycled after exit and a command line is commonly identical across sandboxes. register_launch now rebinds by_pid on reuse and clears the stale last_pid_sid hint (warns if the PID still pointed at a different, leaked sandbox); command line is held in by_cmd only while unique and demoted to a new ambiguous_cmds set on a second owner, so a duplicate command refuses to resolve rather than misroute.

Unit tests: buffer replay (direct + cross-link), buffer bound, PID-reuse rebind, duplicate-cmd non-resolution. Box-verified on 7F203-MXC-001 (5 sandboxes, identical cmd -> 5 isolated sandbox_ids, 50/50 OCSF/JSONL, BuffersLost=0).

Signed-off-by: Akber Raza <akberr@nvidia.com>

* docs(mxc-etw): note cmd_line is captured raw with no privacy filtering

Review item #3 (Shailendra): add a PRIVACY NOTE on map_process_launch stating cmd_line is copied verbatim into OCSF process.cmd_line with no redaction, so secrets/PII on a command line land unredacted in the durable audit trail (deliberate audit-fidelity trade-off; treat the log as sensitive). Redaction is owned by an upstream privacy layer, not this path; no general audit-output PII scrubber exists today (openshell_core::secrets [CREDENTIAL] redaction is scoped to the proxy HTTP-target logging, a separate egress path).

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc-etw): open ETW trace on caller thread so start_session reports real status

Review item #4 (Shailendra): start_session previously returned Ok(EtwSession) as soon as the pump thread was spawned, but OpenTraceW ran later inside that thread; if it failed we still handed back a live-looking session and logged 'consumer started' (silent failure = false audit coverage).

Split the two Win32 calls instead of adding a channel handshake (avoids any lost-wakeup/hang risk): the quick, synchronous OpenTraceW now runs on the caller thread (open_trace), and only the blocking ProcessTrace runs on the pump thread (run_trace). start_session returns Err if OpenTraceW fails (reclaiming the boxed Sender so the consumer disconnects, stopping the session, joining the consumer) and returns Ok/logs 'started' only once capture is genuinely open. Opened handle + LoggerName buffer + boxed Sender are carried to the pump via a Send OpenedTrace so they outlive ProcessTrace.

Box-verified on 7F203-MXC-001: consumer started=True, failed-to-start=False, 50 OCSF rows / 50 JSONL, BuffersLost=0 (no regression to capture/emit).

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc-etw): guard pending-event replay against PID recycling

CodeRabbit flagged that drain_resolved() re-resolved buffered events
against the live by_pid map, so if Windows recycled a wxc-exec PID within
PENDING_TTL a stale event from the dead sandbox could be emitted under the
new owner.

Stamp each by_pid registration with its Instant and add resolve_replay(),
used only on the buffered/replay path. It (a) never falls back to the
recycle-/ambiguity-prone by_cmd or last_pid_sid keys, and (b) trusts a PID
match only when the registration is not newer than the buffered event by
more than REPLAY_PID_GRACE (2s) - a recycled PID's registration lands well
outside that window, so the stale event ages out instead of misattributing.
The legitimate #2 seed race (registration lands ~immediately) still replays.

Adds unit tests for the recycle-refusal, in-grace acceptance, and
weak-fallback exclusion.

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc-etw): surface unexpected ProcessTrace termination (review #4)

start_session already returns Err on OpenTraceW failure (runs on the
caller thread since e41a7701), closing the first half of Shailendra's #4.
This closes the second half: ProcessTrace's result was discarded, so if
capture died mid-run the backend had no way to know.

Add a shared CaptureHealth (stopped/stopping/exit_code) between the pump
thread and EtwSession. run_trace now records ProcessTrace's WIN32_ERROR
and, when the pump returns without a deliberate stop, logs at ERROR that
MXC OCSF capture is no longer running. EtwSession::stop() sets `stopping`
before teardown so a normal shutdown isn't misreported, and
EtwSession::is_capture_alive() exposes the state for status/diagnostics.

Box-verified on 7F203-MXC-001: 5 sandboxes, 50 attributed OCSF rows,
JSONL parity 50/50, BuffersLost=0, clean start/stop (no false failure).

Signed-off-by: Akber Raza <akberr@nvidia.com>

* feat(mxc-ocsf): add ETW->OCSF audit-trail example kit; fix proxy-configured message

Add a runnable OCSF audit-trail example under examples/ (run-ocsf-audit.ps1,
mxc-ocsf-audit.toml, ocsf-audit.yaml, README) that spins up sandboxes with the
in-process ETW consumer and egress proxy on, emitting a full OCSF JSONL audit
trail across all four classes (6002/5019/1007/2004).

Fix SandboxProxyConfigured mapping to log "MXC sandbox proxy configured" instead
of a misleading "(no proxy)" when the provider reports proxyPort=0; the event's
presence already indicates proxy configuration. Verified on-box: 26 events, all
mapped ETW event types present.

Signed-off-by: Akber Raza <akberr@nvidia.com>

* feat(mxc-ocsf): clearer audit report + client-safe run-ocsf-audit.ps1

Improve the ETW to OCSF audit-trail example output and make it safe to ship.

Report:
- Add an event-type coverage count ("N of M expected event types fired");
  the denominator auto-adjusts (8 with proxy on, 7 with -NoProxy).
- Split the checklist into expected event types vs anomaly findings
  (ActivityError/FallbackError), which are reported separately and not
  counted toward coverage (a clean run may emit none).
- Verdict is now coverage-based (all expected types must fire) instead of
  the looser "at least 3 OCSF classes".
- Call out the absolute path to the durable OCSF JSONL log prominently.

Client-safety:
- Default -ShareOut to empty (no auto-copy); pass -ShareOut a UNC path to
  opt in. Removes a hardcoded internal share path from a published example.
- Drop internal-team wording ("Hand that zip back for evaluation", "BUNDLE:")
  in favor of neutral "Results bundle:".
- Update README-ocsf-audit.txt to match the opt-in -ShareOut behavior.

Verified on both MXC boxes: 7F203-MXC-001 (base-container) -> PASS, 8 of 8
event types, 26 OCSF events across 4 classes; 7F203-MXC-003 (AppContainer
fallback) -> reduced set as expected, clean output.

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc): configure OCSF audit workloads per sandbox

- remove unsupported gateway-scoped workload fields from the shipped MXC audit example.
- build the command, working directory, and filesystem grant from each run's ShareDir
- pass the workload through --driver-config-json.
- preserve the host CONNECT proxy configuration and conditional audit coverage for the future host_connect_proxy merge
- require the workload output when determining the audit verdict.

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc): omit command arguments from OCSF audit logs

- record only the executable basename for MXC CreateProcessInSandbox audit events
- leave process.cmd_line unset so workload arguments cannot reach shorthand or JSONL logs
- cover tokens, passwords, signed URLs, and PII with a secret-leak regression test
- update the audit example, architecture guidance, and published logging documentation
- preserve ETW attribution and future host_connect_proxy enforcement behavior

Signed-off-by: Akber Raza <akberr@nvidia.com>

* feat(etw): enhance ETW session management with distinct naming for concurrent gateways

* fix(etw): bound the audit queue during overload

- replace the unbounded ETW callback channel with count- and byte-bounded buffering

- keep the ETW callback non-blocking and count records rejected during overload

- emit immediate, rate-limited warnings that identify resulting audit coverage gaps

- make the audit example fail when queue overload causes dropped ETW records

- cover stalled consumers, oversized events, and warning throttling with unit tests

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(etw): harden sandbox audit attribution

- remove command-line and persistent per-PID fallback keys from live and replay resolution

- retire the driver-owned wxc-exec PID before publishing child completion

- retain established identity, activity, and correlation-vector links only for the five-second late-event window

- prevent buffered records from crossing rapid PID retirement and reuse boundaries

- add resolver and lifecycle coverage and document the attribution trust boundary

Signed-off-by: Akber Raza <akberr@nvidia.com>

* chore(mxc): address rebase follow-ups

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(mxc): align OCSF audit example with driver config

- remove unsupported egress proxy settings

- stop requiring the unavailable proxy audit event

- update example documentation for supported event coverage

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(etw): redact command-line secrets in DecodedEtwEvent summary

* fix(etw): enhance PID resolution and event attribution logic for ETW records

* fix(ocsf): restrict gateway-local JSONL sink to Windows/MXC path with opt-in configuration

* address rebase issues

* fix(mxc): address ETW audit review feedback

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(mxc): fail closed across ambiguous PID reuse

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(mxc): bind ETW attribution to process generation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Akber Raza <akberr@nvidia.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Jamie King <jamiek@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
2026-09-11 02:04:26 +00:00
Simon Scatton 48c449d8c8 chore(deps): replace ring with AWS-LC (#3243)
* chore(deps): replace ring with AWS-LC

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* fix(lint): address warnings after dependency upgrades

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* fix(tls): limit provider initialization to reqwest clients

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

---------

Signed-off-by: Simon Scatton <sscatton@nvidia.com>
2026-09-09 17:57:42 +00:00
Evan Lezar e4369adcd0 chore(deps): replace serde_yml with noyalib (#3031)
* chore(deps): replace serde_yml with noyalib

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(providers): annotate generic YAML test values

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-09-07 13:22:27 +00:00
Evan Lezar 8a13bc1298 chore(deps): remove legacy rustls webpki path (#3013)
Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-09-01 09:49:01 +00:00
grs 40d1b48666 feat(provider): support for SPIFFE backed token exchange (#1970)
* feat(provider): add ability to request token exchange instead of client credentials as OAuth grant_type

Signed-off-by: Gordon Sim <gsim@redhat.com>

* test(proxy): add further tests for token exchange

Signed-off-by: Gordon Sim <gsim@redhat.com>

* test(provider): add runnable example for token exchange

Signed-off-by: Gordon Sim <gsim@redhat.com>

* test(e2e): cover Podman token exchange grants

Signed-off-by: Gordon Sim <gsim@redhat.com>

* refactor(oauth): extract duplicated functionality from server and supervisor

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(provider): evict nearest-to-expiry entry from intermediate token cache

Signed-off-by: Gordon Sim <gsim@redhat.com>

* doc(supervisor): add podman example for token exchange

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(provider): withhold token-exchange subject credentials

Signed-off-by: Gordon Sim <gsim@redhat.com>

---------

Signed-off-by: Gordon Sim <gsim@redhat.com>
2026-08-24 05:42:30 +00:00
John T. Myers 4d7f402ce2 feat(policy): establish direct TCP egress foundation (#2711)
* feat(policy): accept explicit tcp endpoint protocol

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* refactor(network): snapshot authoritative egress decisions

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(policy): document explicit tcp protocol

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(policy): defer transparent TCP release guidance

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* chore(go): regenerate sandbox protobuf bindings

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): complete tcp egress foundation

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(policy): document explicit tcp contract

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): fail closed on authorization errors

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(podman): fence delayed exit events before restart

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(policy): validate network endpoint destinations

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(providers): opt in tcp credential fixture

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(policy): require dns host for transparent tcp

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-08-20 16:43:09 +00:00
2f96c53b8c feat(gateway,cli): windows compilation support (#2496)
* chore(windows): gate Unix-only workspace code for MSVC

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* feat(windows): stub unsupported compute drivers

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* ci(windows): add MSVC mise build lane

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* docs(windows): document MSVC build-only design

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* docs(agent): add Windows MSVC build skill

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* feat(windows): add Windows build support

Signed-off-by: Akber Raza <akberr@nvidia.com>

* refactor(windows): consolidate Windows-specific dependencies and improve build logic

Signed-off-by: Akber Raza <akberr@nvidia.com>

* feat(windows): add libclang path resolution and update cargo commands with bundled Z3 features

Signed-off-by: Akber Raza <akberr@nvidia.com>

* chore(tooling): lock Windows tool artifacts

Signed-off-by: Giedrius Burachas <gburachas@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* feat(windows): enhance libclang path resolution to support architecture-specific subdirectories

Signed-off-by: Akber Raza <akberr@nvidia.com>

* Fix Windows dependency gating after sync merge

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(z3): update Z3 header path requirements in Windows build documentation and scripts

Signed-off-by: Akber Raza <akberr@nvidia.com>

* docs(windows): relocate Windows MSVC build design to architecture/

Why: windows-msvc-build-design.mdx is a design document ("design decisions for
the native Windows MSVC build lane"), but it lived in the published, user-facing
docs/reference/ tree. Per AGENTS.md (Documentation) and architecture/README.md
("rfc/ vs architecture/"), design content belongs in architecture/ (or rfc/),
not in published reference. It also shared Fern sidebar "position: 6" with the
MXC compute-driver design page, colliding in the Reference nav ordering.

What:
- Move docs/reference/windows-msvc-build-design.mdx ->
  architecture/windows-msvc-build.md.
- Strip the Fern publish frontmatter and add a plain H1, matching the other
  architecture docs.
- Register it in the architecture doc index in architecture/README.md.
- Repoint the inbound references (build-openshell-mxc-windows skill + reference,
  implement-openshell-mxc-driver skill) to the new path.

With both design pages moved out of docs/reference/, the duplicate position-6
sidebar collision is resolved.

Signed-off-by: Akber Raza <akberr@nvidia.com>

* remove openshell-supervisor-network from unsupported driver package test exclusion list

Signed-off-by: Akber Raza <akberr@nvidia.com>

# Conflicts:
#	tasks/scripts/windows-msvc.ps1

* fix(interceptors): gate unix-only imports so the crate builds on Windows

openshell-gateway-interceptors failed to compile on Windows (E0432: no UnixStream in tokio::net), breaking any Windows build of openshell-server (which depends on it unconditionally). The connect_unix_endpoint fn was already #[cfg(unix)]-gated, but the imports it uses (UnixStream, TokioIo, Uri, service_fn) were left ungated. Gate those four imports with #[cfg(unix)] too. No behavior change on unix; Windows now compiles (no errors, no unused-import warnings).

Signed-off-by: Akber Raza <akberr@nvidia.com>

* feat(windows): add native ARM64 test support

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mise): skip Skaffold on Windows

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(windows): harden ARM64 toolchain discovery

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(windows): scope ARM64 toolchain preflight

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(windows): restore compatibility after GitHub sync

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(windows): avoid rate-limited Z3 source lookup

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mise): skip Helm checks on Windows

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(windows): support repository pre-commit checks

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(windows): stabilize native MSVC validation

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(windows): harden shared Z3 source cache

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>

* fix(windows): avoid leaking MSVC flags into clang-cl

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(windows): complete ARM64 migration audit

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(windows): restore ARM64 Ninja discovery

Signed-off-by: Akber Raza <akberr@nvidia.com>

* refactor(windows): separate platform crate roots

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(windows): restore proto include cfg gating

Signed-off-by: Akber Raza <akberr@nvidia.com>

* refactor: address lint errors

* fix(windows): add preflight check for proxy auth file path

* docs(windows): update GitHub checkout guidance

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(windows): restore CI after dependency updates

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mise): repair Windows sccache lock entry

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(windows): reconcile validation after rebase

Signed-off-by: Akber Raza <akberr@nvidia.com>

* refactor(server): exclude unsupported drivers on Windows

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* refactor(server): isolate platform driver config

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(windows): repair unsupported driver contract test

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(sandbox): remove stale dependencies

Signed-off-by: Akber Raza <akberr@nvidia.com>

* ci(windows): pin x64 workflow actions

Signed-off-by: Akber Raza <akberr@nvidia.com>

* ci(windows): align x64 Rust toolchain

Signed-off-by: Akber Raza <akberr@nvidia.com>

* ci(windows): align ARM64 workflow setup

Signed-off-by: Akber Raza <akberr@nvidia.com>

* refactor(windows): exclude unsupported runtime crates

Signed-off-by: Akber Raza <akberr@nvidia.com>

* refactor(windows): exclude unsupported crates at workspace boundary

Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com>

* refactor(server): gate builtin driver config by platform

Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com>

* fix(sandbox): restore crate documentation

Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com>

* ci(windows): make build workflow manual

Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com>

* ci(windows): temporarily enable pull request builds

Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com>

* ci(windows): cache Rust dependencies

Signed-off-by: Akber Raza <akberr@nvidia.com>

* refactor(windows): remove unnecessary platform changes

Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com>

* ci(windows): make build workflow manual

Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com>

* fix(ci): synchronize mise lockfile

Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com>

* fix(ci): normalize mise provenance metadata

Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com>

* refactor(python): isolate Windows atomic replace retry

Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com>

* fix(python): type Windows permission test errors

Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com>

---------

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>
Signed-off-by: Giedrius Burachas <gburachas@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com>
Co-authored-by: Shailendra Singh <shailendras@nvidia.com>
Co-authored-by: Giedrius Burachas <gburachas@nvidia.com>
Co-authored-by: Jamie King <jamiek@nvidia.com>
Co-authored-by: Piotr Mlocek <pmlocek@nvidia.com>
Co-authored-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com>
2026-08-11 21:00:36 +00:00
alangou 0c7e59a953 fix(deps): bump russh, jsonwebtoken, tar and npm lint deps (#2617)
Signed-off-by: Adrien Langou <alangou@nvidia.com>
2026-08-06 09:39:43 +00:00
5548405fcb feat(credentials): add provider credential storage drivers (#2437)
* feat(credentials): add provider credential storage drivers

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(credentials): harden credential update handling

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(credentials): harden credential driver security, correctness, and performance

Address review findings from the credential storage drivers PR:

- Route additional_credentials through the driver on refresh to prevent
  silent data loss for multi-credential providers (e.g. AWS STS)
- Clean up stored credential handles on CAS failure during refresh to
  prevent orphaned secrets in external backends
- Enforce namespace validation in the Kubernetes Secrets driver to
  prevent cross-namespace credential access when allow_reference_namespace
  is not enabled
- Cache Vault Kubernetes auth tokens with 80% TTL to avoid re-authenticating
  on every credential operation
- Parallelize resolve_credentials in all three drivers using try_join_all
  for faster sandbox startup
- Add existingSecret support for the KEK Secret to fix helm template/GitOps
  workflows where lookup returns empty and regenerates the key
- Document RBAC blast radius for the Kubernetes Secrets credential driver
  and recommend a dedicated namespace

Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com>
Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>

* fix(credentials): add optimistic concurrency, fix thundering herd, parallelize operations

Use resourceVersion optimistic concurrency with retry loop for K8s
Secret ownership checks to prevent TOCTOU races. Switch Vault token
cache from RwLock to Mutex with double-check pattern to prevent
thundering herd on cache miss. Parallelize credential store and delete
operations across independent keys using try_join_all.

Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com>
Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>

* fix(credentials): handle partial failures, add delete retry, consolidate cleanup

Replace try_join_all with join_all in credential store/delete operations
to handle partial failures — successfully-stored handles are cleaned up
when another key fails. Add retry loop with conflict detection to
db-credstore delete_credential, matching the K8s driver pattern.
Consolidate 4 manual cleanup_pre_stored_provider_credentials call sites
into a single error handler using an async block. Remove inconsistent
.trim() from db-credstore validate_handle_owner.

Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com>
Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>

* fix(credentials): fix retry loop guard and remove unprotected validation

Remove attempt-count guard from 409/Aborted match arms in retry loops
so the post-loop Status::aborted error is reachable after exhausting
retries. Previously, last-attempt conflicts fell through to the
catch-all error arm, producing misleading Status::unavailable errors.

Remove duplicate validation calls that ran after
prepare_provider_credential_update but outside the cleanup-protected
async block, which would leak pre-stored handles on failure.

Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com>
Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>

* fix(credentials): add workspace/provider UUID to credential backend paths

Include workspace and provider ID in credential backend object paths to ensure
cross-workspace uniqueness and prevent credential collision (GATOR-1806c9be-01).

- Updated credential driver proto to include workspace and provider_id fields
- Modified Vault driver to include workspace/provider_id in managed_secret_path
- Modified Kubernetes Secrets driver to include workspace/provider_id in
  credential_owner_id and managed_secret_name
- Updated all credential runtime calls to pass workspace/provider_id
- Updated tests to use the new signatures

This prevents two workspaces sharing the same external credential store from
colliding on provider names, which was a critical security issue (CWE-639).

* fix(credentials): preserve provider-level expiration for handle-backed credentials

Compute effective expiration from both provider and driver values using the
earliest non-zero timestamp and skip expired values before insertion
(GATOR-1806c9be-02).

- Modified resolve_provider_handles to check provider credential_expires_at_ms
- Skip expired credentials during resolution instead of returning them
- Use effective expiration (min of provider and driver) in resolution results
- Fix inference.rs to preserve earliest expiration when merging

This ensures handle-backed credentials respect the same expiration semantics
as inline credentials.

* fix(credentials): stage refresh changes under new handles before validation

Stage credential replacements under new immutable handles instead of reusing
existing handles to prevent overwriting committed values before validation/CAS
(GATOR-1806c9be-03).

- Stage credentials with empty existing_handles map to force new handle creation
- Validate and CAS before the new values are committed to backend storage
- Delete old handles only after successful CAS
- On CAS failure, delete only the newly staged handles
- This prevents CWE-362/CWE-367 race conditions where failed refreshes could
  still modify or delete the active credential

The fix ensures that a rejected refresh cannot modify the backend object still
referenced by the committed provider record.

* fix(credentials): add timeouts to credential driver RPCs

Apply configured timeouts to both startup capability negotiation and runtime
RPCs to prevent indefinite hangs (GATOR-1806c9be-05).

- Add DEFAULT_CREDENTIAL_DRIVER_RPC_TIMEOUT_SECS constant (30s)
- Apply timeout to GetCapabilities during startup connection
- Apply timeout to all runtime RPCs (store, delete, resolve)
- Use tokio::time::timeout to bound the entire GetCapabilities operation
  during startup, not just the socket connection
- Return contextual deadline errors on timeout

This prevents a faulty or overloaded driver from hanging gateway operations
indefinitely.

* fix(credentials): fix test to use consistent workspace/provider identity

The Kubernetes auth Vault resolve test was constructing a managed path with
test-workspace/test-provider-id but sending default/prov-123 in the request,
causing validation to reject the request (GATOR-18e32351-01).

- Update test to use test-workspace and test-provider-id in the request to
  match the logical_path construction
- This ensures the test exercises the intended code path and validates
  Kubernetes auth resolution properly

The test now passes and correctly validates identity enforcement.

* fix(credentials): use unique staging ID for refresh to avoid overwrites

Stage refresh replacements under genuinely distinct immutable handles using
a unique staging ID to prevent overwriting committed values (GATOR-1806c9be-03).

- Generate a unique staging ID using UUID for each refresh operation
- Use this staging ID when storing credentials instead of the real provider ID
- Pass the same staging ID during cleanup on failure to delete only staged objects
- This ensures deterministic paths (Vault) and object names (K8s) don't collide
  with the committed provider's credentials

The fix prevents failed refreshes from silently replacing active credentials
or breaking providers by deleting still-referenced backend objects.

* fix(credentials): wrap credential driver RPCs in local timeouts

Add local tokio::time::timeout wrappers around credential driver RPCs to
bound non-compliant or stalled UDS peers (GATOR-1806c9be-05).

- Wrap StoreCredential, DeleteCredential, and ResolveCredentials in local timeouts
- Return contextual deadline_exceeded errors when timeouts occur
- Keep existing gRPC timeout metadata for compliant implementations
- GetCapabilities during startup was already wrapped in previous commit

This ensures a faulty local driver cannot hang gateway operations indefinitely,
even if it accepts the connection but never responds to the RPC.

* fix(credentials): preserve ownership for staged refreshes

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(credentials): bound startup capability probe

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* test(provider): authenticate credential handler requests

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(ci): grant actions read to credential driver e2e

Signed-off-by: Seth Jennings <sjenning@redhat.com>

---------

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com>
Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>
Signed-off-by: Seth Jennings <sjenning@redhat.com>
Co-authored-by: Taylor Mutch <taylormutch@gmail.com>
Co-authored-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>
2026-08-05 16:29:41 +00:00
Polite_realismandAdam Miller 0e9a44cfa9 feat(build): add system CA root mode (#2324)
* feat(build): add system CA root mode

Allow distro builds to use native trust stores for supervisor upstream TLS while keeping bundled Mozilla roots as the default. Avoid bundled root crates in system-ca-roots builds by using native-root TLS features and z3 0.20.

Signed-off-by: Adam Miller <admiller@redhat.com>

* fix(build): keep CA root feature in telemetry-off verification

The telemetry-off task uses --no-default-features which now disables
bundled-ca-roots in addition to telemetry, triggering the compile_error
guard. Re-enable bundled-ca-roots explicitly so the task verifies only
telemetry compilation.

Signed-off-by: Scott Burdine <sburdine@nvidia.com>
Signed-off-by: politerealism <burdcat17@gmail.com>

* fix(sdk): disable oauth2 default features to prevent webpki-roots leak

The bare `oauth2 = "5"` dependency re-enabled default features
(rustls-tls → reqwest/rustls-tls → webpki-roots), defeating the
system-ca-roots feature gate. Mirror the CLI fix: disable defaults
and enable only the `reqwest` feature.

Signed-off-by: Quinn Burdine <sburdine@redhat.com>
Signed-off-by: politerealism <burdcat17@gmail.com>

* refactor(build): simplify CA root selection to single feature toggle

Replace mutually exclusive bundled-ca-roots / system-ca-roots features
with a single bundled-ca-roots toggle. Disabling it implies system roots
via rustls-native-certs, which is now a regular (non-optional) dependency.
This fixes cargo --all-features and simplifies the distro build interface
from --no-default-features --features system-ca-roots to just
--no-default-features.

Signed-off-by: Quinn Burdine <sburdine@redhat.com>
Signed-off-by: politerealism <burdcat17@gmail.com>

* feat(build): add system-ca-roots convenience alias and fix verify task

Add a system-ca-roots feature alias on openshell-sandbox that includes
all other defaults (telemetry) except bundled-ca-roots, so distro
builds can use --no-default-features --features system-ca-roots without
manually re-adding unrelated defaults. Update the verify CI task to use
the alias and scope checks to the sandbox package. Fix task description
to use "build mode" terminology instead of implying a Cargo feature.

Signed-off-by: Quinn Burdine <sburdine@redhat.com>
Signed-off-by: politerealism <burdcat17@gmail.com>

* ci: fix system CA roots step name to use build mode terminology

Signed-off-by: Quinn Burdine <sburdine@redhat.com>
Signed-off-by: politerealism <burdcat17@gmail.com>

* refactor(sandbox): reorder features to place system-ca-roots alias near default

Signed-off-by: Quinn Burdine <sburdine@redhat.com>
Signed-off-by: politerealism <burdcat17@gmail.com>

* fix(proxy): unwrap Result from build_upstream_client_config in tests

The function signature changed to return Result but the test call sites
were not updated, causing type mismatch compilation errors in CI.

Signed-off-by: Quinn Burdine <sburdine@redhat.com>
Signed-off-by: politerealism <burdcat17@gmail.com>

---------

Signed-off-by: Adam Miller <admiller@redhat.com>
Signed-off-by: Scott Burdine <sburdine@nvidia.com>
Signed-off-by: politerealism <burdcat17@gmail.com>
Signed-off-by: Quinn Burdine <sburdine@redhat.com>
Co-authored-by: Adam Miller <admiller@redhat.com>
2026-08-04 15:55:44 +00:00
Evan LezarandDrew Newberry d220d89468 feat(compute): negotiate gateway callback listeners (#2492)
* feat(compute): query gateway listener requirements

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* feat(compute): add Podman listener requirements

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(docker): use default gateway bind address

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(gateway): avoid wildcard primary listener

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(podman): validate callback listener discovery

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(server): support split dual-stack listeners

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(podman): support legacy rootless listener discovery

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(e2e): accept loopback plaintext rejection

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* docs(agent): add callback listener diagnostics

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(server): restrict compute callback listeners

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(podman): validate local callback port

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(server): clarify callback listener contract

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(podman): require pasta for local callbacks

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* docs(gateway): document RPM listener default

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* refactor(server): keep listener provenance diagnostic-only

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(compute): preserve callback listener isolation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(e2e): remove Podman callback relay

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(packaging): preserve Podman callback loopback

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci(e2e): run VM smoke on nested-virt runner

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci(e2e): gate VM smoke on usable KVM

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci(e2e): probe KVM through VM driver

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci(e2e): tolerate hosted KVM denial

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(server): close traced futures before assertions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* revert: remove tracing test stabilization

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Evan Lezar <elezar@nvidia.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
2026-07-31 16:41:06 +00:00
krishicks fa24299092 feat(gateway): export traces over OTLP (#2534)
Add an opt-in OTLP/gRPC trace exporter to the gateway. Export is enabled
by the presence of an `[openshell.gateway.otlp]` table with an endpoint;
there is no separate toggle.

Instrumented:

- Inbound request server spans, named for the RPC (`$service/$method`) or
  `{method} {path}` for plain HTTP. They continue valid W3C `traceparent`
  context when present and start a new trace otherwise. gRPC spans also
  carry `rpc.system`, `rpc.service`, `rpc.method`, and trailer-derived
  `rpc.grpc.status_code`.
- Compute driver calls (create, delete, list, get, validate, watch) as
  client spans anchored on the `ComputeDriver` contract.
- Store reads and writes as children of the current request or loop span.
- Work with no inbound request: compute driver initialization, the sandbox
  reconcile sweep, provider credential refresh tick, and driver watch events.
  Each roots one operation trace so its child work does not arrive as anonymous
  single-span traces.

This is deliberately not exhaustive. Auth, policy evaluation, and
middleware remain uninstrumented, as do store lifecycle calls (`ping`,
`close`) that a readiness poll would turn into a span per tick. The aim is
a useful trace tree at a reviewable size; coverage can grow against real
traces.

Design notes:

- The TOML table owns whether and where to export. The SDK `OTEL_*`
  variables own how; sampling, batching, and limits are not mirrored into
  gateway config.
- The OpenTelemetry layer exports spans only. Existing `tracing` events
  remain on the stdout and sandbox-log paths and are not copied into trace
  payloads.
- Telemetry never blocks the gateway. A malformed endpoint logs an error
  and disables export rather than failing startup, and buffered spans are
  drained during graceful shutdown.
- Failed spans carry error status without a separate `error.type` attribute.
  Request spans use HTTP status and gRPC response trailers; driver spans use
  the returned gRPC status; autonomous loop spans record failed results
  explicitly. Store spans exempt `UniqueViolation` and `Conflict`, because
  those errors report expected contention such as a held lease or an
  optimistic-concurrency retry.
- The compute driver is reachable only through `TracedDriver::call`, so a
  call cannot skip its span. This is the client half of a client/server pair
  and the single place to inject context if drivers move out of process.
- Tests share one process-wide subscriber and in-memory exporter because
  `tracing` caches callsite interest globally.

Inbound W3C trace context is propagated into gateway request spans. Context
is not yet injected into outbound driver calls, so a future out-of-process
driver would still need propagation at the `TracedDriver` seam.

The Helm chart is intentionally unchanged, so OTLP export cannot yet be
enabled on a chart-deployed gateway.

Refs #2507

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-07-30 20:29:42 +00:00
Artem Lytvyn 5432d01d5a feat(tui): add config key support to provider create/update forms (#2224)
* feat(tui): add config key editing to create provider form

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* feat(tui): add config key editing to update provider form

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* fix(tui): add Up/Down navigation for config_cursor in provider forms

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* fix(tui): show config values in provider detail view

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* fix(tui): improve provider config form validation and reduce duplication

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* fix(tui): fix config deletion and invisible cursor in provider forms

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* test(tui): add tests for config deletion tombstones and cursor scroll window

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* fix(tui): send only config delta on provider update

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* fix(tui): flush pending config input on provider update submit

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* fix(tui): fix provider config form loading, focus reset, and overflow

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* fix(tui): separate config entry navigation from key input focus

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* fix(tui): move config cursor to newly added entry after flush

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* fix(tui): always populate provider_entries regardless of providers_v2_enabled

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

---------

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>
2026-07-22 18:26:22 +00:00
Piotr Mlocek d556748771 feat(supervisor-middleware): add network egress middleware (#2027)
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-07-16 17:47:49 -07:00
Russell Bryant aa483ecb9a feat(providers): AWS STS AssumeRole refresh strategy and aws-s3 profile (#1782)
Add gateway-managed AWS STS credential refresh (provider-v2, #1576). The
gateway calls sts:AssumeRole and writes three short-lived credentials
(AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, AWS_SESSION_TOKEN) to the
provider record; the proxy re-signs requests with SigV4. Adds the aws and
aws-s3 provider profiles and a declarative multi-output refresh model
(additional_outputs) so one AssumeRole co-mints all three credentials.

Signed-off-by: Russell Bryant <rbryant@redhat.com>
2026-07-16 13:28:01 -07:00
Drew Newberry 83003e80fc feat(interceptors): initial gateway interceptor implementation and reference example (#2005)
* feat(gateway): add descriptor-driven interceptors

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(gateway): add service-reflected interceptors

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* wip

* fix(gateway): harden interceptor evaluation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(interceptors): label metrics and harden governance smoke

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* remove on_error: ignore

* feat(gateway-interceptors): emit log annotations

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(examples): govern provider profiles in interceptor

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(gateway): preserve update config annotations

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(providers): support interceptor profile catalogs

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* wip

* feat(governance-interceptor): sign provider profiles

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(providers): use configured profile sources for refresh updates

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(gateway-interceptors): add phase-specific evaluation payloads

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(providers): compose provider profile sources

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(gateway-interceptors): preserve committed responses

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(gateway-interceptors): reject ambiguous protobuf oneofs

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(gateway-interceptors): validate patch candidates per binding

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(gateway-interceptors): use reflected protobuf codec

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(governance-example): canonicalize signed protobuf hashes

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(gateway): commit policy provenance atomically

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(gateway): close signed governance bypasses

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(gateway): isolate interceptor secrets and authority

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(gateway): snapshot provider profiles per request

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* chore(gateway): resolve server clippy warnings

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(server): satisfy provider source clippy lint

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(server): initialize policy test annotations

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(gateway-interceptors): require explicit route allowlist

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(proto): clarify update annotation semantics

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-07-14 20:58:36 -07:00
krishicksandddurst 7bce1223dc feat(policy): add JSON-RPC and MCP L7 policies (#1865)
Add policy schema, proto, provider profile, OPA, and L7 proxy support for
`protocol: json-rpc` and `protocol: mcp`. Generic JSON-RPC endpoints match
exact method names only, with `method: "*"` as the all-method sentinel;
wildcard/glob methods and params matchers are rejected.

Parse JSON-RPC request bodies and batches in the forward proxy, deny
response-shaped client frames, limit receive-stream GET allowance to MCP
endpoints, and redact params in decision logs. Preserve L7 rule params on the
proto load path so MCP `tools/call` tool filters behave like YAML-loaded
policies.

Add MCP conformance coverage, JSON-RPC L7 e2e coverage, and docs for the new
protocols and current matcher limitations.

Signed-off-by: Kris Hicks <khicks@nvidia.com>
Co-authored-by: ddurst <267424412+ddurst-nvidia@users.noreply.github.com>
2026-06-26 15:52:16 -07:00
Taylor Mutch 702cbc4f63 feat(providers): support SPIFFE-backed token grants (#1784)
* feat(providers): support SPIFFE-backed token grants

Add provider profile token_grant metadata and expand endpoint-specific
dynamic credentials so sandbox supervisors can request SPIFFE JWT-SVIDs,
exchange them with an OAuth-style token endpoint, cache returned access
tokens, and inject bearer tokens into matching HTTP requests.

Wire Kubernetes and Helm deployments to mount the provider SPIFFE Workload
API socket into sandbox pods for token grant exchange.

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
Signed-off-by: Gordon Sim <gsim@redhat.com>

* test(examples): add SPIFFE token grant demo

Add a reusable alpha/beta demo that deploys a SPIFFE-verifying token issuer
and protected services, imports a token-grant provider profile, creates a
sandbox, and verifies endpoint-specific bearer tokens.

The script leaves Kubernetes workloads in place, deletes sandboxes through
openshell unless KEEP_SANDBOX=1, and prints protected service logs as proof
of life.

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(providers): harden SPIFFE token grants

* fix(providers): harden dynamic token grants

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(providers): harden token grant handling

---------

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
Signed-off-by: Gordon Sim <gsim@redhat.com>
2026-06-10 10:54:39 -07:00
Colin Walters e7f965a988 refactor(sandbox,driver-vm): Start moving to rustix (esp over libc unsafe) (#1505)
In the Rust ecosystem there's largely three ways to do system calls:

- raw libc
- nix
- rustix

Of the three, libc is almost all `unsafe` and really 95% of use
cases should be either nix or rustix. nix is the original one,
but after having looked at the code of both, I think rustix
is just better designed and organized. It's also reached 1.0,
whereas nix is still making semver-breaking changes (in fact
we're behind here in this project).

Now in practice, we have both *transitively* in the depchain
already, and that's true for quite a lot of projects.

But I think rustix is better, so let's add rustix as
a workspace dependency (process feature) and migrate
a few use cases to it - it's especially better than the raw
libc which is suprisingly widespread.

If we agree to do this, then many other calls can be ported.

Signed-off-by: Colin Walters <walters@verbum.org>
2026-05-21 15:20:21 -05:00
Taylor MutchandDrew Newberry b61a98dbad feat(gateway): add TOML configuration file (RFC 0003) (#1317)
* feat(gateway): add TOML configuration file (RFC 0003)

Introduces an opt-in --config / OPENSHELL_GATEWAY_CONFIG flag that loads a
TOML file with gateway-wide settings and per-driver tables. Source
precedence is CLI > env > file > built-in default, implemented via clap's
ValueSource so existing flags and env vars keep their priority.

Driver crates (kubernetes, docker, podman, vm) now derive Deserialize on
their config structs. SupervisorSideloadMethod gains Deserialize with
kebab-case rename. A per-driver inheritance allowlist on the loader side
overlays [openshell.gateway] shared defaults (default_image,
supervisor_image, image_pull_policy, guest_tls_*, ssh_handshake_skew_secs,
client_tls_secret_name, host_gateway_ip, enable_user_namespaces) onto
each [openshell.drivers.<name>] table before deserialization.

The Helm chart renders a new gateway-config ConfigMap and mounts it at
/etc/openshell/gateway.toml. The migrated OPENSHELL_* env entries are
dropped from the StatefulSet — only the Secret-backed
OPENSHELL_SSH_HANDSHAKE_SECRET remains. database_url stays on --db-url.

Adds examples/gateway/gateway.example.toml and updates architecture/gateway.md
with the source precedence and inheritance rules.

* docs(gateway): drop ssh_handshake_skew_secs and ssh_handshake_secret from examples

Both fields are scheduled for removal. Remove the example values and the
env-only note so the gateway.toml example and the architecture doc stop
recommending settings that will not exist much longer.

* docs(rfc): correct OPENSHELL_CONFIG to OPENSHELL_GATEWAY_CONFIG in RFC 0003

* docs(gateway): add per-driver TOML example configurations

Adds focused single-driver examples next to the comprehensive
gateway.example.toml: kubernetes, docker, podman, and microvm. Each one
demonstrates the realistic settings for that driver plus how shared
[openshell.gateway] defaults inherit into the driver table.

A new unit test (`checked_in_examples_parse`) loads every example through
the config_file loader so schema drift fails CI rather than silently
shipping a broken example.

* refactor(gateway): drop image_pull_policy from shared inheritance

Kubernetes and Podman use mutually-incompatible vocabularies for the same
TOML key:

  - Kubernetes: `Always | IfNotPresent | Never` (free-form string passed
    verbatim to the K8s API).
  - Podman: `always | missing | never | newer` (strict lowercase enum
    deserialised into `ImagePullPolicy`).

No value means the same thing in both drivers. Sharing the key at
`[openshell.gateway]` scope and inheriting it into every active driver's
table meant any value safe for one driver was either wrong or silently
dropped for the other (`IfNotPresent` → `ImagePullPolicy::Missing` after
`.unwrap_or_default()`). Operators run one driver per gateway, so the
"shared default" never pays for itself.

Make `image_pull_policy` driver-local:

  - Remove the field from `GatewayFileSection` and from
    `inheritable_keys()` for both Kubernetes and Podman.
  - Drop the file→`RunArgs` merge for the gateway-scope key.
  - Stop unconditionally clobbering the driver value with
    `config.sandbox_image_pull_policy` in the runtime wiring — only apply
    the CLI/env override when it was set (and, for Podman, only when it
    parses into the lowercase enum).
  - Move the key under `[openshell.drivers.kubernetes]` and
    `[openshell.drivers.podman]` in every example, the RFC, the
    architecture doc, and the Helm-rendered gateway ConfigMap.

The supervisor pull policy follows the same shape: it is K8s-only and
moves into `[openshell.drivers.kubernetes]` alongside `image_pull_policy`
in the Helm template.

* fix(gateway): address review feedback on TOML configuration

Resolves the P1 and P2 issues raised in PR #1317:

- Helm gateway ConfigMap moves `grpc_endpoint` under
  `[openshell.drivers.kubernetes]` so the default install no longer fails
  the gateway's `deny_unknown_fields` schema check.
- `kubernetes_config_from_file` and `podman_config_from_file` only let
  the gateway-wide CLI/env `grpc_endpoint` overwrite the driver-table
  value when it was actually supplied, preserving file-only configs.
- Kubernetes driver default `image_pull_policy` is now empty (was Podman
  vocabulary "missing"), so default deployments let the Kubernetes API
  apply its own policy instead of being rejected.
- New `disable_tls` gateway field plumbs `.Values.server.disableTls`
  through the TOML ConfigMap instead of relying on env vars dropped from
  the StatefulSet.
- StatefulSet pod template now carries a `checksum/gateway-config`
  annotation so `helm upgrade` rolls pods when the ConfigMap changes.
- Auxiliary listener resolution preserves the full `SocketAddr` from
  `health_bind_address` / `metrics_bind_address`, so a loopback-pinned
  health port is not silently relocated onto the public bind address.
- `ssh_session_ttl_secs` from the file is now applied to `Config` (it
  was previously accepted by the loader but never read).

New regression coverage: cli-level merge tests for the new fields plus
helm-unittest assertions for the ConfigMap shape, checksum annotation,
and `disable_tls` rendering.

* docs(gateway): consolidate gateway TOML examples into docs reference

Replaces the per-driver example files under examples/gateway/ with a
single published reference page at docs/reference/gateway-config.mdx
covering source precedence, layout, the full example, and the four
per-driver examples (Kubernetes, Docker, Podman, microVM). Drew flagged
during PR #1317 review that the examples belong with the user-facing
docs rather than in a sibling examples/ directory.

The cross-references in architecture/gateway.md and RFC 0003 are updated
to point at the new docs page; the round-trip test in config_file.rs is
removed (schema coverage stays on the inline parses_full_example test
and per-field merge tests — doc-snippet drift belongs in a separate
docs-lint, not in a cross-tree Rust unit test).

* refactor(core): move DEFAULT_K8S_NAMESPACE into K8s driver

The constant is Kubernetes-specific (used only by KubernetesComputeConfig's
Default impl) and does not belong in openshell-core. Relocate it to the
driver crate that owns the K8s vocabulary; openshell-core retains only
truly cross-cutting defaults.

* refactor(core): move Podman bridge default into Podman driver

DEFAULT_NETWORK_NAME is Podman vocabulary, consumed only by the Podman
driver. Also drops the unused DEFAULT_IMAGE_PULL_POLICY constant.

* docs(auth): scrub remaining SSH handshake secret references

Sweeps the trailing mentions left after the rebase: the gateway
config-file module doc, the Helm gateway-config ConfigMap header,
the gateway-config.mdx env-only note, and the RPM systemd unit
comment for init-gateway-env.sh.

* docs(gateway): clarify OPENSHELL_GRPC_ENDPOINT applies to all drivers

The previous comment implied the callback endpoint was Kubernetes-only,
but the value is propagated to every compute driver (Kubernetes, Docker,
Podman, VM) and must be reachable from wherever the sandbox runs.

* refactor(gateway): move driver options into config (#1394)

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(e2e): regenerate gateway config via TOML for docker + podman harnesses

The gateway CLI flags moved into TOML config tables in 560550d2 (#1394),
which made every existing e2e/with-{docker,podman}-gateway.sh invocation
fail with "unexpected argument '--sandbox-namespace'" (and a long tail
of similar driver-specific options) before the gateway could even bind.

Replace the obsolete CLI flags with a synthesized
`[openshell.drivers.<driver>]` table written to `${STATE_DIR}/gateway.toml`
and passed via the new `--config` flag. Only the gateway-wide flags
that survived 560550d2 (bind-address, port, drivers, db-url, tls-*,
disable-tls, log-level, health-port) stay on the command line.

Both scripts get a small `toml_string` helper to properly TOML-quote
the values (the previous `%q` printf format produced bash-escape, not
TOML-escape). The Docker harness also corrects two field names that
diverged from the driver schema: `docker_network_name` →
`network_name`, and the supervisor binary/image plumbing now reads
through to `supervisor_bin` / `supervisor_image` in the same table.

The Podman harness drops `--ssh-gateway-port` (deleted in 560550d2 —
gRPC + SSH are multiplexed on the same port now) and substitutes
`network_name` + `gateway_port` for the obsolete `--sandbox-namespace`
(which the Podman driver never had as a typed field).

* fix(core): swap bind-only 0.0.0.0 SSH gateway host for cluster URL host

CLI's resolve_ssh_gateway treated 0.0.0.0 as a loopback "keep as-is"
when the cluster URL was also loopback, so the SSH proxy connected to
0.0.0.0:port. The unspecified address is never a valid connect target
and is not present in any TLS cert SAN, which produced BadCertificate
TLS handshake failures during `openshell sandbox create -- ...` in
docker/podman e2e (e.g. bypass_detection).

Resolution: when the server returns 0.0.0.0 or :: as the gateway host
and both endpoints are loopback, fall back to the cluster URL's host
(which the CLI is already using to reach the gateway, so it must
resolve and match the cert).

* fix(e2e): repair podman harness on macOS

Podman 5.x with the applehv/libkrun provider no longer creates the legacy
~/.local/share/containers/podman/machine/podman.sock symlink, and
`podman system service` is a Linux-only subcommand — the macOS client
delegates the API service to the VM. Both assumptions in the harness
were stale, so the script tried to start a temporary service that podman
rejected with "unknown flag: --time".

- Discover the macOS socket via `podman machine inspect` instead of the
  hardcoded path.
- On Darwin, fail fast with a "start podman machine" message rather than
  attempting the Linux-only `podman system service` fallback.
- Write socket_path into [openshell.drivers.podman] so the in-process
  driver picks up the discovered socket; the driver reads TOML only
  after the config refactor (560550d2), so OPENSHELL_PODMAN_SOCKET alone
  was no longer enough.

* fix(server): use clone_from for TLS client CA assignment

clippy 1.95.0 rejects assigning the result of `Clone::clone()` to an
existing variable under `-D warnings` (`assigning_clones`). Switch to
`clone_from(&...)` to satisfy the lint and avoid the redundant
allocation.

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
2026-05-15 12:43:48 -07:00
Drew Newberry f5b546e41b Revert "perf(build): speed up local CLI rebuilds (#1387)" (#1395)
This reverts commit 668c712b63.
2026-05-14 17:45:37 -07:00
John T. Myers 668c712b63 perf(build): speed up local CLI rebuilds (#1387) 2026-05-14 11:44:36 -07:00
Seth Jennings afcd3a9ec4 fix(cli): use OS trust store for reqwest TLS verification (#1342)
Switch reqwest from `rustls-tls` (bundled Mozilla webpki roots) to
`rustls-tls-native-roots` (OS certificate store) so that the CLI
trusts the same CAs as the rest of the system. This fixes OIDC
discovery failures against Keycloak instances using certificates
signed by internal CAs present in the system trust bundle.
2026-05-12 16:41:01 -07:00
Drew Newberry 028763d4db refactor(vm): remove legacy openshell-vm crate (#1239)
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-05-07 10:50:20 -07:00
Drew Newberry 4d388d2677 ci(vm): cleanup vm build infra (#1186) 2026-05-06 08:58:02 -07:00
Saurabh Agarwal 04e48d585c feat(server): add request-ID middleware for request correlation (#1082)
Add a UUID-based request-ID middleware using tower-http's request-id
feature. Each inbound request receives a unique x-request-id header
(or preserves a client-supplied one), which is recorded in the tracing
span and propagated to the response.

This enables operators to correlate log lines across the middleware
stack for a single request under concurrent load, and lets clients
reference specific requests in bug reports.

Signed-off-by: sauagarwa <sauagarw@redhat.com>
2026-05-04 16:14:58 -07:00
John T. Myers 6b21804258 feat(policy): add GraphQL L7 inspection (#1083)
Support GraphQL L7 policies
2026-05-04 11:50:09 -07:00
Mrunal Patel 084505425b feat(auth): add OIDC/Keycloak authentication with RBAC and scope-based permissions (#935)
* feat(auth): add OIDC/Keycloak authentication with RBAC

Add OAuth2/OIDC authentication to the gateway server with role-based
access control, CLI login flows, and full deployment plumbing.

Server: JWT validation against configurable OIDC issuer (oidc.rs),
JWKS key caching with TTL and rotation handling, method classification
(unauthenticated/sandbox-secret/dual-auth/bearer), identity extraction
with provider-agnostic Identity type, and RBAC enforcement via
AuthzPolicy with configurable admin/user roles and auth-only mode.

CLI: browser-based Authorization Code + PKCE flow, Client Credentials
flow for CI/automation, token storage with refresh, gateway add/login/
logout commands, OIDC bearer token injection over mTLS transport,
discovery endpoint for auto-configuration.

Security: sandbox-secret scope restriction on UpdateConfig (policy
sync only), anti-spoofing header stripping, dual-auth fallthrough
from sandbox-secret to Bearer token.

Deployment: OIDC config wired through DeployOptions, Docker env vars,
Helm values/templates, HelmChart manifest, cluster-entrypoint.sh, and
bootstrap scripts. Keycloak dev server script with pre-configured
realm (test users, roles, PKCE client, CI client).

Tested with Keycloak. The roles claim path and role names are
configurable to support other OIDC providers.

* feat(auth): add OAuth2 scope-based fine-grained permissions

Add opt-in scope enforcement on top of existing OIDC role-based access
control. When --oidc-scopes-claim is set, the server extracts scopes
from the JWT and checks them per-method against an exhaustive scope map.

Scopes: sandbox:read, sandbox:write, provider:read, provider:write,
config:read, config:write, inference:read, inference:write, and
openshell:all (wildcard). Methods not in the scope map require
openshell:all. Scopes layer on top of roles and cannot escalate
privilege. Auth-only mode (empty role names) still enforces scopes
when enabled.

Server: scopes_claim in OidcConfig, scope extraction from JWT
(space-delimited and JSON array formats), standard OIDC scope
filtering, scope check in AuthzPolicy after role check.

CLI: --oidc-scopes on gateway add/start stored in metadata and
consumed by gateway login, --oidc-scopes-claim on gateway start
forwarded to server, scopes parameter in browser and client
credentials OAuth2 flows with openid deduplication.

Deployment: oidc_scopes_claim wired through DeployOptions, docker.rs,
Helm, bootstrap scripts, and cluster entrypoint.

Keycloak: realm config updated with built-in OIDC scopes and 9
OpenShell client scopes as optional on openshell-cli and openshell:all
as default on openshell-ci.

* fix(auth): address branch review findings

Add GetInferenceBundle to sandbox-secret methods so sandbox inference
route refresh works under OIDC. Make GetSandboxConfig dual-auth so CLI
users can read sandbox settings with Bearer tokens.

Preserve OIDC gateway metadata on restart — a bare gateway start
without --oidc-* flags no longer erases the stored OIDC registration.

Document CI client ID requirement (openshell-ci vs openshell-cli) in
the testing guide. Add security note about auth-only mode blast radius
for GitHub Actions.

* fix(auth): complete review findings for OIDC auth boundary

Move OpenShell/GetSandboxConfig from sandbox-secret-only to dual-auth
so CLI users can read sandbox settings with Bearer tokens while sandbox
supervisors continue using the shared secret.

Add sandbox secret interceptor to the inference bundle fetch path so
GetInferenceBundle works under OIDC-enabled gateways. Extract shared
interceptor constructor to avoid duplication.

Add GetSandboxConfig to the config:read scope map so scope enforcement
applies consistently when scopes are enabled.

Refactor OIDC metadata preservation into apply_oidc_gateway_metadata()
with explicit resume semantics — only preserve existing OIDC metadata
on real resume paths, not on fresh deployments.

Update architecture docs and testing guide to reflect the corrected
method classifications and add new test coverage for interceptor
injection, scope requirements, metadata preservation, and dual-auth
classification.

* refactor(auth): use oauth2 crate for CLI OIDC flows

Replace hand-written PKCE generation, authorization URL construction,
token exchange, client credentials, and token refresh with the oauth2
crate's typed API.

Eliminates sha2, hex, and getrandom dependencies from the CLI. The
custom urlencoded() helper and manual form POST logic are replaced by
BasicClient methods with proper type-state safety.

Discovery and the callback server remain custom since the oauth2 crate
does not provide OIDC discovery or a localhost redirect listener.

* refactor(auth): move server auth modules into auth/ directory

Group oidc.rs, authz.rs, identity.rs, and the auth HTTP endpoints
under src/auth/ module directory. No behavioral changes.

  auth/mod.rs      — module root, re-exports HTTP router
  auth/oidc.rs     — JWT validation, JWKS caching, method classification
  auth/authz.rs    — role and scope authorization policy
  auth/identity.rs — provider-agnostic Identity type
  auth/http.rs     — /auth/connect and /auth/oidc-config endpoints

* fix(auth): use RequestBody auth type for client credentials flow

The oauth2 crate defaults to BasicAuth (HTTP Basic header) but Keycloak
and most OIDC providers expect client_secret_post (credentials in the
request body). Set AuthType::RequestBody explicitly to match the
pre-refactor behavior.

Also re-export Identity, IdentityProvider, and JwksCache from the auth
module so ServerState's public API remains nameable by external consumers.

* fix(auth): forward OPENSHELL_OIDC_SCOPES through cluster bootstrap

Pass --oidc-scopes to gateway start so the metadata includes requested
scopes after cluster bootstrap. Without this, users had to manually
edit metadata.json to set scopes for gateway login.

Usage: OPENSHELL_OIDC_SCOPES="openshell:all" mise run cluster

* test(auth): add OIDC e2e tests for RBAC, scopes, and client credentials

Add 10 end-to-end tests covering OIDC authentication against a live
K3s cluster with Keycloak:

RBAC (5 tests): admin can create providers, user cannot, user can list
sandboxes, unauthenticated requests rejected, health probe works
without auth.

Scopes (4 tests): sandbox-scoped token can list sandboxes but not
providers, openshell:all grants full access, no-scopes token denied.

Client credentials (1 test): CI token via client_credentials grant.

Tests are opt-in via OPENSHELL_E2E_OIDC=1 and OPENSHELL_E2E_OIDC_SCOPES=1
env vars. They derive the Keycloak URL from gateway metadata to match
the server's configured issuer.

Run with:

  OPENSHELL_E2E_OIDC=1 OPENSHELL_E2E_OIDC_SCOPES=1 \
  PYTHONPATH=python uv run pytest e2e/python/oidc/ -v

* fix(docs): fix markdown lint errors in OIDC architecture docs

Add blank lines before lists and fenced code blocks to satisfy
markdownlint MD031 and MD032 rules.
2026-04-30 10:37:23 -07:00
Seth Jennings 3b7d30934e feat(server): add Prometheus metrics infrastructure and gRPC/HTTP request metrics (#920)
* feat(server): add Prometheus metrics infrastructure and gRPC/HTTP request metrics

Add metrics exposition via a dedicated server port (--metrics-port,
default disabled) following the same optional-listener pattern as the
health port. The metrics crate facade records counters and histograms
in the MultiplexedService hot path, and a PrometheusHandle renders them
at GET /metrics on the dedicated port.

Metrics added:
- openshell_server_grpc_requests_total (counter: method, code)
- openshell_server_grpc_request_duration_seconds (histogram: method, code)
- openshell_server_http_requests_total (counter: path, status)
- openshell_server_http_request_duration_seconds (histogram: path, status)

* feat(helm): add metrics port to openshell-server helm chart

Expose the Prometheus metrics endpoint in the helm chart by adding
service.metricsPort (default 9090) to values, the --metrics-port arg
and container port to the statefulset, and a metrics service port.
Set metricsPort to 0 to disable.
2026-04-23 12:53:13 -07:00
Alexander Watson 7c314e7405 feat(prover): add native Rust policy prover with Z3 solver (#741)
* feat(prover): add native Rust policy prover with Z3 solver

Add openshell-prover crate implementing formal policy verification
using Z3 SMT solving. Answers two questions about any sandbox policy:
"Can data leave?" and "Can the agent write despite read-only intent?"

Native Rust — no Python subprocess, no PYTHONPATH, no uv dependency.
Z3 bundled via z3-sys for self-contained builds.

Replaces the Python prototype from #703.

Closes #699

Signed-off-by: Alexander Watson <zredlined@gmail.com>

* fix(prover): skip L7-write in exfil, write bypass only on read-only intent

Port two fixes from the Python branch:
- Exfil query skips endpoints where L7 is enforced and working
- Write bypass only fires on explicit read-only intent, not L4-only

Signed-off-by: Alexander Watson <zredlined@gmail.com>

* chore(prover): add missing SPDX license headers to registry and testdata YAML files

* fix(prover): revert serde_yaml to serde_yml to match workspace dependency on main

* fix(prover): apply cargo fmt formatting to prover and cli source files

* ci: add clang and libclang-dev to CI image for z3-sys bindgen

z3-sys requires libclang at build time for bindgen to generate FFI
bindings. Without it, the Rust CI jobs fail on the prover crate.

* fix(prover): use system libz3 instead of compiling from source

Add libz3-dev to the CI image and drop the z3 `bundled` feature from
the workspace dependency. This eliminates the ~30 min z3 C++ build
that ran on every CI cache miss.

sccache only wraps rustc — it does not intercept the cmake C++ build
that z3-sys runs when `bundled` is enabled. The cargo target cache
helped on warm runs but evicts on toolchain bumps and fork PRs.
With the system library pre-installed in the CI image, z3 link time
is always zero.

A `bundled-z3` opt-in feature is added to the prover crate for local
development without system z3:

  cargo build -p openshell-prover --features bundled-z3

Regular local dev: brew install z3 (macOS) or apt install libz3-dev
(Linux), then cargo build just works.

Signed-off-by: Alexander Watson <zredlined@gmail.com>
Made-with: Cursor

* fix(prover): use u16 for ports, align include_workdir default with runtime

Port fields changed from u32 to u16 across all prover types (policy,
model, finding, queries). Prevents the prover from silently accepting
port values >65535 that the runtime rejects, which would produce
misleading PASS results on invalid policies.

Change include_workdir serde default from true to false to match the
runtime (openshell-policy uses #[serde(default)] which gives false).
The previous mismatch caused the prover to model a /sandbox path that
does not exist at runtime, producing false analysis.

Signed-off-by: Alexander Watson <zredlined@gmail.com>
Made-with: Cursor

* refactor(prover): embed registry at compile time, replace hand-rolled glob

Embed binary and API capability registry YAML files into the binary at
compile time using include_dir!. The previous approach used
env!("CARGO_MANIFEST_DIR") which bakes in the build machine's source
path — works in tests but breaks for installed binaries. The CWD
fallback was equally fragile.

The --registry CLI flag still works as a filesystem override for custom
registries. Credentials remain filesystem-loaded (user-supplied data).

Replace the hand-rolled glob matching (~60 lines) with the glob crate's
Pattern::matches(), which is already a transitive dependency.

Remove unused dependencies: openshell-policy, openshell-core, thiserror,
tracing. The prover parses policy YAML directly into its own types and
does not use the policy or core crate APIs.

Signed-off-by: Alexander Watson <zredlined@gmail.com>
Made-with: Cursor

* docs: add Z3 system library to prerequisites

The prover crate now links against system libz3 instead of compiling
from source. Document the install steps for macOS, Ubuntu, and Fedora,
and note the bundled-z3 feature flag as a fallback.

Signed-off-by: Alexander Watson <zredlined@gmail.com>
Made-with: Cursor

* fix(prover): apply cargo fmt formatting

Made-with: Cursor

---------

Signed-off-by: Alexander Watson <zredlined@gmail.com>
2026-04-09 15:26:11 -07:00
John T. MyersandJohn Myers eea495e6b9 fix: remediate 9 security findings from external audit (OS-15 through OS-23) (#744)
* fix(install): restrict tar extraction to expected binary member

Prevents CWE-22 path traversal by extracting only the expected APP_NAME
member instead of the full archive contents. Adds --no-same-owner and
--no-same-permissions for defense-in-depth.

OS-20

* fix(deploy): quote registry credentials in YAML heredocs

Wraps username/password values with a yaml_quote helper to prevent YAML
injection from special characters in registry credentials (CWE-94).
Applied to all three heredoc blocks that emit registries.yaml auth.

OS-23

* fix(server): redact session token in SSH tunnel rate-limit log

Logs only the last 4 characters of bearer tokens to prevent credential
exposure in log aggregation systems (CWE-532).

OS-18

* fix(server): escape gateway_display in auth connect page

Applies html_escape() to the Host/X-Forwarded-Host header value before
rendering it into the HTML template, preventing HTML injection (CWE-79).

OS-17

* fix(server): prevent XSS via code param with validation and proper JS escaping

Adds server-side validation rejecting confirmation codes that do not
match the CLI-generated format, replaces manual JS string escaping with
serde_json serialization (handling U+2028/U+2029 line terminators), and
adds a Content-Security-Policy header with nonce-based script-src.

OS-16

* fix(sandbox): add byte cap and idle timeout to streaming inference relay

Prevents resource exhaustion from upstream inference endpoints that stream
indefinitely or hold connections open. Adds a 32 MiB total body limit
and 30-second per-chunk idle timeout (CWE-400).

OS-21

* fix(policy): narrow port field from u32 to u16 to reject invalid values

Prevents meaningless port values >65535 from being accepted in policy
YAML definitions. The proto field remains uint32 (protobuf has no u16)
with validation at the conversion boundary.

OS-22

* fix(deps): migrate from archived serde_yaml to serde_yml

Replaces serde_yaml 0.9 (archived, RUSTSEC-2024-0320) with serde_yml
0.0.12, a maintained API-compatible fork. All import sites updated
across openshell-policy, openshell-sandbox, and openshell-router.

OS-19

* fix(server): re-validate sandbox-submitted security_notes and cap hit_count

The gateway now re-runs security heuristics on proposed policy chunks
instead of trusting sandbox-provided security_notes, validates host
wildcards, caps hit_count at 100, and clamps confidence to [0,1]. The
TUI approve-all path is updated to use ApproveAllDraftChunks RPC which
respects the security_notes filtering gate (CWE-284, confused deputy).

OS-15

* chore: apply cargo fmt and update Cargo.lock for serde_yml

---------

Co-authored-by: John Myers <johntmyers@users.noreply.github.com>
2026-04-02 20:32:59 -07:00
Drew Newberry a912848217 refactor(build): unify image build graph for cache reuse (#390) 2026-03-18 15:01:04 -07:00
John T. Myers 764fac79aa feat(tui): add log copy and visual selection mode (#276)
* feat(tui): add log copy and visual selection mode

Add clipboard support for the TUI log viewer using OSC 52 escape
sequences. Users can copy individual log lines (y), the visible
viewport (Y), or enter visual selection mode (v) to select a range
of lines with j/k and yank with y.

- Add clipboard module with OSC 52 base64 encoding
- Add format_log_line_plain() for plain-text log rendering
- Add visual selection mode with highlight styling and status bar
- Context-switch nav bar hints between normal and visual modes
- Add 7 unit tests for plain-text log formatting

Closes #274

* fix(tui): write OSC 52 to /dev/tty instead of stdout

ratatui owns stdout via the alternate screen buffer, so OSC 52
escape sequences written to stdout are swallowed by the backend.
Write directly to /dev/tty to bypass the buffer and reach the
terminal emulator.
2026-03-13 09:34:54 -07:00
John T. Myers fcf12dff62 feat(tui): support light terminal backgrounds with adaptive theme (#265)
* feat(tui): support light terminal backgrounds with adaptive theme

Replace the hardcoded dark-only color constants in theme.rs with a
runtime Theme struct that has dark() and light() factory constructors.
The active theme is stored on App and threaded through all draw
functions, enabling the TUI to render correctly on both dark and light
terminal backgrounds.

- Add Theme struct with 16 semantic style fields and ThemeMode enum
- Detect terminal background via COLORFGBG env var at startup
- Add --theme dark|light|auto CLI flag and OPENSHELL_THEME env var
- Migrate all 499 styles:: references across 11 UI files to app.theme
- Clean up 9 inline Style constructions that bypassed the theme system
- Add unit tests verifying dark/light palettes and legacy regression

Closes #264

* fix(tui): use OSC 11 query for reliable light/dark terminal detection

Replace COLORFGBG env var heuristic with terminal-colorsaurus crate,
which sends an OSC 11 query to read the actual background color. This
fixes auto-detection in iTerm2 and other terminals that don't set
COLORFGBG.
2026-03-12 13:15:57 -07:00
Drew Newberry 984d1a6e5c chore: rename project from NemoClaw to OpenShell (#198) 2026-03-10 11:49:09 -07:00
Drew Newberry a2de1f24e5 feat: add Cloudflare tunnel auth support (#178) 2026-03-09 18:39:42 -07:00
Drew Newberry 4a78865b90 feat(ci): add CLI binary builds and snapshot release to publish workflow (#110) 2026-03-04 23:57:22 -08:00
Alexander Watson 1d7909cb38 chore: add open-source compliance files and SPDX headers (#71)
Add Apache 2.0 licensing, SPDX copyright headers on all source files,
DCO enforcement, third-party notices, and CI enforcement.

- LICENSE: Apache License 2.0 full text
- DCO: Developer Certificate of Origin 1.1
- SPDX headers on all 176 source files (.rs, .py, .proto, .rego, .sh,
  .toml, .yaml, Dockerfiles)
- scripts/update_license_headers.py: header management with --check mode
- scripts/generate_third_party_notices.py: dependency license aggregation
- THIRD-PARTY-NOTICES: generated listing of all Rust and Python deps
- build/license.toml: mise tasks for license:check and license:update
- CI: license-headers job in checks.yml, DCO check workflow
- CONTRIBUTING.md: DCO sign-off requirement and license header docs
- Cargo.toml: license changed to Apache-2.0, repository URL updated
- pyproject.toml: license field added

Closes #58
2026-03-03 09:30:56 -08:00
Piotr Mlocek 1c5051209c feat(cli): add dynamic shell completion support (!59)
> **🏗️ build-from-issue-agent**

Closes #71

## Summary
Add dynamic shell completion support using `clap_complete`'s `CompleteEnv`, plus a `navigator completions <shell>` subcommand (rustup-style) that outputs completion scripts to stdout with detailed per-shell setup instructions in `--help`.

Completions are dynamic — the shell calls back into the `navigator` binary on each tab-press, so they are always in sync with the installed version. This also lays the groundwork for future runtime completers (e.g., listing sandbox names on tab, tracked in #83).

## Changes Made
- `Cargo.toml`: Added `clap_complete = { version = "4.5", features = ["unstable-dynamic"] }` to workspace dependencies
- `crates/navigator-cli/Cargo.toml`: Added `clap_complete` dependency
- `crates/navigator-cli/src/main.rs`:
  - Added `CompleteEnv::with_factory(Cli::command).complete()` at the top of `main()` (before TLS/argument parsing)
  - Added `Completions { shell }` subcommand that re-invokes the binary with `COMPLETE=<shell>` and pipes the output
  - Added `CompletionShell` enum (bash, fish, zsh, powershell)
  - Added `COMPLETIONS_HELP` with per-shell setup instructions (similar to `rustup completions --help`)
  - Added 3 unit tests
- `CONTRIBUTING.md`: Added shell completions section documenting `nav` wrapper setup

## Usage

```bash
# Generate completions
navigator completions fish > ~/.config/fish/completions/navigator.fish
navigator completions bash > ~/.local/share/bash-completion/completions/navigator
navigator completions zsh > ~/.zfunc/_navigator
navigator completions powershell >> $PROFILE

# See full per-shell instructions
navigator completions --help

# For the nav wrapper (rewrite registration target)
navigator completions fish | sed 's/--command navigator/--command nav/' > ~/.config/fish/completions/nav.fish
```

## Design Decision: Why stdout-only (no auto-install)

We researched how 12 popular CLI tools handle completion installation:

| Tool | Command | Writes files? | Modifies RC? |
|------|---------|---------------|-------------|
| gh, glab, kubectl, docker, rustup, mise, starship | `completion <shell>` | stdout only | No |
| atuin | `gen-completions --shell <shell>` | stdout (`--out-dir` option) | No |
| gcloud | `install.sh` | Yes | **Yes** |
| Amazon Q | `q integrations install` | Yes | **Yes** |

**10 of 12 tools** follow the same pattern: output to stdout, never modify RC files, explicit shell selection, clear `--help` instructions. Only gcloud and Amazon Q modify RC files (and both have been criticized for it). Our approach matches the industry consensus.

## Deviations from Plan
- The original issue proposed a static `completions <shell>` subcommand using `clap_complete::generate()`. After discussion, we switched to dynamic `CompleteEnv` which is always up-to-date and enables future runtime completers.
- Added a `navigator completions <shell>` subcommand (not in original plan) for discoverability — it wraps the `COMPLETE` env var mechanism with a familiar CLI UX.

## Tests Added
- **Unit:** `cli_debug_assert` — validates CLI structure; `completions_engine_returns_candidates` — smoke test confirming the completion engine returns subcommand candidates; `completions_subcommand_appears_in_candidates` — verifies `completions` appears in tab-completion results
- **Integration:** N/A
- **E2E:** N/A

## Verification
- [x] All tests passing
- [x] Pre-commit checks passing (clippy warnings are pre-existing, not from this change)
- [x] CONTRIBUTING.md updated with nav wrapper completion setup
2026-02-26 10:33:00 -08:00
Piotr Mlocek f869182d8a feat(sandbox): move inference execution to sandbox-local routing (!79) (!52)
> **🏗️ build-from-issue-agent**

Closes #79

## Summary

Replaces the gateway-proxied inference model (`ProxyInference` RPC) with sandbox-local execution. The sandbox now resolves routes from either a standalone YAML route file or a cluster bundle fetched via the new `GetSandboxInferenceBundle` RPC, then forwards requests directly to inference backends using `navigator-router`.

### Benefits of removing the gRPC hop

Moving inference execution from the gateway to the sandbox eliminates the gRPC round-trip for every inference request:

- **No payload size limits**: Direct HTTP forwarding replaces protobuf-over-gRPC serialization.
- **Streaming-ready**: Preserves connection semantics for future SSE support (e.g., `stream: true`).
- **Lower latency**: Direct sandbox → backend path instead of sandbox → gateway → backend round-trip.
- **Reduced gateway load**: Gateway only serves the control-plane bundle delivery RPC.
- **Simpler error model**: HTTP status codes flow through directly without gRPC wrapping.

## Changes Made

- **Proto**: Removed `ProxyInference` RPC. Added `GetSandboxInferenceBundle` RPC for route bundle delivery.
- **Gateway**: Replaced `proxy_inference` with `get_sandbox_inference_bundle`. Removed `Router` and `navigator-router` dependency — gateway is now control-plane only for inference.
- **Router**: Switched config file format from TOML to YAML. Added `Debug` impl that redacts API keys.
- **Sandbox**: New `InferenceContext` with local `Router` + route cache. Routes loaded from `--inference-routes` file (standalone) or cluster bundle via gRPC, with 30s background refresh. New `--inference-routes` / `NAVIGATOR_INFERENCE_ROUTES` CLI arg.
- **Dev sandbox**: Added `inference-routes.yaml` at repo root with default NVIDIA NIM route. `mise run sandbox` now mounts it automatically and supports `-e` flag to forward env vars (e.g., `mise run sandbox`).
- **Example**: Added standalone example (`examples/inference/routes.yaml`) and rewrote README to cover both standalone and cluster workflows.
- **E2E tests**: Updated existing test names/docs. Added tests for Anthropic messages protocol and multi-route policy filtering.

## Tests Added

- **Unit (21 new):** `router_error_to_http` (6 variants), `load_from_file` YAML round-trip (5), `resolve_sandbox_inference_bundle` gRPC handler (6), `build_inference_context` route loading (4)
- **E2E (2 new):** Anthropic messages protocol routing, route filtering by `allowed_routes`
- **Total:** 188 unit tests passing across 3 crates

## Documentation Updated

- `architecture/gateway.md`: Removed router, updated Inference Service
- `architecture/sandbox.md`: Updated orchestration flow, InferenceContext, route loading
- `architecture/inference-routing.md`: Complete rewrite for sandbox-local architecture

## Verification

- [x] All unit tests passing (188 across 3 crates)
- [x] Pre-commit Rust checks passing (fmt, clippy, tests)
- [x] Architecture documentation updated
- [x] E2E tests updated with new test cases
- [x] Standalone example and documentation added
- [x] Default inference-routes.yaml with env passthrough for mise run sandbox
2026-02-25 09:57:42 -08:00
John Myers f1439727b9 feat(sandbox): L7 protocol-aware inspection with TLS termination (!29)
## Summary

Implements L7 (application-layer) protocol-aware policy enforcement for the sandbox proxy, enabling per-request allow/deny decisions based on HTTP method and path — not just host:port.

### Phase 1: HTTP/REST L7 Inspection (plaintext)
- New `l7/` module with provider trait, REST parser, and relay loop
- Parses HTTP/1.1 requests inside CONNECT tunnels, evaluates method+path against OPA/Rego policy
- Supports `access` presets (`full`, `read-only`, `read-write`) and explicit `rules` with glob path matching
- `enforcement` modes: `enforce` (deny with 403 JSON) vs `audit` (log only, forward traffic)
- Structured logging for every L7 decision (`L7_REQUEST` with protocol, action, target, decision)

### Phase 2: MITM TLS Termination for HTTPS
- Ephemeral CA generated at sandbox startup (`rcgen`)
- Per-hostname leaf certificate cache (256 cap) for dynamic cert presentation
- Client-side TLS termination (sandbox trusts ephemeral CA via `SSL_CERT_FILE`, `NODE_EXTRA_CA_CERTS`, etc.)
- Upstream TLS connection with `webpki-roots` verification
- Endpoints with `tls: terminate` get full L7 inspection; others remain passthrough

### Additional
- Benign TLS/connection errors (close_notify, handshake EOF, reset) downgraded from WARN to DEBUG
- Generic `AsyncRead + AsyncWrite` bounds throughout L7 module (works with both `TcpStream` and `TlsStream`)
- Proto schema extended with `protocol`, `tls`, `enforcement`, `access`, `rules` fields on `NetworkEndpoint`
- CLI `nav sandbox run` wired with `--rego-policy`/`--rego-data` flags and L7 policy display

### E2E Test Coverage
- 7 L4 tests: no-matching-policy deny, wildcard binary, binary-restricted, wrong port, cross-policy isolation, non-CONNECT 405, structured log fields
- 7 L7 TLS tests: full access allow, read-only deny POST, audit mode, explicit path rules, CA trust store injection, deny response JSON format, structured log fields

## Key Files

| Area | Files |
|------|-------|
| L7 core | `l7/mod.rs`, `l7/provider.rs`, `l7/relay.rs`, `l7/rest.rs` |
| TLS termination | `l7/tls.rs` |
| Proxy integration | `proxy.rs`, `lib.rs`, `main.rs` |
| Env/process | `process.rs`, `ssh.rs` |
| OPA policy | `opa.rs`, `dev-sandbox-policy.rego` |
| Proto schema | `proto/sandbox.proto` |
| CLI | `navigator-cli/src/run.rs` |
| E2E tests | `e2e/python/test_sandbox_policy.py` |

## Test plan

- [x] 67 unit tests pass (`cargo test --workspace`)
- [x] `cargo clippy --workspace --all-targets` clean
- [x] 16 e2e policy tests pass (`mise run test:e2e:sandbox`)
- [x] Manual smoke test: sandbox with `tls: terminate` on `api.anthropic.com:443`

Closes #25
2026-02-19 09:01:26 -08:00
Piotr Mlocek c094769a43 feat(server): add an inference router (!13)
Closes #3
2026-02-11 19:09:14 -08:00
Drew Newberry e34444adf9 feat(sandbox): add ssh connect to sandbox + build agent harness 2026-02-05 07:58:20 -08:00
John Myers b1b5d90275 fix(sandbox): dynamically create and chown read_write directories
The sandbox now automatically creates directories listed in the policy's
`read_write` section and sets ownership to the configured sandbox user/group.
This runs in the supervisor process before forking the child.

Previously, read_write directories had to be pre-created in the Dockerfile
with correct ownership. Now arbitrary paths can be added to the policy
and they will be prepared at runtime.

Also adds a hardcoded chown for /sandbox in Dockerfile.factory as a
belt-and-suspenders approach for the default case.
2026-02-04 13:10:01 -08:00
Drew Newberry d5d3c71e9c feat(sandboxes): initial kube sandbox impl 2026-02-03 22:13:20 -08:00
Drew Newberry 5a15de63df feat(server): add support for entity persistence 2026-02-02 18:18:04 -08:00