mirror of
https://github.com/NVIDIA/OpenShell.git
synced 2026-10-04 16:39:35 +08:00
codex/mxc-https-test-socket-owner
52
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
ab64e84bfc |
fix(core): enforce owner-only Windows ACLs on sensitive files and dirs (#3495)
* fix(core): enforce owner-only Windows ACLs on sensitive files and dirs set_dir_owner_only/set_file_owner_only were unconditional no-ops on Windows, so the CLI's mTLS client private key, OIDC/edge tokens, cached SSH keys, and the gateway's key-encryption key relied entirely on inherited NTFS ACLs with no OpenShell-applied restriction. Apply an owner-only DACL via SetEntriesInAclW/SetNamedSecurityInfoW with PROTECTED_DACL_SECURITY_INFORMATION to strip inherited ACEs, matching the 0700/0600 guarantee already provided on Unix. is_file_permissions_too_open now also works on Windows instead of being Unix-only, closing the detection gap alongside the prevention gap. Signed-off-by: Prashant Khodade <pkhodade@nvidia.com> (cherry picked from commit 71560e947f85819efbcddf70ddda94befab62b0b) * fix(core): treat a NULL DACL as too open in is_file_permissions_too_open has_foreign_trustee conflated a NULL DACL with an unreadable/invalid ACL and returned Some(false) (not too open) for both. Per the Win32 contract, a NULL DACL means the object grants full access to everyone -- the most permissive state possible -- so it must be flagged as too open. Split the null and invalid-ACL branches: null now returns Some(true), invalid ACL keeps the existing unreadable-ACL fallback (None, which the caller maps to false via unwrap_or). Adds a regression test that constructs a real NULL DACL via a SetNamedSecurityInfoW helper confined to the windows_acl module, consistent with the existing unsafe-FFI confinement in that module. Found by CodeRabbit review on MR !113. Signed-off-by: Prashant Khodade <pkhodade@nvidia.com> (cherry picked from commit 46e635a4ef1d6937cdb088f46aa85baee3d6ad28) * fix(core): close three false-negative gaps in the Windows ACL audit restrict_to_current_user() updated only the DACL, leaving a foreign owner's implicit WRITE_DAC right intact -- they could later replace the DACL we just set. Query OWNER_SECURITY_INFORMATION and take ownership in the same SetNamedSecurityInfoW call; if the caller can't (a genuinely foreign-owned object), the call now fails instead of silently leaving the object insecure. is_file_permissions_too_open() mapped every Win32 inspection failure (missing READ_CONTROL, an invalid ACL, a token-query failure) to "not too open" via unwrap_or(false). Fail closed instead: an inspection failure is a security false-negative risk, not a green light. has_foreign_trustee()'s ACE loop only recognized plain ACCESS_ALLOWED_ACE_TYPE and treated every other type as non-granting. Windows also defines access-allowed object, callback, and callback-object ACE variants that can grant rights to a foreign trustee; this audit doesn't parse their wider layouts, so their mere presence is now conservatively flagged as too open instead of silently skipped. Also updates architecture/gateway.md, which still described the SQLite file-tightening behavior only in terms of Unix mode 0o600, to distinguish it from the owner-only DACL behavior on Windows. Addresses review comments on PR #3495. Signed-off-by: Prashant Khodade <pkhodade@nvidia.com> * fix(core): conditional owner claim and audit owner in Windows ACL helpers restrict_to_current_user: query the current owner before calling SetNamedSecurityInfoW. Include OWNER_SECURITY_INFORMATION only when the path has a foreign owner -- requesting it unconditionally fails with ACCESS_DENIED (0x80070005) on standard credentials even when the current user is already the owner, because WRITE_OWNER is not implied by object ownership. A foreign-owned path still triggers an ownership claim and fails hard if the claim is denied, preserving the security contract. has_foreign_trustee: request OWNER_SECURITY_INFORMATION alongside DACL_SECURITY_INFORMATION and reject paths with a foreign owner immediately, before inspecting the DACL. A foreign owner has implicit WRITE_DAC rights and can replace any DACL we set, so a clean DACL is not sufficient evidence of safety on a foreign-owned object. architecture/gateway.md: clarify that the Windows path-hardening behavior sets mode 0o600 on Unix and applies a protected owner-only DACL on Windows, with conditional ownership claim and fail-hard semantics for foreign-owned objects. Signed-off-by: Prashant Khodade <pkhodade@nvidia.com> --------- Signed-off-by: Prashant Khodade <pkhodade@nvidia.com> |
||
|
|
2d5c9bfaab |
fix(network): scope Windows egress by socket owner (#3491)
* fix(ocsf): attribute MXC proxy events to sandboxes NVBug 6783086 Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com> * fix(network): scope Windows egress by socket owner Resolve each accepted Windows proxy connection to its unique owning PID and executable before evaluating binary-scoped network policy. Fail closed when ownership or process identity cannot be established, and cover allowed and undeclared child processes with a real MXC regression. * fix(network): preserve sandbox context with socket owners Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com> (cherry picked from commit 32ea316deb6c83c157fdf2de4a2cdc3a7d2da1e3) * fix(network): harden Windows process identity Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com> --------- Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com> |
||
|
|
b2d6280794 | fix(windows): repair OpenClaw ProcessContainer startup (NVBug 6782898) (#3475) | ||
|
|
39cf4823f7 |
feat(api): add structured gateway errors and SDK decoding (#3313)
* feat(api): expose structured gateway errors across SDKs Refs #3051. Add standard validation, conflict, and retry details; preserve raw transport status in Rust, Go, TypeScript, and Python; document status and recovery guidance. This is the structured-error foundation only. Mutation result shapes, allow_missing, durable request deduplication, and exec retry semantics remain follow-up work. Signed-off-by: Mrunal Patel <mrunalp@gmail.com> * fix(python): preserve wrapped RPC cleanup handling Inspect the original gRPC call when handling missing sandboxes during deletion waits and managed cleanup. Add intercepted cleanup regressions and clarify the error-wrapper migration contract. Addresses the cleanup review on #3313; part of #3051. Signed-off-by: Mrunal Patel <mrunalp@gmail.com> --------- Signed-off-by: Mrunal Patel <mrunalp@gmail.com> |
||
|
|
ae57979b03 |
feat(mxc): add Windows ETW-to-OCSF audit trail (#3015)
* feat(mxc): ETW->OCSF audit consumer + Windows OCSF JSONL parity (cp6 P1) Add a Windows MXC ETW->OCSF audit trail in openshell-driver-mxc: a real-time Sandboxing-provider ETW consumer that decodes events (TDH), attributes each to an OpenShell sandbox_id, and maps them to OCSF (lifecycle 6002, config 5019, process 1007, finding 2004). cp6 Phase 1 - durable OCSF JSONL audit-file parity with Linux: - openshell-ocsf: add emit_ocsf_event_routed (populates the event-bridge thread-local AND stamps sandbox_id+message in one dispatch) plus public set/clear_current_event; OS-aware device (Device::windows/for_current_os) so device.os.name reflects the host instead of a hardcoded Linux stub. - etw_consumer: emit via the routed emit (previously fired a bare info! that never populated the bridge, so the structured event was dropped). - openshell-server: install OcsfJsonlLayer over a synchronous daily-rotated appender (durable under force-kill), gated by OPENSHELL_OCSF_JSON, path via %PROGRAMDATA%\OpenShell\logs (override OPENSHELL_OCSF_LOG_DIR). - device.hostname now resolves to the real gateway machine name. Box-proven on 7F203-MXC-001: JSONL lines == shorthand OCSF rows, all valid OCSF JSON, per-sandbox attribution intact, disabled state writes nothing. Signed-off-by: Akber Raza <akberr@nvidia.com> * feat(mxc): map remaining Sandboxing ETW events to OCSF Close the last three ETW->OCSF gaps so the audit trail covers the full set of events the Sandboxing provider emits (12/12): - ProcessLaunched -> Process Activity [1007] "Launch" (confirmed start; carries the real processId/threadId, the twin of CreateProcessInSandbox which only has the request + command line). - SandboxProxyConfigured -> Device Config State Change [5019] (the one network-plane setup event; surfaces proxyPort, "no proxy" when 0). - SandboxConsoleReferencePlumbed -> Device Config State Change [5019] (console-handle plumbing). map_config_state now handles the full config/hardening/setup family and carries proxyPort/hasConsoleReference/creationFlags as unmapped fields. Verified on 7F203-MXC-001: 11/12 event types emit OCSF without a proxy (SandboxProxyConfigured requires proxy config to fire). Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(mxc): seed ETW attribution under registry lock + Device tests Address CodeRabbit review on !31: - Prevent stale ETW attribution on a delete/launch race: register the wxc-exec pid while holding the registry lock, and bail if the sandbox entry is already gone. Previously the attribution key could be seeded after `delete` had removed the sandbox, leaving a stale key that could misroute later Sandboxing ETW events to a dead sandbox_id. Lock order (registry -> attribution) matches the delete path, so no deadlock. - Add unit tests for the new Device::windows and Device::for_current_os constructors to harden Windows/Linux OCSF device parity. Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(mxc-etw): buffer+replay racing events and harden attribution keys Addresses two ETW->OCSF attribution review items (Shailendra #1, #2). #2 early-event loss: ETW delivers the sandbox create/config burst the instant wxc-exec starts, which can beat the driver's register_launch (now under the registry lock post-Ready). process_event previously dropped anything unresolved, losing the racing burst. Add a bounded, time-bounded pending buffer (PENDING_MAX=4096, PENDING_TTL=5s): unresolved events are held and replayed once attribution lands, aged-out ones dropped. Consumer switched to a timed recv_timeout(200ms) so the buffer is re-driven after each event and on a tick. Emit path factored into shared emit_resolved(). #1 attribution collisions: a Windows PID is recycled after exit and a command line is commonly identical across sandboxes. register_launch now rebinds by_pid on reuse and clears the stale last_pid_sid hint (warns if the PID still pointed at a different, leaked sandbox); command line is held in by_cmd only while unique and demoted to a new ambiguous_cmds set on a second owner, so a duplicate command refuses to resolve rather than misroute. Unit tests: buffer replay (direct + cross-link), buffer bound, PID-reuse rebind, duplicate-cmd non-resolution. Box-verified on 7F203-MXC-001 (5 sandboxes, identical cmd -> 5 isolated sandbox_ids, 50/50 OCSF/JSONL, BuffersLost=0). Signed-off-by: Akber Raza <akberr@nvidia.com> * docs(mxc-etw): note cmd_line is captured raw with no privacy filtering Review item #3 (Shailendra): add a PRIVACY NOTE on map_process_launch stating cmd_line is copied verbatim into OCSF process.cmd_line with no redaction, so secrets/PII on a command line land unredacted in the durable audit trail (deliberate audit-fidelity trade-off; treat the log as sensitive). Redaction is owned by an upstream privacy layer, not this path; no general audit-output PII scrubber exists today (openshell_core::secrets [CREDENTIAL] redaction is scoped to the proxy HTTP-target logging, a separate egress path). Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(mxc-etw): open ETW trace on caller thread so start_session reports real status Review item #4 (Shailendra): start_session previously returned Ok(EtwSession) as soon as the pump thread was spawned, but OpenTraceW ran later inside that thread; if it failed we still handed back a live-looking session and logged 'consumer started' (silent failure = false audit coverage). Split the two Win32 calls instead of adding a channel handshake (avoids any lost-wakeup/hang risk): the quick, synchronous OpenTraceW now runs on the caller thread (open_trace), and only the blocking ProcessTrace runs on the pump thread (run_trace). start_session returns Err if OpenTraceW fails (reclaiming the boxed Sender so the consumer disconnects, stopping the session, joining the consumer) and returns Ok/logs 'started' only once capture is genuinely open. Opened handle + LoggerName buffer + boxed Sender are carried to the pump via a Send OpenedTrace so they outlive ProcessTrace. Box-verified on 7F203-MXC-001: consumer started=True, failed-to-start=False, 50 OCSF rows / 50 JSONL, BuffersLost=0 (no regression to capture/emit). Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(mxc-etw): guard pending-event replay against PID recycling CodeRabbit flagged that drain_resolved() re-resolved buffered events against the live by_pid map, so if Windows recycled a wxc-exec PID within PENDING_TTL a stale event from the dead sandbox could be emitted under the new owner. Stamp each by_pid registration with its Instant and add resolve_replay(), used only on the buffered/replay path. It (a) never falls back to the recycle-/ambiguity-prone by_cmd or last_pid_sid keys, and (b) trusts a PID match only when the registration is not newer than the buffered event by more than REPLAY_PID_GRACE (2s) - a recycled PID's registration lands well outside that window, so the stale event ages out instead of misattributing. The legitimate #2 seed race (registration lands ~immediately) still replays. Adds unit tests for the recycle-refusal, in-grace acceptance, and weak-fallback exclusion. Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(mxc-etw): surface unexpected ProcessTrace termination (review #4) start_session already returns Err on OpenTraceW failure (runs on the caller thread since e41a7701), closing the first half of Shailendra's #4. This closes the second half: ProcessTrace's result was discarded, so if capture died mid-run the backend had no way to know. Add a shared CaptureHealth (stopped/stopping/exit_code) between the pump thread and EtwSession. run_trace now records ProcessTrace's WIN32_ERROR and, when the pump returns without a deliberate stop, logs at ERROR that MXC OCSF capture is no longer running. EtwSession::stop() sets `stopping` before teardown so a normal shutdown isn't misreported, and EtwSession::is_capture_alive() exposes the state for status/diagnostics. Box-verified on 7F203-MXC-001: 5 sandboxes, 50 attributed OCSF rows, JSONL parity 50/50, BuffersLost=0, clean start/stop (no false failure). Signed-off-by: Akber Raza <akberr@nvidia.com> * feat(mxc-ocsf): add ETW->OCSF audit-trail example kit; fix proxy-configured message Add a runnable OCSF audit-trail example under examples/ (run-ocsf-audit.ps1, mxc-ocsf-audit.toml, ocsf-audit.yaml, README) that spins up sandboxes with the in-process ETW consumer and egress proxy on, emitting a full OCSF JSONL audit trail across all four classes (6002/5019/1007/2004). Fix SandboxProxyConfigured mapping to log "MXC sandbox proxy configured" instead of a misleading "(no proxy)" when the provider reports proxyPort=0; the event's presence already indicates proxy configuration. Verified on-box: 26 events, all mapped ETW event types present. Signed-off-by: Akber Raza <akberr@nvidia.com> * feat(mxc-ocsf): clearer audit report + client-safe run-ocsf-audit.ps1 Improve the ETW to OCSF audit-trail example output and make it safe to ship. Report: - Add an event-type coverage count ("N of M expected event types fired"); the denominator auto-adjusts (8 with proxy on, 7 with -NoProxy). - Split the checklist into expected event types vs anomaly findings (ActivityError/FallbackError), which are reported separately and not counted toward coverage (a clean run may emit none). - Verdict is now coverage-based (all expected types must fire) instead of the looser "at least 3 OCSF classes". - Call out the absolute path to the durable OCSF JSONL log prominently. Client-safety: - Default -ShareOut to empty (no auto-copy); pass -ShareOut a UNC path to opt in. Removes a hardcoded internal share path from a published example. - Drop internal-team wording ("Hand that zip back for evaluation", "BUNDLE:") in favor of neutral "Results bundle:". - Update README-ocsf-audit.txt to match the opt-in -ShareOut behavior. Verified on both MXC boxes: 7F203-MXC-001 (base-container) -> PASS, 8 of 8 event types, 26 OCSF events across 4 classes; 7F203-MXC-003 (AppContainer fallback) -> reduced set as expected, clean output. Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(mxc): configure OCSF audit workloads per sandbox - remove unsupported gateway-scoped workload fields from the shipped MXC audit example. - build the command, working directory, and filesystem grant from each run's ShareDir - pass the workload through --driver-config-json. - preserve the host CONNECT proxy configuration and conditional audit coverage for the future host_connect_proxy merge - require the workload output when determining the audit verdict. Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(mxc): omit command arguments from OCSF audit logs - record only the executable basename for MXC CreateProcessInSandbox audit events - leave process.cmd_line unset so workload arguments cannot reach shorthand or JSONL logs - cover tokens, passwords, signed URLs, and PII with a secret-leak regression test - update the audit example, architecture guidance, and published logging documentation - preserve ETW attribution and future host_connect_proxy enforcement behavior Signed-off-by: Akber Raza <akberr@nvidia.com> * feat(etw): enhance ETW session management with distinct naming for concurrent gateways * fix(etw): bound the audit queue during overload - replace the unbounded ETW callback channel with count- and byte-bounded buffering - keep the ETW callback non-blocking and count records rejected during overload - emit immediate, rate-limited warnings that identify resulting audit coverage gaps - make the audit example fail when queue overload causes dropped ETW records - cover stalled consumers, oversized events, and warning throttling with unit tests Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(etw): harden sandbox audit attribution - remove command-line and persistent per-PID fallback keys from live and replay resolution - retire the driver-owned wxc-exec PID before publishing child completion - retain established identity, activity, and correlation-vector links only for the five-second late-event window - prevent buffered records from crossing rapid PID retirement and reuse boundaries - add resolver and lifecycle coverage and document the attribution trust boundary Signed-off-by: Akber Raza <akberr@nvidia.com> * chore(mxc): address rebase follow-ups Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(mxc): align OCSF audit example with driver config - remove unsupported egress proxy settings - stop requiring the unavailable proxy audit event - update example documentation for supported event coverage Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(etw): redact command-line secrets in DecodedEtwEvent summary * fix(etw): enhance PID resolution and event attribution logic for ETW records * fix(ocsf): restrict gateway-local JSONL sink to Windows/MXC path with opt-in configuration * address rebase issues * fix(mxc): address ETW audit review feedback Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(mxc): fail closed across ambiguous PID reuse Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(mxc): bind ETW attribution to process generation Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Akber Raza <akberr@nvidia.com> Signed-off-by: Drew Newberry <anewberry@nvidia.com> Co-authored-by: Jamie King <jamiek@nvidia.com> Co-authored-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
48c449d8c8 |
chore(deps): replace ring with AWS-LC (#3243)
* chore(deps): replace ring with AWS-LC Signed-off-by: Simon Scatton <sscatton@nvidia.com> * fix(lint): address warnings after dependency upgrades Signed-off-by: Simon Scatton <sscatton@nvidia.com> * fix(tls): limit provider initialization to reqwest clients Signed-off-by: Simon Scatton <sscatton@nvidia.com> --------- Signed-off-by: Simon Scatton <sscatton@nvidia.com> |
||
|
|
e4369adcd0 |
chore(deps): replace serde_yml with noyalib (#3031)
* chore(deps): replace serde_yml with noyalib Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(providers): annotate generic YAML test values Signed-off-by: Evan Lezar <elezar@nvidia.com> --------- Signed-off-by: Evan Lezar <elezar@nvidia.com> |
||
|
|
8a13bc1298 |
chore(deps): remove legacy rustls webpki path (#3013)
Signed-off-by: Evan Lezar <elezar@nvidia.com> |
||
|
|
40d1b48666 |
feat(provider): support for SPIFFE backed token exchange (#1970)
* feat(provider): add ability to request token exchange instead of client credentials as OAuth grant_type Signed-off-by: Gordon Sim <gsim@redhat.com> * test(proxy): add further tests for token exchange Signed-off-by: Gordon Sim <gsim@redhat.com> * test(provider): add runnable example for token exchange Signed-off-by: Gordon Sim <gsim@redhat.com> * test(e2e): cover Podman token exchange grants Signed-off-by: Gordon Sim <gsim@redhat.com> * refactor(oauth): extract duplicated functionality from server and supervisor Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(provider): evict nearest-to-expiry entry from intermediate token cache Signed-off-by: Gordon Sim <gsim@redhat.com> * doc(supervisor): add podman example for token exchange Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(provider): withhold token-exchange subject credentials Signed-off-by: Gordon Sim <gsim@redhat.com> --------- Signed-off-by: Gordon Sim <gsim@redhat.com> |
||
|
|
4d7f402ce2 |
feat(policy): establish direct TCP egress foundation (#2711)
* feat(policy): accept explicit tcp endpoint protocol Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * refactor(network): snapshot authoritative egress decisions Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * docs(policy): document explicit tcp protocol Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * docs(policy): defer transparent TCP release guidance Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * chore(go): regenerate sandbox protobuf bindings Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(network): complete tcp egress foundation Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * docs(policy): document explicit tcp contract Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(network): fail closed on authorization errors Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(podman): fence delayed exit events before restart Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(policy): validate network endpoint destinations Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(providers): opt in tcp credential fixture Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(policy): require dns host for transparent tcp Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> --------- Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> |
||
|
|
2f96c53b8c |
feat(gateway,cli): windows compilation support (#2496)
* chore(windows): gate Unix-only workspace code for MSVC Signed-off-by: Shailendra Singh <shailendras@nvidia.com> Signed-off-by: Akber Raza <akberr@nvidia.com> * feat(windows): stub unsupported compute drivers Signed-off-by: Shailendra Singh <shailendras@nvidia.com> Signed-off-by: Akber Raza <akberr@nvidia.com> * ci(windows): add MSVC mise build lane Signed-off-by: Shailendra Singh <shailendras@nvidia.com> Signed-off-by: Akber Raza <akberr@nvidia.com> * docs(windows): document MSVC build-only design Signed-off-by: Shailendra Singh <shailendras@nvidia.com> Signed-off-by: Akber Raza <akberr@nvidia.com> * docs(agent): add Windows MSVC build skill Signed-off-by: Shailendra Singh <shailendras@nvidia.com> Signed-off-by: Akber Raza <akberr@nvidia.com> * feat(windows): add Windows build support Signed-off-by: Akber Raza <akberr@nvidia.com> * refactor(windows): consolidate Windows-specific dependencies and improve build logic Signed-off-by: Akber Raza <akberr@nvidia.com> * feat(windows): add libclang path resolution and update cargo commands with bundled Z3 features Signed-off-by: Akber Raza <akberr@nvidia.com> * chore(tooling): lock Windows tool artifacts Signed-off-by: Giedrius Burachas <gburachas@nvidia.com> Signed-off-by: Akber Raza <akberr@nvidia.com> * feat(windows): enhance libclang path resolution to support architecture-specific subdirectories Signed-off-by: Akber Raza <akberr@nvidia.com> * Fix Windows dependency gating after sync merge Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(z3): update Z3 header path requirements in Windows build documentation and scripts Signed-off-by: Akber Raza <akberr@nvidia.com> * docs(windows): relocate Windows MSVC build design to architecture/ Why: windows-msvc-build-design.mdx is a design document ("design decisions for the native Windows MSVC build lane"), but it lived in the published, user-facing docs/reference/ tree. Per AGENTS.md (Documentation) and architecture/README.md ("rfc/ vs architecture/"), design content belongs in architecture/ (or rfc/), not in published reference. It also shared Fern sidebar "position: 6" with the MXC compute-driver design page, colliding in the Reference nav ordering. What: - Move docs/reference/windows-msvc-build-design.mdx -> architecture/windows-msvc-build.md. - Strip the Fern publish frontmatter and add a plain H1, matching the other architecture docs. - Register it in the architecture doc index in architecture/README.md. - Repoint the inbound references (build-openshell-mxc-windows skill + reference, implement-openshell-mxc-driver skill) to the new path. With both design pages moved out of docs/reference/, the duplicate position-6 sidebar collision is resolved. Signed-off-by: Akber Raza <akberr@nvidia.com> * remove openshell-supervisor-network from unsupported driver package test exclusion list Signed-off-by: Akber Raza <akberr@nvidia.com> # Conflicts: # tasks/scripts/windows-msvc.ps1 * fix(interceptors): gate unix-only imports so the crate builds on Windows openshell-gateway-interceptors failed to compile on Windows (E0432: no UnixStream in tokio::net), breaking any Windows build of openshell-server (which depends on it unconditionally). The connect_unix_endpoint fn was already #[cfg(unix)]-gated, but the imports it uses (UnixStream, TokioIo, Uri, service_fn) were left ungated. Gate those four imports with #[cfg(unix)] too. No behavior change on unix; Windows now compiles (no errors, no unused-import warnings). Signed-off-by: Akber Raza <akberr@nvidia.com> * feat(windows): add native ARM64 test support Signed-off-by: Shailendra Singh <shailendras@nvidia.com> Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(mise): skip Skaffold on Windows Signed-off-by: Shailendra Singh <shailendras@nvidia.com> Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(windows): harden ARM64 toolchain discovery Signed-off-by: Shailendra Singh <shailendras@nvidia.com> Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(windows): scope ARM64 toolchain preflight Signed-off-by: Shailendra Singh <shailendras@nvidia.com> Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(windows): restore compatibility after GitHub sync Signed-off-by: Shailendra Singh <shailendras@nvidia.com> Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(windows): avoid rate-limited Z3 source lookup Signed-off-by: Shailendra Singh <shailendras@nvidia.com> Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(mise): skip Helm checks on Windows Signed-off-by: Shailendra Singh <shailendras@nvidia.com> Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(windows): support repository pre-commit checks Signed-off-by: Shailendra Singh <shailendras@nvidia.com> Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(windows): stabilize native MSVC validation Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(windows): harden shared Z3 source cache Signed-off-by: Shailendra Singh <shailendras@nvidia.com> * fix(windows): avoid leaking MSVC flags into clang-cl Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(windows): complete ARM64 migration audit Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(windows): restore ARM64 Ninja discovery Signed-off-by: Akber Raza <akberr@nvidia.com> * refactor(windows): separate platform crate roots Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(windows): restore proto include cfg gating Signed-off-by: Akber Raza <akberr@nvidia.com> * refactor: address lint errors * fix(windows): add preflight check for proxy auth file path * docs(windows): update GitHub checkout guidance Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(windows): restore CI after dependency updates Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(mise): repair Windows sccache lock entry Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(windows): reconcile validation after rebase Signed-off-by: Akber Raza <akberr@nvidia.com> * refactor(server): exclude unsupported drivers on Windows Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * refactor(server): isolate platform driver config Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(windows): repair unsupported driver contract test Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(sandbox): remove stale dependencies Signed-off-by: Akber Raza <akberr@nvidia.com> * ci(windows): pin x64 workflow actions Signed-off-by: Akber Raza <akberr@nvidia.com> * ci(windows): align x64 Rust toolchain Signed-off-by: Akber Raza <akberr@nvidia.com> * ci(windows): align ARM64 workflow setup Signed-off-by: Akber Raza <akberr@nvidia.com> * refactor(windows): exclude unsupported runtime crates Signed-off-by: Akber Raza <akberr@nvidia.com> * refactor(windows): exclude unsupported crates at workspace boundary Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com> * refactor(server): gate builtin driver config by platform Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com> * fix(sandbox): restore crate documentation Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com> * ci(windows): make build workflow manual Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com> * ci(windows): temporarily enable pull request builds Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com> * ci(windows): cache Rust dependencies Signed-off-by: Akber Raza <akberr@nvidia.com> * refactor(windows): remove unnecessary platform changes Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com> * ci(windows): make build workflow manual Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com> * fix(ci): synchronize mise lockfile Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com> * fix(ci): normalize mise provenance metadata Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com> * refactor(python): isolate Windows atomic replace retry Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com> * fix(python): type Windows permission test errors Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com> --------- Signed-off-by: Shailendra Singh <shailendras@nvidia.com> Signed-off-by: Akber Raza <akberr@nvidia.com> Signed-off-by: Giedrius Burachas <gburachas@nvidia.com> Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com> Co-authored-by: Shailendra Singh <shailendras@nvidia.com> Co-authored-by: Giedrius Burachas <gburachas@nvidia.com> Co-authored-by: Jamie King <jamiek@nvidia.com> Co-authored-by: Piotr Mlocek <pmlocek@nvidia.com> Co-authored-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com> |
||
|
|
0c7e59a953 |
fix(deps): bump russh, jsonwebtoken, tar and npm lint deps (#2617)
Signed-off-by: Adrien Langou <alangou@nvidia.com> |
||
|
|
5548405fcb |
feat(credentials): add provider credential storage drivers (#2437)
* feat(credentials): add provider credential storage drivers Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(credentials): harden credential update handling Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(credentials): harden credential driver security, correctness, and performance Address review findings from the credential storage drivers PR: - Route additional_credentials through the driver on refresh to prevent silent data loss for multi-credential providers (e.g. AWS STS) - Clean up stored credential handles on CAS failure during refresh to prevent orphaned secrets in external backends - Enforce namespace validation in the Kubernetes Secrets driver to prevent cross-namespace credential access when allow_reference_namespace is not enabled - Cache Vault Kubernetes auth tokens with 80% TTL to avoid re-authenticating on every credential operation - Parallelize resolve_credentials in all three drivers using try_join_all for faster sandbox startup - Add existingSecret support for the KEK Secret to fix helm template/GitOps workflows where lookup returns empty and regenerates the key - Document RBAC blast radius for the Kubernetes Secrets credential driver and recommend a dedicated namespace Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com> Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> * fix(credentials): add optimistic concurrency, fix thundering herd, parallelize operations Use resourceVersion optimistic concurrency with retry loop for K8s Secret ownership checks to prevent TOCTOU races. Switch Vault token cache from RwLock to Mutex with double-check pattern to prevent thundering herd on cache miss. Parallelize credential store and delete operations across independent keys using try_join_all. Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com> Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> * fix(credentials): handle partial failures, add delete retry, consolidate cleanup Replace try_join_all with join_all in credential store/delete operations to handle partial failures — successfully-stored handles are cleaned up when another key fails. Add retry loop with conflict detection to db-credstore delete_credential, matching the K8s driver pattern. Consolidate 4 manual cleanup_pre_stored_provider_credentials call sites into a single error handler using an async block. Remove inconsistent .trim() from db-credstore validate_handle_owner. Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com> Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> * fix(credentials): fix retry loop guard and remove unprotected validation Remove attempt-count guard from 409/Aborted match arms in retry loops so the post-loop Status::aborted error is reachable after exhausting retries. Previously, last-attempt conflicts fell through to the catch-all error arm, producing misleading Status::unavailable errors. Remove duplicate validation calls that ran after prepare_provider_credential_update but outside the cleanup-protected async block, which would leak pre-stored handles on failure. Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com> Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> * fix(credentials): add workspace/provider UUID to credential backend paths Include workspace and provider ID in credential backend object paths to ensure cross-workspace uniqueness and prevent credential collision (GATOR-1806c9be-01). - Updated credential driver proto to include workspace and provider_id fields - Modified Vault driver to include workspace/provider_id in managed_secret_path - Modified Kubernetes Secrets driver to include workspace/provider_id in credential_owner_id and managed_secret_name - Updated all credential runtime calls to pass workspace/provider_id - Updated tests to use the new signatures This prevents two workspaces sharing the same external credential store from colliding on provider names, which was a critical security issue (CWE-639). * fix(credentials): preserve provider-level expiration for handle-backed credentials Compute effective expiration from both provider and driver values using the earliest non-zero timestamp and skip expired values before insertion (GATOR-1806c9be-02). - Modified resolve_provider_handles to check provider credential_expires_at_ms - Skip expired credentials during resolution instead of returning them - Use effective expiration (min of provider and driver) in resolution results - Fix inference.rs to preserve earliest expiration when merging This ensures handle-backed credentials respect the same expiration semantics as inline credentials. * fix(credentials): stage refresh changes under new handles before validation Stage credential replacements under new immutable handles instead of reusing existing handles to prevent overwriting committed values before validation/CAS (GATOR-1806c9be-03). - Stage credentials with empty existing_handles map to force new handle creation - Validate and CAS before the new values are committed to backend storage - Delete old handles only after successful CAS - On CAS failure, delete only the newly staged handles - This prevents CWE-362/CWE-367 race conditions where failed refreshes could still modify or delete the active credential The fix ensures that a rejected refresh cannot modify the backend object still referenced by the committed provider record. * fix(credentials): add timeouts to credential driver RPCs Apply configured timeouts to both startup capability negotiation and runtime RPCs to prevent indefinite hangs (GATOR-1806c9be-05). - Add DEFAULT_CREDENTIAL_DRIVER_RPC_TIMEOUT_SECS constant (30s) - Apply timeout to GetCapabilities during startup connection - Apply timeout to all runtime RPCs (store, delete, resolve) - Use tokio::time::timeout to bound the entire GetCapabilities operation during startup, not just the socket connection - Return contextual deadline errors on timeout This prevents a faulty or overloaded driver from hanging gateway operations indefinitely. * fix(credentials): fix test to use consistent workspace/provider identity The Kubernetes auth Vault resolve test was constructing a managed path with test-workspace/test-provider-id but sending default/prov-123 in the request, causing validation to reject the request (GATOR-18e32351-01). - Update test to use test-workspace and test-provider-id in the request to match the logical_path construction - This ensures the test exercises the intended code path and validates Kubernetes auth resolution properly The test now passes and correctly validates identity enforcement. * fix(credentials): use unique staging ID for refresh to avoid overwrites Stage refresh replacements under genuinely distinct immutable handles using a unique staging ID to prevent overwriting committed values (GATOR-1806c9be-03). - Generate a unique staging ID using UUID for each refresh operation - Use this staging ID when storing credentials instead of the real provider ID - Pass the same staging ID during cleanup on failure to delete only staged objects - This ensures deterministic paths (Vault) and object names (K8s) don't collide with the committed provider's credentials The fix prevents failed refreshes from silently replacing active credentials or breaking providers by deleting still-referenced backend objects. * fix(credentials): wrap credential driver RPCs in local timeouts Add local tokio::time::timeout wrappers around credential driver RPCs to bound non-compliant or stalled UDS peers (GATOR-1806c9be-05). - Wrap StoreCredential, DeleteCredential, and ResolveCredentials in local timeouts - Return contextual deadline_exceeded errors when timeouts occur - Keep existing gRPC timeout metadata for compliant implementations - GetCapabilities during startup was already wrapped in previous commit This ensures a faulty local driver cannot hang gateway operations indefinitely, even if it accepts the connection but never responds to the RPC. * fix(credentials): preserve ownership for staged refreshes Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(credentials): bound startup capability probe Signed-off-by: Seth Jennings <sjenning@redhat.com> * test(provider): authenticate credential handler requests Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(ci): grant actions read to credential driver e2e Signed-off-by: Seth Jennings <sjenning@redhat.com> --------- Signed-off-by: Taylor Mutch <taylormutch@gmail.com> Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com> Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> Signed-off-by: Seth Jennings <sjenning@redhat.com> Co-authored-by: Taylor Mutch <taylormutch@gmail.com> Co-authored-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> |
||
|
|
0e9a44cfa9 |
feat(build): add system CA root mode (#2324)
* feat(build): add system CA root mode Allow distro builds to use native trust stores for supervisor upstream TLS while keeping bundled Mozilla roots as the default. Avoid bundled root crates in system-ca-roots builds by using native-root TLS features and z3 0.20. Signed-off-by: Adam Miller <admiller@redhat.com> * fix(build): keep CA root feature in telemetry-off verification The telemetry-off task uses --no-default-features which now disables bundled-ca-roots in addition to telemetry, triggering the compile_error guard. Re-enable bundled-ca-roots explicitly so the task verifies only telemetry compilation. Signed-off-by: Scott Burdine <sburdine@nvidia.com> Signed-off-by: politerealism <burdcat17@gmail.com> * fix(sdk): disable oauth2 default features to prevent webpki-roots leak The bare `oauth2 = "5"` dependency re-enabled default features (rustls-tls → reqwest/rustls-tls → webpki-roots), defeating the system-ca-roots feature gate. Mirror the CLI fix: disable defaults and enable only the `reqwest` feature. Signed-off-by: Quinn Burdine <sburdine@redhat.com> Signed-off-by: politerealism <burdcat17@gmail.com> * refactor(build): simplify CA root selection to single feature toggle Replace mutually exclusive bundled-ca-roots / system-ca-roots features with a single bundled-ca-roots toggle. Disabling it implies system roots via rustls-native-certs, which is now a regular (non-optional) dependency. This fixes cargo --all-features and simplifies the distro build interface from --no-default-features --features system-ca-roots to just --no-default-features. Signed-off-by: Quinn Burdine <sburdine@redhat.com> Signed-off-by: politerealism <burdcat17@gmail.com> * feat(build): add system-ca-roots convenience alias and fix verify task Add a system-ca-roots feature alias on openshell-sandbox that includes all other defaults (telemetry) except bundled-ca-roots, so distro builds can use --no-default-features --features system-ca-roots without manually re-adding unrelated defaults. Update the verify CI task to use the alias and scope checks to the sandbox package. Fix task description to use "build mode" terminology instead of implying a Cargo feature. Signed-off-by: Quinn Burdine <sburdine@redhat.com> Signed-off-by: politerealism <burdcat17@gmail.com> * ci: fix system CA roots step name to use build mode terminology Signed-off-by: Quinn Burdine <sburdine@redhat.com> Signed-off-by: politerealism <burdcat17@gmail.com> * refactor(sandbox): reorder features to place system-ca-roots alias near default Signed-off-by: Quinn Burdine <sburdine@redhat.com> Signed-off-by: politerealism <burdcat17@gmail.com> * fix(proxy): unwrap Result from build_upstream_client_config in tests The function signature changed to return Result but the test call sites were not updated, causing type mismatch compilation errors in CI. Signed-off-by: Quinn Burdine <sburdine@redhat.com> Signed-off-by: politerealism <burdcat17@gmail.com> --------- Signed-off-by: Adam Miller <admiller@redhat.com> Signed-off-by: Scott Burdine <sburdine@nvidia.com> Signed-off-by: politerealism <burdcat17@gmail.com> Signed-off-by: Quinn Burdine <sburdine@redhat.com> Co-authored-by: Adam Miller <admiller@redhat.com> |
||
|
|
d220d89468 |
feat(compute): negotiate gateway callback listeners (#2492)
* feat(compute): query gateway listener requirements Signed-off-by: Evan Lezar <elezar@nvidia.com> * feat(compute): add Podman listener requirements Signed-off-by: Evan Lezar <elezar@nvidia.com> * test(docker): use default gateway bind address Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(gateway): avoid wildcard primary listener Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(podman): validate callback listener discovery Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(server): support split dual-stack listeners Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(podman): support legacy rootless listener discovery Signed-off-by: Evan Lezar <elezar@nvidia.com> * test(e2e): accept loopback plaintext rejection Signed-off-by: Evan Lezar <elezar@nvidia.com> * docs(agent): add callback listener diagnostics Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(server): restrict compute callback listeners Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(podman): validate local callback port Signed-off-by: Evan Lezar <elezar@nvidia.com> * test(server): clarify callback listener contract Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(podman): require pasta for local callbacks Signed-off-by: Evan Lezar <elezar@nvidia.com> * docs(gateway): document RPM listener default Signed-off-by: Evan Lezar <elezar@nvidia.com> * refactor(server): keep listener provenance diagnostic-only Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(compute): preserve callback listener isolation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(e2e): remove Podman callback relay Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(packaging): preserve Podman callback loopback Signed-off-by: Drew Newberry <anewberry@nvidia.com> * ci(e2e): run VM smoke on nested-virt runner Signed-off-by: Drew Newberry <anewberry@nvidia.com> * ci(e2e): gate VM smoke on usable KVM Signed-off-by: Drew Newberry <anewberry@nvidia.com> * ci(e2e): probe KVM through VM driver Signed-off-by: Drew Newberry <anewberry@nvidia.com> * ci(e2e): tolerate hosted KVM denial Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(server): close traced futures before assertions Signed-off-by: Drew Newberry <anewberry@nvidia.com> * revert: remove tracing test stabilization Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Evan Lezar <elezar@nvidia.com> Signed-off-by: Drew Newberry <anewberry@nvidia.com> Co-authored-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
fa24299092 |
feat(gateway): export traces over OTLP (#2534)
Add an opt-in OTLP/gRPC trace exporter to the gateway. Export is enabled
by the presence of an `[openshell.gateway.otlp]` table with an endpoint;
there is no separate toggle.
Instrumented:
- Inbound request server spans, named for the RPC (`$service/$method`) or
`{method} {path}` for plain HTTP. They continue valid W3C `traceparent`
context when present and start a new trace otherwise. gRPC spans also
carry `rpc.system`, `rpc.service`, `rpc.method`, and trailer-derived
`rpc.grpc.status_code`.
- Compute driver calls (create, delete, list, get, validate, watch) as
client spans anchored on the `ComputeDriver` contract.
- Store reads and writes as children of the current request or loop span.
- Work with no inbound request: compute driver initialization, the sandbox
reconcile sweep, provider credential refresh tick, and driver watch events.
Each roots one operation trace so its child work does not arrive as anonymous
single-span traces.
This is deliberately not exhaustive. Auth, policy evaluation, and
middleware remain uninstrumented, as do store lifecycle calls (`ping`,
`close`) that a readiness poll would turn into a span per tick. The aim is
a useful trace tree at a reviewable size; coverage can grow against real
traces.
Design notes:
- The TOML table owns whether and where to export. The SDK `OTEL_*`
variables own how; sampling, batching, and limits are not mirrored into
gateway config.
- The OpenTelemetry layer exports spans only. Existing `tracing` events
remain on the stdout and sandbox-log paths and are not copied into trace
payloads.
- Telemetry never blocks the gateway. A malformed endpoint logs an error
and disables export rather than failing startup, and buffered spans are
drained during graceful shutdown.
- Failed spans carry error status without a separate `error.type` attribute.
Request spans use HTTP status and gRPC response trailers; driver spans use
the returned gRPC status; autonomous loop spans record failed results
explicitly. Store spans exempt `UniqueViolation` and `Conflict`, because
those errors report expected contention such as a held lease or an
optimistic-concurrency retry.
- The compute driver is reachable only through `TracedDriver::call`, so a
call cannot skip its span. This is the client half of a client/server pair
and the single place to inject context if drivers move out of process.
- Tests share one process-wide subscriber and in-memory exporter because
`tracing` caches callsite interest globally.
Inbound W3C trace context is propagated into gateway request spans. Context
is not yet injected into outbound driver calls, so a future out-of-process
driver would still need propagation at the `TracedDriver` seam.
The Helm chart is intentionally unchanged, so OTLP export cannot yet be
enabled on a chart-deployed gateway.
Refs #2507
Signed-off-by: Kris Hicks <khicks@nvidia.com>
|
||
|
|
5432d01d5a |
feat(tui): add config key support to provider create/update forms (#2224)
* feat(tui): add config key editing to create provider form Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * feat(tui): add config key editing to update provider form Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * fix(tui): add Up/Down navigation for config_cursor in provider forms Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * fix(tui): show config values in provider detail view Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * fix(tui): improve provider config form validation and reduce duplication Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * fix(tui): fix config deletion and invisible cursor in provider forms Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * test(tui): add tests for config deletion tombstones and cursor scroll window Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * fix(tui): send only config delta on provider update Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * fix(tui): flush pending config input on provider update submit Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * fix(tui): fix provider config form loading, focus reset, and overflow Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * fix(tui): separate config entry navigation from key input focus Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * fix(tui): move config cursor to newly added entry after flush Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * fix(tui): always populate provider_entries regardless of providers_v2_enabled Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> --------- Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> |
||
|
|
d556748771 |
feat(supervisor-middleware): add network egress middleware (#2027)
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> |
||
|
|
aa483ecb9a |
feat(providers): AWS STS AssumeRole refresh strategy and aws-s3 profile (#1782)
Add gateway-managed AWS STS credential refresh (provider-v2, #1576). The gateway calls sts:AssumeRole and writes three short-lived credentials (AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, AWS_SESSION_TOKEN) to the provider record; the proxy re-signs requests with SigV4. Adds the aws and aws-s3 provider profiles and a declarative multi-output refresh model (additional_outputs) so one AssumeRole co-mints all three credentials. Signed-off-by: Russell Bryant <rbryant@redhat.com> |
||
|
|
83003e80fc |
feat(interceptors): initial gateway interceptor implementation and reference example (#2005)
* feat(gateway): add descriptor-driven interceptors Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(gateway): add service-reflected interceptors Signed-off-by: Drew Newberry <anewberry@nvidia.com> * wip * fix(gateway): harden interceptor evaluation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(interceptors): label metrics and harden governance smoke Signed-off-by: Drew Newberry <anewberry@nvidia.com> * remove on_error: ignore * feat(gateway-interceptors): emit log annotations Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(examples): govern provider profiles in interceptor Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(gateway): preserve update config annotations Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(providers): support interceptor profile catalogs Signed-off-by: Drew Newberry <anewberry@nvidia.com> * wip * feat(governance-interceptor): sign provider profiles Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(providers): use configured profile sources for refresh updates Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(gateway-interceptors): add phase-specific evaluation payloads Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(providers): compose provider profile sources Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(gateway-interceptors): preserve committed responses Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(gateway-interceptors): reject ambiguous protobuf oneofs Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(gateway-interceptors): validate patch candidates per binding Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(gateway-interceptors): use reflected protobuf codec Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(governance-example): canonicalize signed protobuf hashes Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(gateway): commit policy provenance atomically Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(gateway): close signed governance bypasses Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(gateway): isolate interceptor secrets and authority Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(gateway): snapshot provider profiles per request Signed-off-by: Drew Newberry <anewberry@nvidia.com> * chore(gateway): resolve server clippy warnings Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(server): satisfy provider source clippy lint Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(server): initialize policy test annotations Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(gateway-interceptors): require explicit route allowlist Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(proto): clarify update annotation semantics Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
7bce1223dc |
feat(policy): add JSON-RPC and MCP L7 policies (#1865)
Add policy schema, proto, provider profile, OPA, and L7 proxy support for `protocol: json-rpc` and `protocol: mcp`. Generic JSON-RPC endpoints match exact method names only, with `method: "*"` as the all-method sentinel; wildcard/glob methods and params matchers are rejected. Parse JSON-RPC request bodies and batches in the forward proxy, deny response-shaped client frames, limit receive-stream GET allowance to MCP endpoints, and redact params in decision logs. Preserve L7 rule params on the proto load path so MCP `tools/call` tool filters behave like YAML-loaded policies. Add MCP conformance coverage, JSON-RPC L7 e2e coverage, and docs for the new protocols and current matcher limitations. Signed-off-by: Kris Hicks <khicks@nvidia.com> Co-authored-by: ddurst <267424412+ddurst-nvidia@users.noreply.github.com> |
||
|
|
702cbc4f63 |
feat(providers): support SPIFFE-backed token grants (#1784)
* feat(providers): support SPIFFE-backed token grants Add provider profile token_grant metadata and expand endpoint-specific dynamic credentials so sandbox supervisors can request SPIFFE JWT-SVIDs, exchange them with an OAuth-style token endpoint, cache returned access tokens, and inject bearer tokens into matching HTTP requests. Wire Kubernetes and Helm deployments to mount the provider SPIFFE Workload API socket into sandbox pods for token grant exchange. Signed-off-by: Taylor Mutch <taylormutch@gmail.com> Signed-off-by: Gordon Sim <gsim@redhat.com> * test(examples): add SPIFFE token grant demo Add a reusable alpha/beta demo that deploys a SPIFFE-verifying token issuer and protected services, imports a token-grant provider profile, creates a sandbox, and verifies endpoint-specific bearer tokens. The script leaves Kubernetes workloads in place, deletes sandboxes through openshell unless KEEP_SANDBOX=1, and prints protected service logs as proof of life. Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(providers): harden SPIFFE token grants * fix(providers): harden dynamic token grants Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(providers): harden token grant handling --------- Signed-off-by: Taylor Mutch <taylormutch@gmail.com> Signed-off-by: Gordon Sim <gsim@redhat.com> |
||
|
|
e7f965a988 |
refactor(sandbox,driver-vm): Start moving to rustix (esp over libc unsafe) (#1505)
In the Rust ecosystem there's largely three ways to do system calls: - raw libc - nix - rustix Of the three, libc is almost all `unsafe` and really 95% of use cases should be either nix or rustix. nix is the original one, but after having looked at the code of both, I think rustix is just better designed and organized. It's also reached 1.0, whereas nix is still making semver-breaking changes (in fact we're behind here in this project). Now in practice, we have both *transitively* in the depchain already, and that's true for quite a lot of projects. But I think rustix is better, so let's add rustix as a workspace dependency (process feature) and migrate a few use cases to it - it's especially better than the raw libc which is suprisingly widespread. If we agree to do this, then many other calls can be ported. Signed-off-by: Colin Walters <walters@verbum.org> |
||
|
|
b61a98dbad |
feat(gateway): add TOML configuration file (RFC 0003) (#1317)
* feat(gateway): add TOML configuration file (RFC 0003)
Introduces an opt-in --config / OPENSHELL_GATEWAY_CONFIG flag that loads a
TOML file with gateway-wide settings and per-driver tables. Source
precedence is CLI > env > file > built-in default, implemented via clap's
ValueSource so existing flags and env vars keep their priority.
Driver crates (kubernetes, docker, podman, vm) now derive Deserialize on
their config structs. SupervisorSideloadMethod gains Deserialize with
kebab-case rename. A per-driver inheritance allowlist on the loader side
overlays [openshell.gateway] shared defaults (default_image,
supervisor_image, image_pull_policy, guest_tls_*, ssh_handshake_skew_secs,
client_tls_secret_name, host_gateway_ip, enable_user_namespaces) onto
each [openshell.drivers.<name>] table before deserialization.
The Helm chart renders a new gateway-config ConfigMap and mounts it at
/etc/openshell/gateway.toml. The migrated OPENSHELL_* env entries are
dropped from the StatefulSet — only the Secret-backed
OPENSHELL_SSH_HANDSHAKE_SECRET remains. database_url stays on --db-url.
Adds examples/gateway/gateway.example.toml and updates architecture/gateway.md
with the source precedence and inheritance rules.
* docs(gateway): drop ssh_handshake_skew_secs and ssh_handshake_secret from examples
Both fields are scheduled for removal. Remove the example values and the
env-only note so the gateway.toml example and the architecture doc stop
recommending settings that will not exist much longer.
* docs(rfc): correct OPENSHELL_CONFIG to OPENSHELL_GATEWAY_CONFIG in RFC 0003
* docs(gateway): add per-driver TOML example configurations
Adds focused single-driver examples next to the comprehensive
gateway.example.toml: kubernetes, docker, podman, and microvm. Each one
demonstrates the realistic settings for that driver plus how shared
[openshell.gateway] defaults inherit into the driver table.
A new unit test (`checked_in_examples_parse`) loads every example through
the config_file loader so schema drift fails CI rather than silently
shipping a broken example.
* refactor(gateway): drop image_pull_policy from shared inheritance
Kubernetes and Podman use mutually-incompatible vocabularies for the same
TOML key:
- Kubernetes: `Always | IfNotPresent | Never` (free-form string passed
verbatim to the K8s API).
- Podman: `always | missing | never | newer` (strict lowercase enum
deserialised into `ImagePullPolicy`).
No value means the same thing in both drivers. Sharing the key at
`[openshell.gateway]` scope and inheriting it into every active driver's
table meant any value safe for one driver was either wrong or silently
dropped for the other (`IfNotPresent` → `ImagePullPolicy::Missing` after
`.unwrap_or_default()`). Operators run one driver per gateway, so the
"shared default" never pays for itself.
Make `image_pull_policy` driver-local:
- Remove the field from `GatewayFileSection` and from
`inheritable_keys()` for both Kubernetes and Podman.
- Drop the file→`RunArgs` merge for the gateway-scope key.
- Stop unconditionally clobbering the driver value with
`config.sandbox_image_pull_policy` in the runtime wiring — only apply
the CLI/env override when it was set (and, for Podman, only when it
parses into the lowercase enum).
- Move the key under `[openshell.drivers.kubernetes]` and
`[openshell.drivers.podman]` in every example, the RFC, the
architecture doc, and the Helm-rendered gateway ConfigMap.
The supervisor pull policy follows the same shape: it is K8s-only and
moves into `[openshell.drivers.kubernetes]` alongside `image_pull_policy`
in the Helm template.
* fix(gateway): address review feedback on TOML configuration
Resolves the P1 and P2 issues raised in PR #1317:
- Helm gateway ConfigMap moves `grpc_endpoint` under
`[openshell.drivers.kubernetes]` so the default install no longer fails
the gateway's `deny_unknown_fields` schema check.
- `kubernetes_config_from_file` and `podman_config_from_file` only let
the gateway-wide CLI/env `grpc_endpoint` overwrite the driver-table
value when it was actually supplied, preserving file-only configs.
- Kubernetes driver default `image_pull_policy` is now empty (was Podman
vocabulary "missing"), so default deployments let the Kubernetes API
apply its own policy instead of being rejected.
- New `disable_tls` gateway field plumbs `.Values.server.disableTls`
through the TOML ConfigMap instead of relying on env vars dropped from
the StatefulSet.
- StatefulSet pod template now carries a `checksum/gateway-config`
annotation so `helm upgrade` rolls pods when the ConfigMap changes.
- Auxiliary listener resolution preserves the full `SocketAddr` from
`health_bind_address` / `metrics_bind_address`, so a loopback-pinned
health port is not silently relocated onto the public bind address.
- `ssh_session_ttl_secs` from the file is now applied to `Config` (it
was previously accepted by the loader but never read).
New regression coverage: cli-level merge tests for the new fields plus
helm-unittest assertions for the ConfigMap shape, checksum annotation,
and `disable_tls` rendering.
* docs(gateway): consolidate gateway TOML examples into docs reference
Replaces the per-driver example files under examples/gateway/ with a
single published reference page at docs/reference/gateway-config.mdx
covering source precedence, layout, the full example, and the four
per-driver examples (Kubernetes, Docker, Podman, microVM). Drew flagged
during PR #1317 review that the examples belong with the user-facing
docs rather than in a sibling examples/ directory.
The cross-references in architecture/gateway.md and RFC 0003 are updated
to point at the new docs page; the round-trip test in config_file.rs is
removed (schema coverage stays on the inline parses_full_example test
and per-field merge tests — doc-snippet drift belongs in a separate
docs-lint, not in a cross-tree Rust unit test).
* refactor(core): move DEFAULT_K8S_NAMESPACE into K8s driver
The constant is Kubernetes-specific (used only by KubernetesComputeConfig's
Default impl) and does not belong in openshell-core. Relocate it to the
driver crate that owns the K8s vocabulary; openshell-core retains only
truly cross-cutting defaults.
* refactor(core): move Podman bridge default into Podman driver
DEFAULT_NETWORK_NAME is Podman vocabulary, consumed only by the Podman
driver. Also drops the unused DEFAULT_IMAGE_PULL_POLICY constant.
* docs(auth): scrub remaining SSH handshake secret references
Sweeps the trailing mentions left after the rebase: the gateway
config-file module doc, the Helm gateway-config ConfigMap header,
the gateway-config.mdx env-only note, and the RPM systemd unit
comment for init-gateway-env.sh.
* docs(gateway): clarify OPENSHELL_GRPC_ENDPOINT applies to all drivers
The previous comment implied the callback endpoint was Kubernetes-only,
but the value is propagated to every compute driver (Kubernetes, Docker,
Podman, VM) and must be reachable from wherever the sandbox runs.
* refactor(gateway): move driver options into config (#1394)
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* fix(e2e): regenerate gateway config via TOML for docker + podman harnesses
The gateway CLI flags moved into TOML config tables in 560550d2 (#1394),
which made every existing e2e/with-{docker,podman}-gateway.sh invocation
fail with "unexpected argument '--sandbox-namespace'" (and a long tail
of similar driver-specific options) before the gateway could even bind.
Replace the obsolete CLI flags with a synthesized
`[openshell.drivers.<driver>]` table written to `${STATE_DIR}/gateway.toml`
and passed via the new `--config` flag. Only the gateway-wide flags
that survived 560550d2 (bind-address, port, drivers, db-url, tls-*,
disable-tls, log-level, health-port) stay on the command line.
Both scripts get a small `toml_string` helper to properly TOML-quote
the values (the previous `%q` printf format produced bash-escape, not
TOML-escape). The Docker harness also corrects two field names that
diverged from the driver schema: `docker_network_name` →
`network_name`, and the supervisor binary/image plumbing now reads
through to `supervisor_bin` / `supervisor_image` in the same table.
The Podman harness drops `--ssh-gateway-port` (deleted in 560550d2 —
gRPC + SSH are multiplexed on the same port now) and substitutes
`network_name` + `gateway_port` for the obsolete `--sandbox-namespace`
(which the Podman driver never had as a typed field).
* fix(core): swap bind-only 0.0.0.0 SSH gateway host for cluster URL host
CLI's resolve_ssh_gateway treated 0.0.0.0 as a loopback "keep as-is"
when the cluster URL was also loopback, so the SSH proxy connected to
0.0.0.0:port. The unspecified address is never a valid connect target
and is not present in any TLS cert SAN, which produced BadCertificate
TLS handshake failures during `openshell sandbox create -- ...` in
docker/podman e2e (e.g. bypass_detection).
Resolution: when the server returns 0.0.0.0 or :: as the gateway host
and both endpoints are loopback, fall back to the cluster URL's host
(which the CLI is already using to reach the gateway, so it must
resolve and match the cert).
* fix(e2e): repair podman harness on macOS
Podman 5.x with the applehv/libkrun provider no longer creates the legacy
~/.local/share/containers/podman/machine/podman.sock symlink, and
`podman system service` is a Linux-only subcommand — the macOS client
delegates the API service to the VM. Both assumptions in the harness
were stale, so the script tried to start a temporary service that podman
rejected with "unknown flag: --time".
- Discover the macOS socket via `podman machine inspect` instead of the
hardcoded path.
- On Darwin, fail fast with a "start podman machine" message rather than
attempting the Linux-only `podman system service` fallback.
- Write socket_path into [openshell.drivers.podman] so the in-process
driver picks up the discovered socket; the driver reads TOML only
after the config refactor (560550d2), so OPENSHELL_PODMAN_SOCKET alone
was no longer enough.
* fix(server): use clone_from for TLS client CA assignment
clippy 1.95.0 rejects assigning the result of `Clone::clone()` to an
existing variable under `-D warnings` (`assigning_clones`). Switch to
`clone_from(&...)` to satisfy the lint and avoid the redundant
allocation.
---------
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
|
||
|
|
f5b546e41b |
Revert "perf(build): speed up local CLI rebuilds (#1387)" (#1395)
This reverts commit
|
||
|
|
668c712b63 | perf(build): speed up local CLI rebuilds (#1387) | ||
|
|
afcd3a9ec4 |
fix(cli): use OS trust store for reqwest TLS verification (#1342)
Switch reqwest from `rustls-tls` (bundled Mozilla webpki roots) to `rustls-tls-native-roots` (OS certificate store) so that the CLI trusts the same CAs as the rest of the system. This fixes OIDC discovery failures against Keycloak instances using certificates signed by internal CAs present in the system trust bundle. |
||
|
|
028763d4db |
refactor(vm): remove legacy openshell-vm crate (#1239)
Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
4d388d2677 | ci(vm): cleanup vm build infra (#1186) | ||
|
|
04e48d585c |
feat(server): add request-ID middleware for request correlation (#1082)
Add a UUID-based request-ID middleware using tower-http's request-id feature. Each inbound request receives a unique x-request-id header (or preserves a client-supplied one), which is recorded in the tracing span and propagated to the response. This enables operators to correlate log lines across the middleware stack for a single request under concurrent load, and lets clients reference specific requests in bug reports. Signed-off-by: sauagarwa <sauagarw@redhat.com> |
||
|
|
6b21804258 |
feat(policy): add GraphQL L7 inspection (#1083)
Support GraphQL L7 policies |
||
|
|
084505425b |
feat(auth): add OIDC/Keycloak authentication with RBAC and scope-based permissions (#935)
* feat(auth): add OIDC/Keycloak authentication with RBAC Add OAuth2/OIDC authentication to the gateway server with role-based access control, CLI login flows, and full deployment plumbing. Server: JWT validation against configurable OIDC issuer (oidc.rs), JWKS key caching with TTL and rotation handling, method classification (unauthenticated/sandbox-secret/dual-auth/bearer), identity extraction with provider-agnostic Identity type, and RBAC enforcement via AuthzPolicy with configurable admin/user roles and auth-only mode. CLI: browser-based Authorization Code + PKCE flow, Client Credentials flow for CI/automation, token storage with refresh, gateway add/login/ logout commands, OIDC bearer token injection over mTLS transport, discovery endpoint for auto-configuration. Security: sandbox-secret scope restriction on UpdateConfig (policy sync only), anti-spoofing header stripping, dual-auth fallthrough from sandbox-secret to Bearer token. Deployment: OIDC config wired through DeployOptions, Docker env vars, Helm values/templates, HelmChart manifest, cluster-entrypoint.sh, and bootstrap scripts. Keycloak dev server script with pre-configured realm (test users, roles, PKCE client, CI client). Tested with Keycloak. The roles claim path and role names are configurable to support other OIDC providers. * feat(auth): add OAuth2 scope-based fine-grained permissions Add opt-in scope enforcement on top of existing OIDC role-based access control. When --oidc-scopes-claim is set, the server extracts scopes from the JWT and checks them per-method against an exhaustive scope map. Scopes: sandbox:read, sandbox:write, provider:read, provider:write, config:read, config:write, inference:read, inference:write, and openshell:all (wildcard). Methods not in the scope map require openshell:all. Scopes layer on top of roles and cannot escalate privilege. Auth-only mode (empty role names) still enforces scopes when enabled. Server: scopes_claim in OidcConfig, scope extraction from JWT (space-delimited and JSON array formats), standard OIDC scope filtering, scope check in AuthzPolicy after role check. CLI: --oidc-scopes on gateway add/start stored in metadata and consumed by gateway login, --oidc-scopes-claim on gateway start forwarded to server, scopes parameter in browser and client credentials OAuth2 flows with openid deduplication. Deployment: oidc_scopes_claim wired through DeployOptions, docker.rs, Helm, bootstrap scripts, and cluster entrypoint. Keycloak: realm config updated with built-in OIDC scopes and 9 OpenShell client scopes as optional on openshell-cli and openshell:all as default on openshell-ci. * fix(auth): address branch review findings Add GetInferenceBundle to sandbox-secret methods so sandbox inference route refresh works under OIDC. Make GetSandboxConfig dual-auth so CLI users can read sandbox settings with Bearer tokens. Preserve OIDC gateway metadata on restart — a bare gateway start without --oidc-* flags no longer erases the stored OIDC registration. Document CI client ID requirement (openshell-ci vs openshell-cli) in the testing guide. Add security note about auth-only mode blast radius for GitHub Actions. * fix(auth): complete review findings for OIDC auth boundary Move OpenShell/GetSandboxConfig from sandbox-secret-only to dual-auth so CLI users can read sandbox settings with Bearer tokens while sandbox supervisors continue using the shared secret. Add sandbox secret interceptor to the inference bundle fetch path so GetInferenceBundle works under OIDC-enabled gateways. Extract shared interceptor constructor to avoid duplication. Add GetSandboxConfig to the config:read scope map so scope enforcement applies consistently when scopes are enabled. Refactor OIDC metadata preservation into apply_oidc_gateway_metadata() with explicit resume semantics — only preserve existing OIDC metadata on real resume paths, not on fresh deployments. Update architecture docs and testing guide to reflect the corrected method classifications and add new test coverage for interceptor injection, scope requirements, metadata preservation, and dual-auth classification. * refactor(auth): use oauth2 crate for CLI OIDC flows Replace hand-written PKCE generation, authorization URL construction, token exchange, client credentials, and token refresh with the oauth2 crate's typed API. Eliminates sha2, hex, and getrandom dependencies from the CLI. The custom urlencoded() helper and manual form POST logic are replaced by BasicClient methods with proper type-state safety. Discovery and the callback server remain custom since the oauth2 crate does not provide OIDC discovery or a localhost redirect listener. * refactor(auth): move server auth modules into auth/ directory Group oidc.rs, authz.rs, identity.rs, and the auth HTTP endpoints under src/auth/ module directory. No behavioral changes. auth/mod.rs — module root, re-exports HTTP router auth/oidc.rs — JWT validation, JWKS caching, method classification auth/authz.rs — role and scope authorization policy auth/identity.rs — provider-agnostic Identity type auth/http.rs — /auth/connect and /auth/oidc-config endpoints * fix(auth): use RequestBody auth type for client credentials flow The oauth2 crate defaults to BasicAuth (HTTP Basic header) but Keycloak and most OIDC providers expect client_secret_post (credentials in the request body). Set AuthType::RequestBody explicitly to match the pre-refactor behavior. Also re-export Identity, IdentityProvider, and JwksCache from the auth module so ServerState's public API remains nameable by external consumers. * fix(auth): forward OPENSHELL_OIDC_SCOPES through cluster bootstrap Pass --oidc-scopes to gateway start so the metadata includes requested scopes after cluster bootstrap. Without this, users had to manually edit metadata.json to set scopes for gateway login. Usage: OPENSHELL_OIDC_SCOPES="openshell:all" mise run cluster * test(auth): add OIDC e2e tests for RBAC, scopes, and client credentials Add 10 end-to-end tests covering OIDC authentication against a live K3s cluster with Keycloak: RBAC (5 tests): admin can create providers, user cannot, user can list sandboxes, unauthenticated requests rejected, health probe works without auth. Scopes (4 tests): sandbox-scoped token can list sandboxes but not providers, openshell:all grants full access, no-scopes token denied. Client credentials (1 test): CI token via client_credentials grant. Tests are opt-in via OPENSHELL_E2E_OIDC=1 and OPENSHELL_E2E_OIDC_SCOPES=1 env vars. They derive the Keycloak URL from gateway metadata to match the server's configured issuer. Run with: OPENSHELL_E2E_OIDC=1 OPENSHELL_E2E_OIDC_SCOPES=1 \ PYTHONPATH=python uv run pytest e2e/python/oidc/ -v * fix(docs): fix markdown lint errors in OIDC architecture docs Add blank lines before lists and fenced code blocks to satisfy markdownlint MD031 and MD032 rules. |
||
|
|
3b7d30934e |
feat(server): add Prometheus metrics infrastructure and gRPC/HTTP request metrics (#920)
* feat(server): add Prometheus metrics infrastructure and gRPC/HTTP request metrics Add metrics exposition via a dedicated server port (--metrics-port, default disabled) following the same optional-listener pattern as the health port. The metrics crate facade records counters and histograms in the MultiplexedService hot path, and a PrometheusHandle renders them at GET /metrics on the dedicated port. Metrics added: - openshell_server_grpc_requests_total (counter: method, code) - openshell_server_grpc_request_duration_seconds (histogram: method, code) - openshell_server_http_requests_total (counter: path, status) - openshell_server_http_request_duration_seconds (histogram: path, status) * feat(helm): add metrics port to openshell-server helm chart Expose the Prometheus metrics endpoint in the helm chart by adding service.metricsPort (default 9090) to values, the --metrics-port arg and container port to the statefulset, and a metrics service port. Set metricsPort to 0 to disable. |
||
|
|
7c314e7405 |
feat(prover): add native Rust policy prover with Z3 solver (#741)
* feat(prover): add native Rust policy prover with Z3 solver Add openshell-prover crate implementing formal policy verification using Z3 SMT solving. Answers two questions about any sandbox policy: "Can data leave?" and "Can the agent write despite read-only intent?" Native Rust — no Python subprocess, no PYTHONPATH, no uv dependency. Z3 bundled via z3-sys for self-contained builds. Replaces the Python prototype from #703. Closes #699 Signed-off-by: Alexander Watson <zredlined@gmail.com> * fix(prover): skip L7-write in exfil, write bypass only on read-only intent Port two fixes from the Python branch: - Exfil query skips endpoints where L7 is enforced and working - Write bypass only fires on explicit read-only intent, not L4-only Signed-off-by: Alexander Watson <zredlined@gmail.com> * chore(prover): add missing SPDX license headers to registry and testdata YAML files * fix(prover): revert serde_yaml to serde_yml to match workspace dependency on main * fix(prover): apply cargo fmt formatting to prover and cli source files * ci: add clang and libclang-dev to CI image for z3-sys bindgen z3-sys requires libclang at build time for bindgen to generate FFI bindings. Without it, the Rust CI jobs fail on the prover crate. * fix(prover): use system libz3 instead of compiling from source Add libz3-dev to the CI image and drop the z3 `bundled` feature from the workspace dependency. This eliminates the ~30 min z3 C++ build that ran on every CI cache miss. sccache only wraps rustc — it does not intercept the cmake C++ build that z3-sys runs when `bundled` is enabled. The cargo target cache helped on warm runs but evicts on toolchain bumps and fork PRs. With the system library pre-installed in the CI image, z3 link time is always zero. A `bundled-z3` opt-in feature is added to the prover crate for local development without system z3: cargo build -p openshell-prover --features bundled-z3 Regular local dev: brew install z3 (macOS) or apt install libz3-dev (Linux), then cargo build just works. Signed-off-by: Alexander Watson <zredlined@gmail.com> Made-with: Cursor * fix(prover): use u16 for ports, align include_workdir default with runtime Port fields changed from u32 to u16 across all prover types (policy, model, finding, queries). Prevents the prover from silently accepting port values >65535 that the runtime rejects, which would produce misleading PASS results on invalid policies. Change include_workdir serde default from true to false to match the runtime (openshell-policy uses #[serde(default)] which gives false). The previous mismatch caused the prover to model a /sandbox path that does not exist at runtime, producing false analysis. Signed-off-by: Alexander Watson <zredlined@gmail.com> Made-with: Cursor * refactor(prover): embed registry at compile time, replace hand-rolled glob Embed binary and API capability registry YAML files into the binary at compile time using include_dir!. The previous approach used env!("CARGO_MANIFEST_DIR") which bakes in the build machine's source path — works in tests but breaks for installed binaries. The CWD fallback was equally fragile. The --registry CLI flag still works as a filesystem override for custom registries. Credentials remain filesystem-loaded (user-supplied data). Replace the hand-rolled glob matching (~60 lines) with the glob crate's Pattern::matches(), which is already a transitive dependency. Remove unused dependencies: openshell-policy, openshell-core, thiserror, tracing. The prover parses policy YAML directly into its own types and does not use the policy or core crate APIs. Signed-off-by: Alexander Watson <zredlined@gmail.com> Made-with: Cursor * docs: add Z3 system library to prerequisites The prover crate now links against system libz3 instead of compiling from source. Document the install steps for macOS, Ubuntu, and Fedora, and note the bundled-z3 feature flag as a fallback. Signed-off-by: Alexander Watson <zredlined@gmail.com> Made-with: Cursor * fix(prover): apply cargo fmt formatting Made-with: Cursor --------- Signed-off-by: Alexander Watson <zredlined@gmail.com> |
||
|
|
eea495e6b9 |
fix: remediate 9 security findings from external audit (OS-15 through OS-23) (#744)
* fix(install): restrict tar extraction to expected binary member Prevents CWE-22 path traversal by extracting only the expected APP_NAME member instead of the full archive contents. Adds --no-same-owner and --no-same-permissions for defense-in-depth. OS-20 * fix(deploy): quote registry credentials in YAML heredocs Wraps username/password values with a yaml_quote helper to prevent YAML injection from special characters in registry credentials (CWE-94). Applied to all three heredoc blocks that emit registries.yaml auth. OS-23 * fix(server): redact session token in SSH tunnel rate-limit log Logs only the last 4 characters of bearer tokens to prevent credential exposure in log aggregation systems (CWE-532). OS-18 * fix(server): escape gateway_display in auth connect page Applies html_escape() to the Host/X-Forwarded-Host header value before rendering it into the HTML template, preventing HTML injection (CWE-79). OS-17 * fix(server): prevent XSS via code param with validation and proper JS escaping Adds server-side validation rejecting confirmation codes that do not match the CLI-generated format, replaces manual JS string escaping with serde_json serialization (handling U+2028/U+2029 line terminators), and adds a Content-Security-Policy header with nonce-based script-src. OS-16 * fix(sandbox): add byte cap and idle timeout to streaming inference relay Prevents resource exhaustion from upstream inference endpoints that stream indefinitely or hold connections open. Adds a 32 MiB total body limit and 30-second per-chunk idle timeout (CWE-400). OS-21 * fix(policy): narrow port field from u32 to u16 to reject invalid values Prevents meaningless port values >65535 from being accepted in policy YAML definitions. The proto field remains uint32 (protobuf has no u16) with validation at the conversion boundary. OS-22 * fix(deps): migrate from archived serde_yaml to serde_yml Replaces serde_yaml 0.9 (archived, RUSTSEC-2024-0320) with serde_yml 0.0.12, a maintained API-compatible fork. All import sites updated across openshell-policy, openshell-sandbox, and openshell-router. OS-19 * fix(server): re-validate sandbox-submitted security_notes and cap hit_count The gateway now re-runs security heuristics on proposed policy chunks instead of trusting sandbox-provided security_notes, validates host wildcards, caps hit_count at 100, and clamps confidence to [0,1]. The TUI approve-all path is updated to use ApproveAllDraftChunks RPC which respects the security_notes filtering gate (CWE-284, confused deputy). OS-15 * chore: apply cargo fmt and update Cargo.lock for serde_yml --------- Co-authored-by: John Myers <johntmyers@users.noreply.github.com> |
||
|
|
a912848217 | refactor(build): unify image build graph for cache reuse (#390) | ||
|
|
764fac79aa |
feat(tui): add log copy and visual selection mode (#276)
* feat(tui): add log copy and visual selection mode Add clipboard support for the TUI log viewer using OSC 52 escape sequences. Users can copy individual log lines (y), the visible viewport (Y), or enter visual selection mode (v) to select a range of lines with j/k and yank with y. - Add clipboard module with OSC 52 base64 encoding - Add format_log_line_plain() for plain-text log rendering - Add visual selection mode with highlight styling and status bar - Context-switch nav bar hints between normal and visual modes - Add 7 unit tests for plain-text log formatting Closes #274 * fix(tui): write OSC 52 to /dev/tty instead of stdout ratatui owns stdout via the alternate screen buffer, so OSC 52 escape sequences written to stdout are swallowed by the backend. Write directly to /dev/tty to bypass the buffer and reach the terminal emulator. |
||
|
|
fcf12dff62 |
feat(tui): support light terminal backgrounds with adaptive theme (#265)
* feat(tui): support light terminal backgrounds with adaptive theme Replace the hardcoded dark-only color constants in theme.rs with a runtime Theme struct that has dark() and light() factory constructors. The active theme is stored on App and threaded through all draw functions, enabling the TUI to render correctly on both dark and light terminal backgrounds. - Add Theme struct with 16 semantic style fields and ThemeMode enum - Detect terminal background via COLORFGBG env var at startup - Add --theme dark|light|auto CLI flag and OPENSHELL_THEME env var - Migrate all 499 styles:: references across 11 UI files to app.theme - Clean up 9 inline Style constructions that bypassed the theme system - Add unit tests verifying dark/light palettes and legacy regression Closes #264 * fix(tui): use OSC 11 query for reliable light/dark terminal detection Replace COLORFGBG env var heuristic with terminal-colorsaurus crate, which sends an OSC 11 query to read the actual background color. This fixes auto-detection in iTerm2 and other terminals that don't set COLORFGBG. |
||
|
|
984d1a6e5c | chore: rename project from NemoClaw to OpenShell (#198) | ||
|
|
a2de1f24e5 | feat: add Cloudflare tunnel auth support (#178) | ||
|
|
4a78865b90 | feat(ci): add CLI binary builds and snapshot release to publish workflow (#110) | ||
|
|
1d7909cb38 |
chore: add open-source compliance files and SPDX headers (#71)
Add Apache 2.0 licensing, SPDX copyright headers on all source files, DCO enforcement, third-party notices, and CI enforcement. - LICENSE: Apache License 2.0 full text - DCO: Developer Certificate of Origin 1.1 - SPDX headers on all 176 source files (.rs, .py, .proto, .rego, .sh, .toml, .yaml, Dockerfiles) - scripts/update_license_headers.py: header management with --check mode - scripts/generate_third_party_notices.py: dependency license aggregation - THIRD-PARTY-NOTICES: generated listing of all Rust and Python deps - build/license.toml: mise tasks for license:check and license:update - CI: license-headers job in checks.yml, DCO check workflow - CONTRIBUTING.md: DCO sign-off requirement and license header docs - Cargo.toml: license changed to Apache-2.0, repository URL updated - pyproject.toml: license field added Closes #58 |
||
|
|
1c5051209c |
feat(cli): add dynamic shell completion support (!59)
> **🏗️ build-from-issue-agent** Closes #71 ## Summary Add dynamic shell completion support using `clap_complete`'s `CompleteEnv`, plus a `navigator completions <shell>` subcommand (rustup-style) that outputs completion scripts to stdout with detailed per-shell setup instructions in `--help`. Completions are dynamic — the shell calls back into the `navigator` binary on each tab-press, so they are always in sync with the installed version. This also lays the groundwork for future runtime completers (e.g., listing sandbox names on tab, tracked in #83). ## Changes Made - `Cargo.toml`: Added `clap_complete = { version = "4.5", features = ["unstable-dynamic"] }` to workspace dependencies - `crates/navigator-cli/Cargo.toml`: Added `clap_complete` dependency - `crates/navigator-cli/src/main.rs`: - Added `CompleteEnv::with_factory(Cli::command).complete()` at the top of `main()` (before TLS/argument parsing) - Added `Completions { shell }` subcommand that re-invokes the binary with `COMPLETE=<shell>` and pipes the output - Added `CompletionShell` enum (bash, fish, zsh, powershell) - Added `COMPLETIONS_HELP` with per-shell setup instructions (similar to `rustup completions --help`) - Added 3 unit tests - `CONTRIBUTING.md`: Added shell completions section documenting `nav` wrapper setup ## Usage ```bash # Generate completions navigator completions fish > ~/.config/fish/completions/navigator.fish navigator completions bash > ~/.local/share/bash-completion/completions/navigator navigator completions zsh > ~/.zfunc/_navigator navigator completions powershell >> $PROFILE # See full per-shell instructions navigator completions --help # For the nav wrapper (rewrite registration target) navigator completions fish | sed 's/--command navigator/--command nav/' > ~/.config/fish/completions/nav.fish ``` ## Design Decision: Why stdout-only (no auto-install) We researched how 12 popular CLI tools handle completion installation: | Tool | Command | Writes files? | Modifies RC? | |------|---------|---------------|-------------| | gh, glab, kubectl, docker, rustup, mise, starship | `completion <shell>` | stdout only | No | | atuin | `gen-completions --shell <shell>` | stdout (`--out-dir` option) | No | | gcloud | `install.sh` | Yes | **Yes** | | Amazon Q | `q integrations install` | Yes | **Yes** | **10 of 12 tools** follow the same pattern: output to stdout, never modify RC files, explicit shell selection, clear `--help` instructions. Only gcloud and Amazon Q modify RC files (and both have been criticized for it). Our approach matches the industry consensus. ## Deviations from Plan - The original issue proposed a static `completions <shell>` subcommand using `clap_complete::generate()`. After discussion, we switched to dynamic `CompleteEnv` which is always up-to-date and enables future runtime completers. - Added a `navigator completions <shell>` subcommand (not in original plan) for discoverability — it wraps the `COMPLETE` env var mechanism with a familiar CLI UX. ## Tests Added - **Unit:** `cli_debug_assert` — validates CLI structure; `completions_engine_returns_candidates` — smoke test confirming the completion engine returns subcommand candidates; `completions_subcommand_appears_in_candidates` — verifies `completions` appears in tab-completion results - **Integration:** N/A - **E2E:** N/A ## Verification - [x] All tests passing - [x] Pre-commit checks passing (clippy warnings are pre-existing, not from this change) - [x] CONTRIBUTING.md updated with nav wrapper completion setup |
||
|
|
f869182d8a |
feat(sandbox): move inference execution to sandbox-local routing (!79) (!52)
> **🏗️ build-from-issue-agent** Closes #79 ## Summary Replaces the gateway-proxied inference model (`ProxyInference` RPC) with sandbox-local execution. The sandbox now resolves routes from either a standalone YAML route file or a cluster bundle fetched via the new `GetSandboxInferenceBundle` RPC, then forwards requests directly to inference backends using `navigator-router`. ### Benefits of removing the gRPC hop Moving inference execution from the gateway to the sandbox eliminates the gRPC round-trip for every inference request: - **No payload size limits**: Direct HTTP forwarding replaces protobuf-over-gRPC serialization. - **Streaming-ready**: Preserves connection semantics for future SSE support (e.g., `stream: true`). - **Lower latency**: Direct sandbox → backend path instead of sandbox → gateway → backend round-trip. - **Reduced gateway load**: Gateway only serves the control-plane bundle delivery RPC. - **Simpler error model**: HTTP status codes flow through directly without gRPC wrapping. ## Changes Made - **Proto**: Removed `ProxyInference` RPC. Added `GetSandboxInferenceBundle` RPC for route bundle delivery. - **Gateway**: Replaced `proxy_inference` with `get_sandbox_inference_bundle`. Removed `Router` and `navigator-router` dependency — gateway is now control-plane only for inference. - **Router**: Switched config file format from TOML to YAML. Added `Debug` impl that redacts API keys. - **Sandbox**: New `InferenceContext` with local `Router` + route cache. Routes loaded from `--inference-routes` file (standalone) or cluster bundle via gRPC, with 30s background refresh. New `--inference-routes` / `NAVIGATOR_INFERENCE_ROUTES` CLI arg. - **Dev sandbox**: Added `inference-routes.yaml` at repo root with default NVIDIA NIM route. `mise run sandbox` now mounts it automatically and supports `-e` flag to forward env vars (e.g., `mise run sandbox`). - **Example**: Added standalone example (`examples/inference/routes.yaml`) and rewrote README to cover both standalone and cluster workflows. - **E2E tests**: Updated existing test names/docs. Added tests for Anthropic messages protocol and multi-route policy filtering. ## Tests Added - **Unit (21 new):** `router_error_to_http` (6 variants), `load_from_file` YAML round-trip (5), `resolve_sandbox_inference_bundle` gRPC handler (6), `build_inference_context` route loading (4) - **E2E (2 new):** Anthropic messages protocol routing, route filtering by `allowed_routes` - **Total:** 188 unit tests passing across 3 crates ## Documentation Updated - `architecture/gateway.md`: Removed router, updated Inference Service - `architecture/sandbox.md`: Updated orchestration flow, InferenceContext, route loading - `architecture/inference-routing.md`: Complete rewrite for sandbox-local architecture ## Verification - [x] All unit tests passing (188 across 3 crates) - [x] Pre-commit Rust checks passing (fmt, clippy, tests) - [x] Architecture documentation updated - [x] E2E tests updated with new test cases - [x] Standalone example and documentation added - [x] Default inference-routes.yaml with env passthrough for mise run sandbox |
||
|
|
f1439727b9 |
feat(sandbox): L7 protocol-aware inspection with TLS termination (!29)
## Summary Implements L7 (application-layer) protocol-aware policy enforcement for the sandbox proxy, enabling per-request allow/deny decisions based on HTTP method and path — not just host:port. ### Phase 1: HTTP/REST L7 Inspection (plaintext) - New `l7/` module with provider trait, REST parser, and relay loop - Parses HTTP/1.1 requests inside CONNECT tunnels, evaluates method+path against OPA/Rego policy - Supports `access` presets (`full`, `read-only`, `read-write`) and explicit `rules` with glob path matching - `enforcement` modes: `enforce` (deny with 403 JSON) vs `audit` (log only, forward traffic) - Structured logging for every L7 decision (`L7_REQUEST` with protocol, action, target, decision) ### Phase 2: MITM TLS Termination for HTTPS - Ephemeral CA generated at sandbox startup (`rcgen`) - Per-hostname leaf certificate cache (256 cap) for dynamic cert presentation - Client-side TLS termination (sandbox trusts ephemeral CA via `SSL_CERT_FILE`, `NODE_EXTRA_CA_CERTS`, etc.) - Upstream TLS connection with `webpki-roots` verification - Endpoints with `tls: terminate` get full L7 inspection; others remain passthrough ### Additional - Benign TLS/connection errors (close_notify, handshake EOF, reset) downgraded from WARN to DEBUG - Generic `AsyncRead + AsyncWrite` bounds throughout L7 module (works with both `TcpStream` and `TlsStream`) - Proto schema extended with `protocol`, `tls`, `enforcement`, `access`, `rules` fields on `NetworkEndpoint` - CLI `nav sandbox run` wired with `--rego-policy`/`--rego-data` flags and L7 policy display ### E2E Test Coverage - 7 L4 tests: no-matching-policy deny, wildcard binary, binary-restricted, wrong port, cross-policy isolation, non-CONNECT 405, structured log fields - 7 L7 TLS tests: full access allow, read-only deny POST, audit mode, explicit path rules, CA trust store injection, deny response JSON format, structured log fields ## Key Files | Area | Files | |------|-------| | L7 core | `l7/mod.rs`, `l7/provider.rs`, `l7/relay.rs`, `l7/rest.rs` | | TLS termination | `l7/tls.rs` | | Proxy integration | `proxy.rs`, `lib.rs`, `main.rs` | | Env/process | `process.rs`, `ssh.rs` | | OPA policy | `opa.rs`, `dev-sandbox-policy.rego` | | Proto schema | `proto/sandbox.proto` | | CLI | `navigator-cli/src/run.rs` | | E2E tests | `e2e/python/test_sandbox_policy.py` | ## Test plan - [x] 67 unit tests pass (`cargo test --workspace`) - [x] `cargo clippy --workspace --all-targets` clean - [x] 16 e2e policy tests pass (`mise run test:e2e:sandbox`) - [x] Manual smoke test: sandbox with `tls: terminate` on `api.anthropic.com:443` Closes #25 |
||
|
|
c094769a43 |
feat(server): add an inference router (!13)
Closes #3 |
||
|
|
e34444adf9 | feat(sandbox): add ssh connect to sandbox + build agent harness | ||
|
|
b1b5d90275 |
fix(sandbox): dynamically create and chown read_write directories
The sandbox now automatically creates directories listed in the policy's `read_write` section and sets ownership to the configured sandbox user/group. This runs in the supervisor process before forking the child. Previously, read_write directories had to be pre-created in the Dockerfile with correct ownership. Now arbitrary paths can be added to the policy and they will be prepared at runtime. Also adds a hardcoded chown for /sandbox in Dockerfile.factory as a belt-and-suspenders approach for the default case. |
||
|
|
d5d3c71e9c | feat(sandboxes): initial kube sandbox impl | ||
|
|
5a15de63df | feat(server): add support for entity persistence |