mirror of
https://github.com/NVIDIA/OpenShell.git
synced 2026-10-03 16:11:17 +08:00
pull-request/3931
31
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
8091f66877 |
feat(sandbox): write agent output to the container log (#4005)
Since #2726 the canonical main process's stdout and stderr are captured in pipes that feed only the in-memory replay buffer used by sandbox connect. Agent output therefore never reaches the container's own stdout and stderr, so it is missing from kubectl logs, docker logs, and podman logs and from anything that collects container logs. Before #2726 the entrypoint inherited the container's descriptors and its output appeared there. Copy the main process's output to the launcher's stdout and stderr in addition to the replay buffer, restoring the earlier behavior: - Output is copied byte for byte to the matching stream from a forwarder thread per stream, after it is published to the replay buffer. When the container runtime falls behind on a stream, that stream's reader waits instead of dropping output, so backpressure reaches the agent as it did with inherited descriptors, while the other stream and attachments keep receiving output. - Before the main process's exit is published, the output readers finish and queued output is drained to the container log, so an agent's final lines are not lost at shutdown. A 30 second deadline covers both; when it expires, readers waiting on the container log are released and drain the pipes into the replay buffer only, so a stalled container log cannot block exit reporting. - PTY-mode processes are not copied. The terminal stream carries escape sequences and echoed input, and terminal commands never reached the container log before #2726. - Exec, SSH, and SFTP sessions are not copied. Launcher log lines keep their existing format and remain in the container's stderr. They are written as whole lines, and a newline is inserted first when the agent left stderr mid-line, so launcher and agent lines do not merge. The Docker and VM drivers appended the tail of the workload's output to failure messages: Docker the workload container's log, and the VM driver the guest console, which carries the launcher's stdout and stderr. Those messages land in the sandbox's Ready condition and in platform events that the gateway republishes to the sandbox event stream. With agent output in that log, those messages would carry arbitrary agent output, including anything sensitive the agent prints, into gateway status and events. The supervisor starts its health endpoint only after the agent starts, so every Docker failure path could include agent output, and the VM driver reports one whenever the VM or host supervisor exits. Forward only the supervisor's log tail, matching the Podman driver, which reads the workload log solely to match fixed launcher markers and never forwards raw workload output. The workload's output remains available through docker logs and the VM's rootfs-console.log. Document where main process output appears in the logging docs and the cluster debugging skill. Closes #3928 Signed-off-by: Kris Hicks <khicks@nvidia.com> |
||
|
|
e21b7fd8cf |
chore(build): remove bundled Z3 support (#3275)
* chore(build): remove bundled Z3 support Signed-off-by: Simon Scatton <sscatton@nvidia.com> Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(build): preserve vendored Z3 for local gateway artifacts Signed-off-by: Simon Scatton <sscatton@nvidia.com> --------- Signed-off-by: Simon Scatton <sscatton@nvidia.com> Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> |
||
|
|
a875add234 |
feat(server): write gateway OCSF events to JSONL (#3264)
* feat(server): write gateway OCSF events to JSONL Previously, gateway security activity was available only in diagnostic output, and events not associated with a sandbox, such as TLS certificate reloads, had no independent structured record. Now, configuring `openshell.gateway.ocsf_log` writes every gateway-produced OCSF record to a bounded JSONL destination independently of `RUST_LOG`. The destination supports daily or disabled rotation, retention limits, queue bounds, and optional schema downgrade to OCSF 1.1 or 1.3. Additionally, existing gateway emitters (TLS reloads, service routing, and policy approval and auto-approval audits) emit structured events, so they reach the JSONL destination, console shorthand, and the affected sandbox's log stream. Records identify the gateway by its configured name in `device.uid` and `device.name`, shared across replicas, with `device.hostname` identifying the replica and `device.os` the gateway's operating system. Metrics and warnings expose known best-effort losses. Refs #2762 Signed-off-by: Kris Hicks <khicks@nvidia.com> * fix(mxc): attribute ETW events to gateway Signed-off-by: Evan Lezar <elezar@nvidia.com> --------- Signed-off-by: Kris Hicks <khicks@nvidia.com> Signed-off-by: Evan Lezar <elezar@nvidia.com> Co-authored-by: Evan Lezar <elezar@nvidia.com> |
||
|
|
45e3308d39 |
fix(network): refuse protocol upgrades on JSON-RPC and MCP endpoints (#3753)
* fix(network): refuse protocol upgrades on JSON-RPC and MCP endpoints JSON-RPC and MCP rules apply to each HTTP request, but the proxy could forward a request that also carried upgrade headers. After an upstream answered 101, route selection and the forward proxy relayed the connection without inspection. Refuse any request that carries an Upgrade header on JSON-RPC-family endpoints before the L7 policy decision, in every enforcement mode. Share the check with the existing h2c refusal and call it from relay_jsonrpc as well. Record the refusal as a policy denial and answer with the unsupported_l7_protocol error, because no policy rule can allow the request. If a JSON-RPC-family endpoint still receives 101, close the connection instead of relaying raw bytes. Document the refusal and the WebSocket alternative. Signed-off-by: Shiju <shiju@nvidia.com> * docs(observability): remove duplicate protocol error definition Keep unsupported_l7_protocol in the response error-code list and retain its explanation in the policy troubleshooting table. Signed-off-by: Shiju <shiju@nvidia.com> --------- Signed-off-by: Shiju <shiju@nvidia.com> |
||
|
|
2f1ea658fe |
docs(policy): clarify sandbox-local loopback access (#3740)
Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
73a181d32f |
docs: streamline README, add policy prover to architecture docs (#3718)
* docs(readme): streamline README and move reference detail to docs Restructure the README as a short path from overview to quickstart to further reading. Move prerelease install steps into the installation guide and telemetry build flags into a new observability page. Fix broken docs links and outdated runtime and credential descriptions. Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(readme): describe 0.1.0 as adding new isolation primitives Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(architecture): add policy prover as a gateway component Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs: describe OpenShell as a runtime for fleets of autonomous AI agents Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(architecture): describe policy prover as formal verification Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(architecture): name OpenShell Sandbox in component table Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(architecture): fold isolation backend into supervisor row Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(architecture): mention formal verification in overview Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(run-agent): use the published OpenCode image in the first-agent guide Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(inference): correct provider examples and readiness Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(providers): correct Google binding and provider selection Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(readme): sharpen value prop, how it works, and explore further Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(run-agent): use example Anthropic profile and add policy advisor step Import the example Anthropic profile, which now allows OpenCode, instead of editing it with sed. Add a step that shows how to review and approve mechanistic policy proposals as the agent needs more access. Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(providers): add OpenRouter example for OpenCode Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(run-agent): run OpenCode against OpenRouter with a free model Add an example OpenRouter provider profile scoped to OpenCode and switch the first-agent guide to it, using a free Nemotron model so readers do not need OpenRouter credits. Revert the OpenCode binary added to the example Anthropic profile. Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(readme): link first-agent guide and add agent skills section Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> Co-authored-by: Piotr Mlocek <pmlocek@nvidia.com> |
||
|
|
d7f921190b |
docs(policy): refresh policy documentation and references (#3563)
* docs(policy): correct schema and default policy guidance Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): add network recipes and update command reference Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): organize lifecycle guidance and troubleshooting Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): split policy overview into concepts and management tasks Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): reorganize network recipes as a cookbook Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): restructure schema reference by field group and protocol Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): align troubleshooting, advisor, and reference pages Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): fix first policy tutorial and security guidance Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): keep overview high level and move network rules to their own page Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): focus policy management on CLI workflows and remove command reference Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): clarify policy views and sandbox deletion in management guide Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): streamline network rule concepts and examples Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): correct request path wildcard semantics Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): rewrite policy advisor guide for clarity Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): clarify policy advisor scope, setup, and review Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): rewrite policy prover guide for clarity Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): explain the two uses of the policy prover Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): describe policy prover uses, boundaries, and coverage Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): place prover before advisor and troubleshooting last Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): remove unsupported CI guidance from prover page Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): tighten policy prover introduction Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): move policy change behavior into management guide Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): name prover check types and note expanding coverage Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): prefix prover and advisor sidebar labels Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): streamline policy schema reference Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): place default policy before schema reference Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): fold troubleshooting into policy management guide Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): correct tutorial log samples and GitHub push policy steps The first policy tutorial said the 403 body begins with error, policy, and rule, but the proxy serializes the body with sorted keys. Its log samples also showed the wrong CONNECT deny reason for a sandbox without network rules, and the L7 deny sample omitted the :443 authority, the `l7` engine, and the reason tag that the shorthand formatter emits. The GitHub tutorial filtered denials with `--level warn`, which hides the INFO level OCSF policy events, and showed the retired key=value log format. Its hand-written policy also omitted /bin from the restrictive default, so `policy set` would reject the file for removing a filesystem path on a live sandbox. Start from `policy get --base` and add only the network rules. Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): improve flow and terminology across policy pages Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): correct network rule matching and protocol details Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): align policy management steps with CLI behavior Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): correct policy advisor proposal and approval details Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): correct policy section, default, and schema details Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): correct prover installation and coverage limits Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): recommend tls skip for server-first protocols Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): fix stale baseline path and interpreter examples Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): move policy pages under how-it-works and fix links Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): align native TCP guidance in security best practices Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): restore policy.local and policy DNS details from main Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): state exact glob matching rules Signed-off-by: Johnny Greco <jogreco@nvidia.com> --------- Signed-off-by: Johnny Greco <jogreco@nvidia.com> |
||
|
|
9244868056 |
docs: refresh architecture and agent guides (#3705)
* docs: refresh architecture and agent guides Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs: describe updated security architecture neutrally Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs: highlight new isolation primitives Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(sandboxes): clarify how to disconnect Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs: align architecture and guides with current navigation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(extensibility): streamline extension authentication guidance Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
470a34635d |
fix(api): make WatchSandbox loss-aware and resumable (#3209)
* fix(api): emit warning on WatchSandbox broadcast lag instead of terminating Broadcast lag on the status, log, and platform receivers was converted to a RESOURCE_EXHAUSTED status that terminated the whole watch stream. Lag is recoverable: the receiver resumes at the oldest surviving message. Emit a SandboxStreamWarning and continue streaming instead; keep terminating on Closed. Add helpers and unit tests covering the warning payload and receiver recovery after lag. Partially addresses #3055 (cursor/resume follow up separately). Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * refactor(server): group per-sandbox log bus state and stamp sequence numbers Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * feat(proto): add resume cursor fields to sandbox watch API Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * feat(server): stamp watch cursors from a shared per-sandbox sequence Allocate cursors from a single SeqAllocator shared by the log and platform event buses, so a sandbox's merged watch stream carries unique, strictly increasing cursors. A single resume_after_cursor can then unambiguously locate a client's position across both sources. Rewrite both publish paths to allocate the sequence, stamp event.cursor, send, and append to the tail under one lock. This removes the previous get_mut().expect() TOCTOU race where a concurrent remove() between the two lock sections could panic. Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * feat(server): serve WatchSandbox resume from cursor with gap detection Add tail_after() to the log and platform event buses, returning every buffered event newer than a client's resume cursor. Each PerSandbox now tracks last_trimmed_seq (the highest seq it has evicted) so a resume is reported as an unrecoverable ResumeGap only when this bus dropped an event the client still needs. Judging gaps by evictions, not by the tail's oldest seq, is required under the shared cursor space: each bus's tail is non-contiguous in the global sequence because the other bus owns the missing seqs, so comparing against tail.front() would flag false gaps. Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * feat(server): resume WatchSandbox from cursor across log and platform buses Wire resume_after_cursor into the watch producer. On a non-zero cursor, replay events strictly after it from both the log and platform buses, merge by shared cursor, and emit in order before entering the live loop. A trimmed range on either bus is an unrecoverable gap and terminates the stream with OUT_OF_RANGE carrying the requested and earliest-available cursors, distinct from recoverable lag which warns and continues. Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * test(server): cover WatchSandbox cursor resume paths Add handler-level tests for the resumable watch stream: replay strictly after the client cursor, merge log and platform events in shared-cursor order, suppress duplicates when resuming at the latest cursor, and terminate with OUT_OF_RANGE when the requested cursor has been trimmed. Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * docs(api): document WatchSandbox loss-awareness and resume Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * fix(server): deliver watch events once and harden cursor teardown Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * feat(sdk): add loss-aware resumable watch_logs to Rust SDK client Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * fix(server): keep watch cursors monotonic across teardown and restart Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * fix(server): merge live watch sources by cursor before emission Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * fix(api): bind watch cursors to a cursor space and merge tail sources Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * fix(server): revalidate the watch cursor space after collecting replay Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * test(server): update the public RPC schema fingerprint for the string cursor Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * fix(server): hold watch events above the publication watermark and emit the watch lag warning before its batch Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * test(server): synchronize the watch live-order test with the end of initialization Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * fix(sdk): use canonical sandbox name in watch_logs Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * test(sdk): guard canonical-name addressing in watch_logs Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * fix(server): fix public rpc schema Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * fix(api): reconcile watch resume rebase Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(server): bound interactive relay cleanup Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> --------- Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> Co-authored-by: John Myers <9696606+johntmyers@users.noreply.github.com> |
||
|
|
cbf026366d |
fix(ocsf): require network activity endpoints (#3355)
Previously, Network Activity could be constructed without a source or destination endpoint, allowing connection, accept, relay, and configuration events to violate the OCSF 1.8 endpoint constraint. Now, NetworkActivityBuilder requires a source or destination endpoint at compile time. Connection failures identify the workload peer or genuine transparent destination, listener failures identify the listening endpoint, and mediation-lane failures use Application Lifecycle rather than fabricated network endpoints. Malformed forward requests use HTTP Activity with a method-only request, generated 400 response, and workload peer. Additionally, Unix relay-channel events use Base Event, policy-validation warnings use Config State Change, and the unused bypass monitor is removed because the current isolation architecture no longer uses it. Signed-off-by: Kris Hicks <khicks@nvidia.com> |
||
|
|
481ce566e1 |
fix(ocsf): correct HTTP activity context (#3316)
Previously, metadata events omitted both HTTP request and response objects, early proxy rejections used HTTP Activity without request context, and unsupported-scheme events did not expose enough safe HTTP context to satisfy the OCSF 1.8 schema. Now, metadata events include a method-only request and their actual HTTP response codes without recording the metadata URL. Unsupported-scheme events also include a method-only request plus the generated 400 response. Authority mismatches and credential-resolution denials use HTTP Activity with their generated 403 or 500 responses, and HTTP activity IDs are derived from the request method. Additionally, HttpActivityBuilder now enforces the OCSF request-or-response constraint at compile time. Signed-off-by: Kris Hicks <khicks@nvidia.com> |
||
|
|
ae57979b03 |
feat(mxc): add Windows ETW-to-OCSF audit trail (#3015)
* feat(mxc): ETW->OCSF audit consumer + Windows OCSF JSONL parity (cp6 P1) Add a Windows MXC ETW->OCSF audit trail in openshell-driver-mxc: a real-time Sandboxing-provider ETW consumer that decodes events (TDH), attributes each to an OpenShell sandbox_id, and maps them to OCSF (lifecycle 6002, config 5019, process 1007, finding 2004). cp6 Phase 1 - durable OCSF JSONL audit-file parity with Linux: - openshell-ocsf: add emit_ocsf_event_routed (populates the event-bridge thread-local AND stamps sandbox_id+message in one dispatch) plus public set/clear_current_event; OS-aware device (Device::windows/for_current_os) so device.os.name reflects the host instead of a hardcoded Linux stub. - etw_consumer: emit via the routed emit (previously fired a bare info! that never populated the bridge, so the structured event was dropped). - openshell-server: install OcsfJsonlLayer over a synchronous daily-rotated appender (durable under force-kill), gated by OPENSHELL_OCSF_JSON, path via %PROGRAMDATA%\OpenShell\logs (override OPENSHELL_OCSF_LOG_DIR). - device.hostname now resolves to the real gateway machine name. Box-proven on 7F203-MXC-001: JSONL lines == shorthand OCSF rows, all valid OCSF JSON, per-sandbox attribution intact, disabled state writes nothing. Signed-off-by: Akber Raza <akberr@nvidia.com> * feat(mxc): map remaining Sandboxing ETW events to OCSF Close the last three ETW->OCSF gaps so the audit trail covers the full set of events the Sandboxing provider emits (12/12): - ProcessLaunched -> Process Activity [1007] "Launch" (confirmed start; carries the real processId/threadId, the twin of CreateProcessInSandbox which only has the request + command line). - SandboxProxyConfigured -> Device Config State Change [5019] (the one network-plane setup event; surfaces proxyPort, "no proxy" when 0). - SandboxConsoleReferencePlumbed -> Device Config State Change [5019] (console-handle plumbing). map_config_state now handles the full config/hardening/setup family and carries proxyPort/hasConsoleReference/creationFlags as unmapped fields. Verified on 7F203-MXC-001: 11/12 event types emit OCSF without a proxy (SandboxProxyConfigured requires proxy config to fire). Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(mxc): seed ETW attribution under registry lock + Device tests Address CodeRabbit review on !31: - Prevent stale ETW attribution on a delete/launch race: register the wxc-exec pid while holding the registry lock, and bail if the sandbox entry is already gone. Previously the attribution key could be seeded after `delete` had removed the sandbox, leaving a stale key that could misroute later Sandboxing ETW events to a dead sandbox_id. Lock order (registry -> attribution) matches the delete path, so no deadlock. - Add unit tests for the new Device::windows and Device::for_current_os constructors to harden Windows/Linux OCSF device parity. Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(mxc-etw): buffer+replay racing events and harden attribution keys Addresses two ETW->OCSF attribution review items (Shailendra #1, #2). #2 early-event loss: ETW delivers the sandbox create/config burst the instant wxc-exec starts, which can beat the driver's register_launch (now under the registry lock post-Ready). process_event previously dropped anything unresolved, losing the racing burst. Add a bounded, time-bounded pending buffer (PENDING_MAX=4096, PENDING_TTL=5s): unresolved events are held and replayed once attribution lands, aged-out ones dropped. Consumer switched to a timed recv_timeout(200ms) so the buffer is re-driven after each event and on a tick. Emit path factored into shared emit_resolved(). #1 attribution collisions: a Windows PID is recycled after exit and a command line is commonly identical across sandboxes. register_launch now rebinds by_pid on reuse and clears the stale last_pid_sid hint (warns if the PID still pointed at a different, leaked sandbox); command line is held in by_cmd only while unique and demoted to a new ambiguous_cmds set on a second owner, so a duplicate command refuses to resolve rather than misroute. Unit tests: buffer replay (direct + cross-link), buffer bound, PID-reuse rebind, duplicate-cmd non-resolution. Box-verified on 7F203-MXC-001 (5 sandboxes, identical cmd -> 5 isolated sandbox_ids, 50/50 OCSF/JSONL, BuffersLost=0). Signed-off-by: Akber Raza <akberr@nvidia.com> * docs(mxc-etw): note cmd_line is captured raw with no privacy filtering Review item #3 (Shailendra): add a PRIVACY NOTE on map_process_launch stating cmd_line is copied verbatim into OCSF process.cmd_line with no redaction, so secrets/PII on a command line land unredacted in the durable audit trail (deliberate audit-fidelity trade-off; treat the log as sensitive). Redaction is owned by an upstream privacy layer, not this path; no general audit-output PII scrubber exists today (openshell_core::secrets [CREDENTIAL] redaction is scoped to the proxy HTTP-target logging, a separate egress path). Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(mxc-etw): open ETW trace on caller thread so start_session reports real status Review item #4 (Shailendra): start_session previously returned Ok(EtwSession) as soon as the pump thread was spawned, but OpenTraceW ran later inside that thread; if it failed we still handed back a live-looking session and logged 'consumer started' (silent failure = false audit coverage). Split the two Win32 calls instead of adding a channel handshake (avoids any lost-wakeup/hang risk): the quick, synchronous OpenTraceW now runs on the caller thread (open_trace), and only the blocking ProcessTrace runs on the pump thread (run_trace). start_session returns Err if OpenTraceW fails (reclaiming the boxed Sender so the consumer disconnects, stopping the session, joining the consumer) and returns Ok/logs 'started' only once capture is genuinely open. Opened handle + LoggerName buffer + boxed Sender are carried to the pump via a Send OpenedTrace so they outlive ProcessTrace. Box-verified on 7F203-MXC-001: consumer started=True, failed-to-start=False, 50 OCSF rows / 50 JSONL, BuffersLost=0 (no regression to capture/emit). Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(mxc-etw): guard pending-event replay against PID recycling CodeRabbit flagged that drain_resolved() re-resolved buffered events against the live by_pid map, so if Windows recycled a wxc-exec PID within PENDING_TTL a stale event from the dead sandbox could be emitted under the new owner. Stamp each by_pid registration with its Instant and add resolve_replay(), used only on the buffered/replay path. It (a) never falls back to the recycle-/ambiguity-prone by_cmd or last_pid_sid keys, and (b) trusts a PID match only when the registration is not newer than the buffered event by more than REPLAY_PID_GRACE (2s) - a recycled PID's registration lands well outside that window, so the stale event ages out instead of misattributing. The legitimate #2 seed race (registration lands ~immediately) still replays. Adds unit tests for the recycle-refusal, in-grace acceptance, and weak-fallback exclusion. Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(mxc-etw): surface unexpected ProcessTrace termination (review #4) start_session already returns Err on OpenTraceW failure (runs on the caller thread since e41a7701), closing the first half of Shailendra's #4. This closes the second half: ProcessTrace's result was discarded, so if capture died mid-run the backend had no way to know. Add a shared CaptureHealth (stopped/stopping/exit_code) between the pump thread and EtwSession. run_trace now records ProcessTrace's WIN32_ERROR and, when the pump returns without a deliberate stop, logs at ERROR that MXC OCSF capture is no longer running. EtwSession::stop() sets `stopping` before teardown so a normal shutdown isn't misreported, and EtwSession::is_capture_alive() exposes the state for status/diagnostics. Box-verified on 7F203-MXC-001: 5 sandboxes, 50 attributed OCSF rows, JSONL parity 50/50, BuffersLost=0, clean start/stop (no false failure). Signed-off-by: Akber Raza <akberr@nvidia.com> * feat(mxc-ocsf): add ETW->OCSF audit-trail example kit; fix proxy-configured message Add a runnable OCSF audit-trail example under examples/ (run-ocsf-audit.ps1, mxc-ocsf-audit.toml, ocsf-audit.yaml, README) that spins up sandboxes with the in-process ETW consumer and egress proxy on, emitting a full OCSF JSONL audit trail across all four classes (6002/5019/1007/2004). Fix SandboxProxyConfigured mapping to log "MXC sandbox proxy configured" instead of a misleading "(no proxy)" when the provider reports proxyPort=0; the event's presence already indicates proxy configuration. Verified on-box: 26 events, all mapped ETW event types present. Signed-off-by: Akber Raza <akberr@nvidia.com> * feat(mxc-ocsf): clearer audit report + client-safe run-ocsf-audit.ps1 Improve the ETW to OCSF audit-trail example output and make it safe to ship. Report: - Add an event-type coverage count ("N of M expected event types fired"); the denominator auto-adjusts (8 with proxy on, 7 with -NoProxy). - Split the checklist into expected event types vs anomaly findings (ActivityError/FallbackError), which are reported separately and not counted toward coverage (a clean run may emit none). - Verdict is now coverage-based (all expected types must fire) instead of the looser "at least 3 OCSF classes". - Call out the absolute path to the durable OCSF JSONL log prominently. Client-safety: - Default -ShareOut to empty (no auto-copy); pass -ShareOut a UNC path to opt in. Removes a hardcoded internal share path from a published example. - Drop internal-team wording ("Hand that zip back for evaluation", "BUNDLE:") in favor of neutral "Results bundle:". - Update README-ocsf-audit.txt to match the opt-in -ShareOut behavior. Verified on both MXC boxes: 7F203-MXC-001 (base-container) -> PASS, 8 of 8 event types, 26 OCSF events across 4 classes; 7F203-MXC-003 (AppContainer fallback) -> reduced set as expected, clean output. Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(mxc): configure OCSF audit workloads per sandbox - remove unsupported gateway-scoped workload fields from the shipped MXC audit example. - build the command, working directory, and filesystem grant from each run's ShareDir - pass the workload through --driver-config-json. - preserve the host CONNECT proxy configuration and conditional audit coverage for the future host_connect_proxy merge - require the workload output when determining the audit verdict. Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(mxc): omit command arguments from OCSF audit logs - record only the executable basename for MXC CreateProcessInSandbox audit events - leave process.cmd_line unset so workload arguments cannot reach shorthand or JSONL logs - cover tokens, passwords, signed URLs, and PII with a secret-leak regression test - update the audit example, architecture guidance, and published logging documentation - preserve ETW attribution and future host_connect_proxy enforcement behavior Signed-off-by: Akber Raza <akberr@nvidia.com> * feat(etw): enhance ETW session management with distinct naming for concurrent gateways * fix(etw): bound the audit queue during overload - replace the unbounded ETW callback channel with count- and byte-bounded buffering - keep the ETW callback non-blocking and count records rejected during overload - emit immediate, rate-limited warnings that identify resulting audit coverage gaps - make the audit example fail when queue overload causes dropped ETW records - cover stalled consumers, oversized events, and warning throttling with unit tests Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(etw): harden sandbox audit attribution - remove command-line and persistent per-PID fallback keys from live and replay resolution - retire the driver-owned wxc-exec PID before publishing child completion - retain established identity, activity, and correlation-vector links only for the five-second late-event window - prevent buffered records from crossing rapid PID retirement and reuse boundaries - add resolver and lifecycle coverage and document the attribution trust boundary Signed-off-by: Akber Raza <akberr@nvidia.com> * chore(mxc): address rebase follow-ups Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(mxc): align OCSF audit example with driver config - remove unsupported egress proxy settings - stop requiring the unavailable proxy audit event - update example documentation for supported event coverage Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(etw): redact command-line secrets in DecodedEtwEvent summary * fix(etw): enhance PID resolution and event attribution logic for ETW records * fix(ocsf): restrict gateway-local JSONL sink to Windows/MXC path with opt-in configuration * address rebase issues * fix(mxc): address ETW audit review feedback Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(mxc): fail closed across ambiguous PID reuse Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(mxc): bind ETW attribution to process generation Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Akber Raza <akberr@nvidia.com> Signed-off-by: Drew Newberry <anewberry@nvidia.com> Co-authored-by: Jamie King <jamiek@nvidia.com> Co-authored-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
67374efdf8 |
fix(ocsf): emit schema-valid event identities (#3247)
Previously, every event from a sandbox reused the sandbox ID as its event ID. Consumers deduplicating security records could mistake separate events for the same record, and missing device types or empty image objects could prevent schema validation. Give each event its own ID, retain the sandbox association separately, and classify the environment as Other/Sandbox while keeping the OS separate. Omit unknown container details instead of emitting empty objects. Security tooling can now distinguish events from the same sandbox and read their identity consistently after serialization. Refs #1055 Signed-off-by: Kris Hicks <khicks@nvidia.com> |
||
|
|
f4dc6be4b2 |
refactor(inference): remove managed inference routes (#3195)
* refactor(inference): remove managed inference routes Closes #3172 Remove the inference route control plane, inference.local data path, built-in router crate, and SDK surface. Move inference workloads to explicitly imported provider profiles and native endpoints, with migration cleanup and updated tests and documentation. Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(policy): preserve alternate upstream isolation Restore the provider policy activation guard so legacy OpenAI and Anthropic providers configured for alternate base URLs do not grant egress to the built-in public vendor endpoints. Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> --------- Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> |
||
|
|
48a8a4bf09 |
feat(ocsf): configurable schema version for SIEM backward compatibility (#2717)
* feat(ocsf): configurable schema version for SIEM backward compatibility Add a gateway-configurable OCSF schema version target that downgrades JSONL output for SIEMs that only support older schema versions. AWS Security Lake requires v1.1.0, Splunk CIM Add-On targets v1.1-v1.3. The downgrade filter strips profile-gated fields (ai_model, container, observation_point_id), removes unknown profiles from metadata.profiles, rewrites metadata.version, and adds an unmapped.downgraded_from breadcrumb so auditors can distinguish "no model involved" from "model attribution stripped." Supported target versions (1.1, 1.3) are enforced by an allow-list in the settings registry. Invalid values are rejected with a clear error. The setting flows to sandboxes via the settings bundle and takes effect on the next poll cycle. The shorthand log output is unaffected. Closes #2662 Signed-off-by: Adel Zaalouk <zanetworker@gmail.com> Signed-off-by: Adel Zaalouk <azaalouk@redhat.com> * fix(ocsf): align downgrade with schema version Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> --------- Signed-off-by: Adel Zaalouk <zanetworker@gmail.com> Signed-off-by: Adel Zaalouk <azaalouk@redhat.com> Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> Co-authored-by: John Myers <9696606+johntmyers@users.noreply.github.com> |
||
|
|
ef296806f5 |
feat(sandbox): add canonical main process (#2726)
* feat(sandbox): add canonical main process Closes #2710 Persist and supervise one canonical workload per sandbox, attach sandbox connect to its retained session, and make every unexpected main-process exit terminal. Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(sandbox): simplify canonical main process contract Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): preserve legacy VM main compatibility Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): preserve main status across driver updates Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): satisfy macOS process lint Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): gate Linux exit acknowledgement publisher Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(sandbox): simplify main process plumbing Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(supervisor): make controlling tty ioctl portable Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(supervisor): initialize canonical process environment Signed-off-by: Drew Newberry <anewberry@nvidia.com> * perf(supervisor): optimize retained main session Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(sandbox): detach main session on ctrl-c Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): use explicit main detach keys Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(dev): atomically stage Docker supervisor Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(test): align Docker main environment assertion Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(sdk): expose canonical main process fields Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
600bbae845 |
feat(ocsf): emit AI inference events via ai_operation profile on ApiActivity (#2664)
Apply the official OCSF ai_operation profile (introduced in v1.8.0) to ApiActivity [6003] events when the inference proxy routes a model call through inference.local. Attaches an ai_model object (name, ai_provider) and puts token counts and latency in unmapped fields. ApiActivity [6003] is the schema-correct class for the ai_operation profile in v1.8.0 (HttpActivity only gets it in v1.9.0). In Splunk CIM, ApiActivity maps to the "Change" data model, naturally separating inference events from regular HTTP proxy traffic. Changes: - Add AiModel object and ai_model field on BaseEventData - Add ApiActivityEvent struct and ApiActivityBuilder - Add emit_ai_inference in proxy.rs using ApiActivity with ai_operation - Vendor OCSF v1.8.0 schemas including api_activity class, ai_model object, and ai_operation profile definitions - Bump OCSF_VERSION to 1.8.0 - Update schema validation to skip profile-gated required fields Shorthand: API:INFERENCE [INFO] claude-3-haiku via anthropic 701ms [POST /v1/messages] Splunk/SIEM backward compatibility (v1.1/v1.3 CIM mapping) is tracked separately in #2662 as a configurable serialization concern. Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
44bf0df485 |
feat(middleware): inspect WebSocket text messages (#2477)
* feat(middleware): inspect websocket text messages Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(middleware): address websocket review feedback Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(network): bound websocket message assembly Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(network): harden websocket upgrade lifecycle Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * refactor(middleware): unify in-process and remote transports Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * feat(middleware): support regex websocket redaction Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(middleware): bound persistent streaming sessions Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(middleware): accept websocket sequence gaps Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * refactor(middleware): refine websocket introspection contract Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(middleware): clarify websocket preflight lifecycle Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(middleware): clarify websocket coverage semantics Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(network): type websocket frame failures Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(network): return 503 when middleware admission is exhausted Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(middleware): align streaming API contract Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(middleware): clarify WebSocket event result scope Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(rfc): simplify middleware revision history Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * feat(examples): add WebSocket content guard support Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(middleware): unify binding payload limits Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * refactor(middleware): align payload limit terminology Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(middleware): address websocket review feedback Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(network): address websocket review findings Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * test(network): allow Linux handler setup in preflight regression Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(network): harden websocket relay finalization Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(network): inspect compressed websocket messages Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * test(network): stabilize compressed websocket regressions Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(go-sdk): regenerate middleware protobuf binding Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(middleware): clarify websocket skip lifecycle Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(middleware): address WebSocket review feedback Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> --------- Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> |
||
|
|
472e23f96a |
fix(proxy): include OPA deny reason in CONNECT 403 response (#2363)
* fix(proxy): include OPA deny reason in CONNECT 403 response When a CONNECT request was denied by OPA policy, the 403 response used a generic "not permitted by policy" message for both "endpoint not in policy" and "endpoint matched but binary didn't match." Users had no way to distinguish the two without reading supervisor logs. The OPA policy already computes a detailed deny_reason (e.g., "binary '/usr/bin/node' not allowed in policy 'X'") but the proxy was not including it in the HTTP response. Now the CONNECT deny response includes a "reason" field with the OPA deny reason when available. When the reason is empty, the field is omitted for backward compatibility. Fixes #2355 Signed-off-by: Adel Zaalouk <azaalouk@redhat.com> * docs(observability): document optional reason field in CONNECT 403 response The proxy now includes a reason field in the JSON body of denied CONNECT responses when the policy engine provides a specific denial cause. Update the Proxy Error Responses section to show the field and describe when it is present vs omitted. Signed-off-by: Adel Zaalouk <azaalouk@redhat.com> --------- Signed-off-by: Adel Zaalouk <azaalouk@redhat.com> |
||
|
|
d556748771 |
feat(supervisor-middleware): add network egress middleware (#2027)
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> |
||
|
|
e98ea3ee93 | feat(policy): add agentic approval loop (#1528) | ||
|
|
8322e4fd00 |
docs: style fixes (#1341)
* docs: style fixes * docs: drop observability section overview page and rename a section title * docs: title updates |
||
|
|
0914f3f4f7 |
fix(sandbox): log L7 parse denials (#1072)
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> |
||
|
|
a6d45528c1 | feat(server,sandbox): supervisor-initiated SSH connect and exec over gRPC-multiplexed relay (#867) | ||
|
|
40e9bf6feb |
feat(policy): add incremental sandbox policy updates (#860)
* feat(policy): add incremental sandbox policy updates Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * docs(policies): expand incremental update guidance Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * feat(policy): audit incremental updates in gateway logs * docs(policy): quote glob specs in shell examples --------- Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> |
||
|
|
b7c763204d |
docs: refresh user-facing docs for recent sandbox and inference changes (#868)
* docs: refresh user-facing docs for recent sandbox and inference changes - architecture: document system CA loading for upstream TLS, `tls: skip` as the opt-out, gateway state persistence across restarts, and OCSF structured logging surface. - inference: document per-provider header allowlist, Authorization stripping, 120s streaming idle tolerance, and extended-thinking timeout guidance. - manage-sandboxes: add "Execute a Command in a Sandbox" section for `openshell sandbox exec` with flag reference. - security best practices: expand seccomp denylist (unconditional and conditional blocks), document two-phase Landlock probe, High-severity `landlock-unavailable` finding, and inference keep-alive closure. - observability logging: document port in HTTP log URLs, `[reason:...]` denial suffixes, proxy 403/502 JSON error bodies, and Landlock CONFIG:ENABLED/CONFIG:OTHER events. Signed-off-by: Miyoung Choi <miyoungc@nvidia.com> * docs(architecture): tighten high-level architecture page Trim implementation detail (CA bundle paths, deprecated TLS keys, SSH handshake secret) from the high-level architecture page, fix accuracy issues surfaced during deep audit, and expand uncommon acronyms on first mention. - Drop unsupported "cost-based routing" claim from Privacy Router row. - Replace "brokers requests across the platform" with auth-boundary description. - Add "inference" to Policy Engine constraint list per AGENTS.md. - Expand Deny rule to include SSRF, blocked control-plane port, and L7 deny paths in addition to deny-by-default. - Switch Allow/Deny labels from hyphen to colon; remove em dashes and a double space. - Expand LLM, SSRF, L7, TLS, CA, PEM, SSH, OCSF, and JSONL on first use. Signed-off-by: Miyoung Choi <miyoungc@nvidia.com> Made-with: Cursor * docs: address audit feedback on refresh PR - observability/logging: rewrite the allowed_ips paragraph after the Denial Reasons table; the previous wording said authors "can use" invalid entries while also stating they were rejected, which was contradictory and conflated load-time validation with the runtime per-CONNECT denial phrases the section documents. - about/architecture: split compound sentences in the new Gateway Lifecycle and Observability sections so each clause stands alone. - inference/about: drop the streaming-tolerance sentence from the prose paragraph since the dedicated Streaming reliability table row already covers it. Signed-off-by: Miyoung Choi <miyoungc@nvidia.com> Made-with: Cursor * docs(architecture): address review feedback on architecture page - Drop the Privacy Router row from the Components table; it is not yet a separately exposed customer-facing component. - Update the page description and intro count to match the remaining three components (gateway, sandbox, policy engine). - Split the policy decision into the three modes that the engine actually implements: Explicit Deny (deny rules and hardening rules, takes precedence), Allow, and Implicit Deny (no rule matched). Signed-off-by: Miyoung Choi <miyoungc@nvidia.com> Made-with: Cursor * docs(architecture): convert policy decisions to a table Promote the three policy decisions (Explicit Deny, Allow, Implicit Deny) to a top-level table with Decision, When it applies, and Outcome columns instead of a nested bulleted list under list item 5. Top-level tables render reliably across markdown renderers, where nested-in-list tables do not. Signed-off-by: Miyoung Choi <miyoungc@nvidia.com> Made-with: Cursor * docs: repharse a bit --------- Signed-off-by: Miyoung Choi <miyoungc@nvidia.com> |
||
|
|
2ca553a4a0 |
fix(sandbox): validate always-blocked IPs at load time, enrich denial logs, and filter un-fixable proposals (#814) (#815)
Policies with allowed_ips entries targeting loopback, link-local, or unspecified ranges now fail at connection time instead of being silently blocked at runtime. The shorthand log format for DENIED events includes a [reason:...] suffix so operators can distinguish 'allowlist miss' from 'structurally un-allowable'. The mechanistic mapper skips proposals for always-blocked destinations, preventing the infinite TUI notification loop. The gateway validates proposed rules on approval as defense-in-depth. - Extract shared IP helpers (is_always_blocked_ip, is_always_blocked_net, is_internal_ip) to openshell_core::net - Reject always-blocked entries in parse_allowed_ips with hard error - Skip implicit allowed_ips synthesis for always-blocked literal IP hosts - Add status_detail to HttpActivityBuilder for denial reason propagation - Enrich NET and HTTP shorthand with [reason:...] for DENIED events - Add engine: tag to HTTP shorthand (consistency with NET shorthand) - Filter always-blocked proposals in mechanistic mapper generate_proposals - Add validate_rule_not_always_blocked server-side defense-in-depth - Update architecture docs, published docs, and E2E test assertions |
||
|
|
8b15ef772e |
docs(fern): move published docs into docs tree (#796)
Remove the legacy Sphinx pipeline and make docs/ the single source of truth so the published site matches the repository layout. |
||
|
|
b7779bdefa |
feat(sandbox): integrate OCSF structured logging for sandbox events (#720)
* feat(sandbox): integrate OCSF structured logging for all sandbox events WIP: Replace ad-hoc tracing calls with OCSF event builders across all sandbox subsystems (network, SSH, process, filesystem, config, lifecycle). - Register ocsf_logging_enabled setting (defaults false) - Replace stdout/file fmt layers with OcsfShorthandLayer - Add conditional OcsfJsonlLayer for /var/log/openshell-ocsf.log - Update LogPushLayer to extract OCSF shorthand for gRPC push - Migrate ~106 log sites to OCSF builders (NetworkActivity, HttpActivity, SshActivity, ProcessActivity, DetectionFinding, ConfigStateChange, AppLifecycle) - Add openshell-ocsf to all Docker build contexts * fix(scripts): attach provider to all smoke test phases to avoid rate limits GitHub's unauthenticated API rate limit (60/hour) causes flaky 403s for Phases 1, 2, and 4. Fix by attaching the provider to all sandboxes and upgrading the Phase 1 policy to L7 so credential injection works. Phase 4 (tls:skip) cannot inject credentials by design, so relax the assertion to accept either 200 or 403 from upstream -- both prove the proxy forwarded the request. * fix(ocsf): remove timestamp from shorthand format to avoid double-timestamp The display layer (gateway logs, TUI, sandbox logs CLI) already prepends a timestamp. Having one in the shorthand output too produces redundant double-timestamps like: 15:49:11 sandbox INFO 15:49:11.649 I NET:OPEN ALLOWED ... Now the shorthand is just the severity + structured content: 15:49:11 sandbox INFO I NET:OPEN ALLOWED ... * refactor(ocsf): replace single-char severity with bracketed labels Replace cryptic single-character severity codes (I/L/M/H/C/F) with readable bracketed labels: [LOW], [MED], [HIGH], [CRIT], [FATAL]. Informational severity (the happy-path default) is omitted entirely to keep normal log output clean and avoid redundancy with the tracing-level INFO that the display layer already provides. Before: sandbox INFO I NET:OPEN ALLOWED ... After: sandbox INFO NET:OPEN ALLOWED ... Before: sandbox INFO M NET:OPEN DENIED ... After: sandbox INFO [MED] NET:OPEN DENIED ... * feat(sandbox): use OCSF level label for structured events in log push Set the level field to 'OCSF' instead of 'INFO' for OCSF events in the gRPC log push. This visually distinguishes structured OCSF events from plain tracing output in the TUI and CLI sandbox logs: sandbox OCSF NET:OPEN [INFO] ALLOWED python3(42) -> api.example.com:443 sandbox OCSF NET:OPEN [MED] DENIED python3(42) -> blocked.com:443 sandbox INFO Fetching sandbox policy via gRPC * fix(sandbox): convert new Landlock path-skip warning to OCSF PR #677 added a warn!() for inaccessible Landlock paths in best-effort mode. Convert to ConfigStateChangeBuilder with degraded state so it flows through the OCSF shorthand format consistently. * fix(sandbox): use rolling appender for OCSF JSONL file Match the main openshell.log rotation mechanics (daily, 3 files max) instead of a single unbounded append-only file. Prevents disk exhaustion when ocsf_logging_enabled is left on in long-running sandboxes. * fix(sandbox): address reviewer warnings for OCSF integration W1: Remove redundant 'OCSF' prefix from shorthand file layer — the class name (NET:OPEN, HTTP:GET) already identifies structured events and the LogPushLayer separately sets the level field. W2: Log a debug message when OCSF_CTX.set() is called a second time instead of silently discarding via let _. W3: Document the boundary between OCSF-migrated events and intentionally plain tracing calls (DEBUG/TRACE, transient, internal plumbing). W4: Migrate remaining iptables LOG rule failure warnings in netns.rs (IPv4 TCP/UDP, IPv6 TCP/UDP) to ConfigStateChangeBuilder for consistency with the IPv4 bypass rule failure already migrated. W5: Migrate malformed inference request warn to NetworkActivity with ActivityId::Refuse and SeverityId::Medium. W6: Use Medium severity for L7 deny decisions (both CONNECT tunnel and FORWARD proxy paths) to match the CONNECT deny severity pattern. Allows and audits remain Informational. * refactor(sandbox): rename ocsf_logging_enabled to ocsf_json_enabled The shorthand logs are already OCSF-structured events. The setting specifically controls the JSONL file export, so the name should reflect that: ocsf_json_enabled. * fix(ocsf): add timestamps to shorthand file layer output The OcsfShorthandLayer writes directly to the log file with no outer display layer to supply timestamps. Add a UTC timestamp prefix to every line so the file output matches what tracing::fmt used to provide. Before: CONFIG:VALIDATED [INFO] Validated 'sandbox' user exists in image After: 2026-04-01T15:49:11.649Z CONFIG:VALIDATED [INFO] Validated ... * fix(docker): touch openshell-ocsf source to invalidate cargo cache The supervisor-workspace stage touches sandbox and core sources to force recompilation over the rust-deps dummy stubs, but openshell-ocsf was missing. This caused the Docker cargo cache to use stale ocsf objects from the deps stage, preventing changes to the ocsf crate (like the timestamp fix) from appearing in the final binary. Also adds a shorthand layer test verifying timestamp output, and drafts the observability docs section. * fix(ocsf): add OCSF level prefix to file layer shorthand output Without a level prefix, OCSF events in the log file have no visual anchor at the position where standard tracing lines show INFO/WARN. This makes scanning the file harder since the eye has nothing consistent to lock onto after the timestamp. Before: 2026-04-01T04:04:13.065Z CONFIG:DISCOVERY [INFO] ... After: 2026-04-01T04:04:13.065Z OCSF CONFIG:DISCOVERY [INFO] ... * fix(ocsf): clean up shorthand formatting for listen and SSH events - Fix double space in NET:LISTEN, SSH:LISTEN, and other events where action is empty (e.g., 'NET:LISTEN [INFO] 10.200.0.1' -> 'NET:LISTEN [INFO] 10.200.0.1') - Add listen address to SSH:LISTEN event (was empty) - Downgrade SSH handshake intermediate steps (reading preface, verifying) from OCSF events to debug!() traces. Only the final verdict (accepted/denied) is an OCSF event now, reducing noise from 3 events to 1 per SSH connection. - Apply same spacing fix to HTTP shorthand for consistency. * docs(observability): update examples with OCSF prefix and formatting fixes Align doc examples with the deployed output: - Add OCSF level prefix to all shorthand examples in the log file - Show mixed OCSF + standard tracing in the file format section - Update listen events (no double space, SSH includes address) - Show one SSH:OPEN per connection instead of three - Update grep patterns to use 'OCSF NET:' etc. * docs(agents): add OCSF logging guidance to AGENTS.md Add a Sandbox Logging (OCSF) section to AGENTS.md so agents have in-context guidance for deciding whether new log emissions should use OCSF structured logging or plain tracing. Covers event class selection, severity guidelines, builder API usage, dual-emit pattern for security findings, and the no-secrets rule. Also adds openshell-ocsf to the Architecture Overview table. * fix: remove workflow files accidentally included during rebase These files were already merged to main in separate PRs. They got pulled into our branch during rebase conflict resolution for the deleted docs-preview-pr.yml file. * docs(observability): use sandbox connect instead of raw SSH Users access sandboxes via 'openshell sandbox connect', not direct SSH. * fix(docs): correct settings CLI syntax in OCSF JSON export page The settings CLI requires --key and --value named flags, not positional arguments. Also fix the per-sandbox form: the sandbox name is a positional argument, not a --sandbox flag. * fix(e2e): update log assertions for OCSF shorthand format The E2E tests asserted on the old tracing::fmt key=value format (action=allow, l7_decision=audit, FORWARD, L7_REQUEST, always-blocked). Update to match the new OCSF shorthand (ALLOWED/DENIED, HTTP:, NET:, engine:ssrf, policy:). * feat(sandbox): convert WebSocket upgrade log calls to OCSF PR #718 added two log calls for WebSocket upgrade handling: - 101 Switching Protocols info → NetworkActivity with Upgrade activity. This is a significant state change (L7 enforcement drops to raw relay). - Unsolicited 101 without client Upgrade header → DetectionFinding with High severity. A non-compliant upstream sending 101 without a client Upgrade request could be attempting to bypass L7 inspection. |
||
|
|
574ef18dfc |
docs: consolidate information architecture and author content (#124)
* initial doc filling * improvements * stage provided get started, and add clean tutorials * pull in Kirit's content and polish * improve observability * moving pieces * move TOC around * drop support matrix from concepts * fix links * minor fixes and fix badges * minor fixes * incorporate missed content * minor improvements * clean up * run dori style guide review * clean up * updates impacting docs * incorporate feedback * minor fix * some edits * enterprise structure * update cards * improve * add some emojis * improve landing page with animated getting started code * fix the animated code * small improvements * refresh content based on PR 156 and 158 * README as the source of truth for quickstart * update README * run edits * change to the new prod name, text only, code swipe later * add Home * incorporate dev feedback on README * krit's edits * edit and improve index pages * fix build * revert README * revert quickstart to not pull from README |
||
|
|
11f795a462 |
docs: setup initial docs/ infrastructure and scaffolding (#94)
* set up docs * rm nv sphinx theme version * rm v* trigger * incorporate feedback with cursor * minor updates * doc docs * more small tweaks * clean contributing --------- Co-authored-by: Drew Newberry <anewberry@nvidia.com> |