Commit Graph
31 Commits
Author SHA1 Message Date
krishicks 8091f66877 feat(sandbox): write agent output to the container log (#4005)
Since #2726 the canonical main process's stdout and stderr are captured
in pipes that feed only the in-memory replay buffer used by sandbox
connect. Agent output therefore never reaches the container's own stdout
and stderr, so it is missing from kubectl logs, docker logs, and podman
logs and from anything that collects container logs. Before #2726 the
entrypoint inherited the container's descriptors and its output appeared
there.

Copy the main process's output to the launcher's stdout and stderr in
addition to the replay buffer, restoring the earlier behavior:

- Output is copied byte for byte to the matching stream from a
  forwarder thread per stream, after it is published to the replay
  buffer. When the container runtime falls behind on a stream, that
  stream's reader waits instead of dropping output, so backpressure
  reaches the agent as it did with inherited descriptors, while the
  other stream and attachments keep receiving output.
- Before the main process's exit is published, the output readers
  finish and queued output is drained to the container log, so an
  agent's final lines are not lost at shutdown. A 30 second deadline
  covers both; when it expires, readers waiting on the container log
  are released and drain the pipes into the replay buffer only, so a
  stalled container log cannot block exit reporting.
- PTY-mode processes are not copied. The terminal stream carries escape
  sequences and echoed input, and terminal commands never reached the
  container log before #2726.
- Exec, SSH, and SFTP sessions are not copied.

Launcher log lines keep their existing format and remain in the
container's stderr. They are written as whole lines, and a newline is
inserted first when the agent left stderr mid-line, so launcher and
agent lines do not merge.

The Docker and VM drivers appended the tail of the workload's output to
failure messages: Docker the workload container's log, and the VM driver
the guest console, which carries the launcher's stdout and stderr. Those
messages land in the sandbox's Ready condition and in platform events
that the gateway republishes to the sandbox event stream. With agent
output in that log, those messages would carry arbitrary agent output,
including anything sensitive the agent prints, into gateway status and
events. The supervisor starts its health endpoint only after the agent
starts, so every Docker failure path could include agent output, and the
VM driver reports one whenever the VM or host supervisor exits. Forward
only the supervisor's log tail, matching the Podman driver, which reads
the workload log solely to match fixed launcher markers and never
forwards raw workload output. The workload's output remains available
through docker logs and the VM's rootfs-console.log.

Document where main process output appears in the logging docs and the
cluster debugging skill.

Closes #3928

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-10-01 17:23:09 +00:00
Simon Scatton e21b7fd8cf chore(build): remove bundled Z3 support (#3275)
* chore(build): remove bundled Z3 support

Signed-off-by: Simon Scatton <sscatton@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(build): preserve vendored Z3 for local gateway artifacts

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

---------

Signed-off-by: Simon Scatton <sscatton@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-10-01 12:08:33 +00:00
krishicksandEvan Lezar a875add234 feat(server): write gateway OCSF events to JSONL (#3264)
* feat(server): write gateway OCSF events to JSONL

Previously, gateway security activity was available only in diagnostic output,
and events not associated with a sandbox, such as TLS certificate reloads, had
no independent structured record.

Now, configuring `openshell.gateway.ocsf_log` writes every gateway-produced
OCSF record to a bounded JSONL destination independently of `RUST_LOG`. The
destination supports daily or disabled rotation, retention limits, queue bounds,
and optional schema downgrade to OCSF 1.1 or 1.3.

Additionally, existing gateway emitters (TLS reloads, service routing, and
policy approval and auto-approval audits) emit structured events, so they reach
the JSONL destination, console shorthand, and the affected sandbox's log
stream. Records identify the gateway by its configured name in `device.uid` and
`device.name`, shared across replicas, with `device.hostname` identifying the
replica and `device.os` the gateway's operating system. Metrics and warnings
expose known best-effort losses.

Refs #2762

Signed-off-by: Kris Hicks <khicks@nvidia.com>

* fix(mxc): attribute ETW events to gateway

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Kris Hicks <khicks@nvidia.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>
Co-authored-by: Evan Lezar <elezar@nvidia.com>
2026-09-29 16:19:53 +00:00
Shiju 45e3308d39 fix(network): refuse protocol upgrades on JSON-RPC and MCP endpoints (#3753)
* fix(network): refuse protocol upgrades on JSON-RPC and MCP endpoints

JSON-RPC and MCP rules apply to each HTTP request, but the proxy could
forward a request that also carried upgrade headers. After an upstream
answered 101, route selection and the forward proxy relayed the
connection without inspection.

Refuse any request that carries an Upgrade header on JSON-RPC-family
endpoints before the L7 policy decision, in every enforcement mode.
Share the check with the existing h2c refusal and call it from
relay_jsonrpc as well. Record the refusal as a policy denial and answer
with the unsupported_l7_protocol error, because no policy rule can
allow the request.

If a JSON-RPC-family endpoint still receives 101, close the connection
instead of relaying raw bytes. Document the refusal and the WebSocket
alternative.

Signed-off-by: Shiju <shiju@nvidia.com>

* docs(observability): remove duplicate protocol error definition

Keep unsupported_l7_protocol in the response error-code list and retain its explanation in the policy troubleshooting table.

Signed-off-by: Shiju <shiju@nvidia.com>

---------

Signed-off-by: Shiju <shiju@nvidia.com>
2026-09-28 16:08:19 +00:00
Drew Newberry 2f1ea658fe docs(policy): clarify sandbox-local loopback access (#3740)
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-26 18:18:28 +00:00
Drew NewberryandPiotr Mlocek 73a181d32f docs: streamline README, add policy prover to architecture docs (#3718)
* docs(readme): streamline README and move reference detail to docs

Restructure the README as a short path from overview to quickstart to
further reading. Move prerelease install steps into the installation
guide and telemetry build flags into a new observability page. Fix
broken docs links and outdated runtime and credential descriptions.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(readme): describe 0.1.0 as adding new isolation primitives

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(architecture): add policy prover as a gateway component

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs: describe OpenShell as a runtime for fleets of autonomous AI agents

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(architecture): describe policy prover as formal verification

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(architecture): name OpenShell Sandbox in component table

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(architecture): fold isolation backend into supervisor row

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(architecture): mention formal verification in overview

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(run-agent): use the published OpenCode image in the first-agent guide

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(inference): correct provider examples and readiness

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(providers): correct Google binding and provider selection

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(readme): sharpen value prop, how it works, and explore further

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(run-agent): use example Anthropic profile and add policy advisor step

Import the example Anthropic profile, which now allows OpenCode, instead
of editing it with sed. Add a step that shows how to review and approve
mechanistic policy proposals as the agent needs more access.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(providers): add OpenRouter example for OpenCode

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(run-agent): run OpenCode against OpenRouter with a free model

Add an example OpenRouter provider profile scoped to OpenCode and switch
the first-agent guide to it, using a free Nemotron model so readers do
not need OpenRouter credits. Revert the OpenCode binary added to the
example Anthropic profile.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(readme): link first-agent guide and add agent skills section

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Co-authored-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-09-25 13:53:45 -07:00
Johnny Greco d7f921190b docs(policy): refresh policy documentation and references (#3563)
* docs(policy): correct schema and default policy guidance

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): add network recipes and update command reference

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): organize lifecycle guidance and troubleshooting

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): split policy overview into concepts and management tasks

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): reorganize network recipes as a cookbook

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): restructure schema reference by field group and protocol

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): align troubleshooting, advisor, and reference pages

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): fix first policy tutorial and security guidance

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): keep overview high level and move network rules to their own page

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): focus policy management on CLI workflows and remove command reference

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): clarify policy views and sandbox deletion in management guide

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): streamline network rule concepts and examples

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): correct request path wildcard semantics

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): rewrite policy advisor guide for clarity

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): clarify policy advisor scope, setup, and review

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): rewrite policy prover guide for clarity

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): explain the two uses of the policy prover

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): describe policy prover uses, boundaries, and coverage

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): place prover before advisor and troubleshooting last

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): remove unsupported CI guidance from prover page

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): tighten policy prover introduction

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): move policy change behavior into management guide

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): name prover check types and note expanding coverage

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): prefix prover and advisor sidebar labels

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): streamline policy schema reference

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): place default policy before schema reference

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): fold troubleshooting into policy management guide

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): correct tutorial log samples and GitHub push policy steps

The first policy tutorial said the 403 body begins with error, policy, and
rule, but the proxy serializes the body with sorted keys. Its log samples also
showed the wrong CONNECT deny reason for a sandbox without network rules, and
the L7 deny sample omitted the :443 authority, the `l7` engine, and the reason
tag that the shorthand formatter emits.

The GitHub tutorial filtered denials with `--level warn`, which hides the INFO
level OCSF policy events, and showed the retired key=value log format. Its
hand-written policy also omitted /bin from the restrictive default, so
`policy set` would reject the file for removing a filesystem path on a live
sandbox. Start from `policy get --base` and add only the network rules.

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): improve flow and terminology across policy pages

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): correct network rule matching and protocol details

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): align policy management steps with CLI behavior

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): correct policy advisor proposal and approval details

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): correct policy section, default, and schema details

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): correct prover installation and coverage limits

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): recommend tls skip for server-first protocols

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): fix stale baseline path and interpreter examples

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): move policy pages under how-it-works and fix links

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): align native TCP guidance in security best practices

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): restore policy.local and policy DNS details from main

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): state exact glob matching rules

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

---------

Signed-off-by: Johnny Greco <jogreco@nvidia.com>
2026-09-25 18:19:16 +00:00
Drew Newberry 9244868056 docs: refresh architecture and agent guides (#3705)
* docs: refresh architecture and agent guides

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs: describe updated security architecture neutrally

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs: highlight new isolation primitives

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(sandboxes): clarify how to disconnect

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs: align architecture and guides with current navigation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(extensibility): streamline extension authentication guidance

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-25 02:38:00 -07:00
Artem LytvynandJohn Myers 470a34635d fix(api): make WatchSandbox loss-aware and resumable (#3209)
* fix(api): emit warning on WatchSandbox broadcast lag instead of terminating

Broadcast lag on the status, log, and platform receivers was converted to a RESOURCE_EXHAUSTED status that terminated the whole watch stream. Lag is recoverable: the receiver resumes at the oldest surviving message. Emit a SandboxStreamWarning and continue streaming instead; keep terminating on Closed. Add helpers and unit tests covering the warning payload and receiver recovery after lag.

Partially addresses #3055 (cursor/resume follow up separately).

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* refactor(server): group per-sandbox log bus state and stamp sequence numbers

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* feat(proto): add resume cursor fields to sandbox watch API

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* feat(server): stamp watch cursors from a shared per-sandbox sequence

Allocate cursors from a single SeqAllocator shared by the log and
platform event buses, so a sandbox's merged watch stream carries
unique, strictly increasing cursors. A single resume_after_cursor can
then unambiguously locate a client's position across both sources.

Rewrite both publish paths to allocate the sequence, stamp
event.cursor, send, and append to the tail under one lock. This
removes the previous get_mut().expect() TOCTOU race where a concurrent
remove() between the two lock sections could panic.

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* feat(server): serve WatchSandbox resume from cursor with gap detection

Add tail_after() to the log and platform event buses, returning every
buffered event newer than a client's resume cursor. Each PerSandbox now
tracks last_trimmed_seq (the highest seq it has evicted) so a resume is
reported as an unrecoverable ResumeGap only when this bus dropped an
event the client still needs.

Judging gaps by evictions, not by the tail's oldest seq, is required
under the shared cursor space: each bus's tail is non-contiguous in the
global sequence because the other bus owns the missing seqs, so
comparing against tail.front() would flag false gaps.

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* feat(server): resume WatchSandbox from cursor across log and platform buses

Wire resume_after_cursor into the watch producer. On a non-zero cursor,
replay events strictly after it from both the log and platform buses,
merge by shared cursor, and emit in order before entering the live loop.
A trimmed range on either bus is an unrecoverable gap and terminates the
stream with OUT_OF_RANGE carrying the requested and earliest-available
cursors, distinct from recoverable lag which warns and continues.

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* test(server): cover WatchSandbox cursor resume paths

Add handler-level tests for the resumable watch stream: replay strictly
after the client cursor, merge log and platform events in shared-cursor
order, suppress duplicates when resuming at the latest cursor, and
terminate with OUT_OF_RANGE when the requested cursor has been trimmed.

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* docs(api): document WatchSandbox loss-awareness and resume

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* fix(server): deliver watch events once and harden cursor teardown

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* feat(sdk): add loss-aware resumable watch_logs to Rust SDK client

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* fix(server): keep watch cursors monotonic across teardown and restart

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* fix(server): merge live watch sources by cursor before emission

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* fix(api): bind watch cursors to a cursor space and merge tail sources

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* fix(server): revalidate the watch cursor space after collecting replay

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* test(server): update the public RPC schema fingerprint for the string cursor

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* fix(server): hold watch events above the publication watermark and emit the watch lag warning before its batch

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* test(server): synchronize the watch live-order test with the end of initialization

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* fix(sdk): use canonical sandbox name in watch_logs

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* test(sdk): guard canonical-name addressing in watch_logs

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* fix(server): fix public rpc schema

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* fix(api): reconcile watch resume rebase

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(server): bound interactive relay cleanup

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-09-22 19:45:55 +00:00
krishicks cbf026366d fix(ocsf): require network activity endpoints (#3355)
Previously, Network Activity could be constructed without a source or
destination endpoint, allowing connection, accept, relay, and configuration
events to violate the OCSF 1.8 endpoint constraint.

Now, NetworkActivityBuilder requires a source or destination endpoint at
compile time. Connection failures identify the workload peer or genuine
transparent destination, listener failures identify the listening endpoint,
and mediation-lane failures use Application Lifecycle rather than fabricated
network endpoints. Malformed forward requests use HTTP Activity with a
method-only request, generated 400 response, and workload peer.

Additionally, Unix relay-channel events use Base Event, policy-validation
warnings use Config State Change, and the unused bypass monitor is removed
because the current isolation architecture no longer uses it.

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-09-21 17:13:41 +00:00
krishicks 481ce566e1 fix(ocsf): correct HTTP activity context (#3316)
Previously, metadata events omitted both HTTP request and response objects,
early proxy rejections used HTTP Activity without request context, and
unsupported-scheme events did not expose enough safe HTTP context to satisfy
the OCSF 1.8 schema.

Now, metadata events include a method-only request and their actual HTTP
response codes without recording the metadata URL. Unsupported-scheme events
also include a method-only request plus the generated 400 response. Authority
mismatches and credential-resolution denials use HTTP Activity with their
generated 403 or 500 responses, and HTTP activity IDs are derived from the
request method.

Additionally, HttpActivityBuilder now enforces the OCSF request-or-response
constraint at compile time.

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-09-15 18:47:55 +00:00
ae57979b03 feat(mxc): add Windows ETW-to-OCSF audit trail (#3015)
* feat(mxc): ETW->OCSF audit consumer + Windows OCSF JSONL parity (cp6 P1)

Add a Windows MXC ETW->OCSF audit trail in openshell-driver-mxc: a real-time
Sandboxing-provider ETW consumer that decodes events (TDH), attributes each to
an OpenShell sandbox_id, and maps them to OCSF (lifecycle 6002, config 5019,
process 1007, finding 2004).

cp6 Phase 1 - durable OCSF JSONL audit-file parity with Linux:
- openshell-ocsf: add emit_ocsf_event_routed (populates the event-bridge
  thread-local AND stamps sandbox_id+message in one dispatch) plus public
  set/clear_current_event; OS-aware device (Device::windows/for_current_os) so
  device.os.name reflects the host instead of a hardcoded Linux stub.
- etw_consumer: emit via the routed emit (previously fired a bare info! that
  never populated the bridge, so the structured event was dropped).
- openshell-server: install OcsfJsonlLayer over a synchronous daily-rotated
  appender (durable under force-kill), gated by OPENSHELL_OCSF_JSON, path via
  %PROGRAMDATA%\OpenShell\logs (override OPENSHELL_OCSF_LOG_DIR).
- device.hostname now resolves to the real gateway machine name.

Box-proven on 7F203-MXC-001: JSONL lines == shorthand OCSF rows, all valid
OCSF JSON, per-sandbox attribution intact, disabled state writes nothing.

Signed-off-by: Akber Raza <akberr@nvidia.com>

* feat(mxc): map remaining Sandboxing ETW events to OCSF

Close the last three ETW->OCSF gaps so the audit trail covers the full
set of events the Sandboxing provider emits (12/12):

- ProcessLaunched -> Process Activity [1007] "Launch" (confirmed start;
  carries the real processId/threadId, the twin of CreateProcessInSandbox
  which only has the request + command line).
- SandboxProxyConfigured -> Device Config State Change [5019] (the one
  network-plane setup event; surfaces proxyPort, "no proxy" when 0).
- SandboxConsoleReferencePlumbed -> Device Config State Change [5019]
  (console-handle plumbing).

map_config_state now handles the full config/hardening/setup family and
carries proxyPort/hasConsoleReference/creationFlags as unmapped fields.
Verified on 7F203-MXC-001: 11/12 event types emit OCSF without a proxy
(SandboxProxyConfigured requires proxy config to fire).

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc): seed ETW attribution under registry lock + Device tests

Address CodeRabbit review on !31:

- Prevent stale ETW attribution on a delete/launch race: register the
  wxc-exec pid while holding the registry lock, and bail if the sandbox
  entry is already gone. Previously the attribution key could be seeded
  after `delete` had removed the sandbox, leaving a stale key that could
  misroute later Sandboxing ETW events to a dead sandbox_id. Lock order
  (registry -> attribution) matches the delete path, so no deadlock.
- Add unit tests for the new Device::windows and Device::for_current_os
  constructors to harden Windows/Linux OCSF device parity.

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc-etw): buffer+replay racing events and harden attribution keys

Addresses two ETW->OCSF attribution review items (Shailendra #1, #2).

#2 early-event loss: ETW delivers the sandbox create/config burst the instant wxc-exec starts, which can beat the driver's register_launch (now under the registry lock post-Ready). process_event previously dropped anything unresolved, losing the racing burst. Add a bounded, time-bounded pending buffer (PENDING_MAX=4096, PENDING_TTL=5s): unresolved events are held and replayed once attribution lands, aged-out ones dropped. Consumer switched to a timed recv_timeout(200ms) so the buffer is re-driven after each event and on a tick. Emit path factored into shared emit_resolved().

#1 attribution collisions: a Windows PID is recycled after exit and a command line is commonly identical across sandboxes. register_launch now rebinds by_pid on reuse and clears the stale last_pid_sid hint (warns if the PID still pointed at a different, leaked sandbox); command line is held in by_cmd only while unique and demoted to a new ambiguous_cmds set on a second owner, so a duplicate command refuses to resolve rather than misroute.

Unit tests: buffer replay (direct + cross-link), buffer bound, PID-reuse rebind, duplicate-cmd non-resolution. Box-verified on 7F203-MXC-001 (5 sandboxes, identical cmd -> 5 isolated sandbox_ids, 50/50 OCSF/JSONL, BuffersLost=0).

Signed-off-by: Akber Raza <akberr@nvidia.com>

* docs(mxc-etw): note cmd_line is captured raw with no privacy filtering

Review item #3 (Shailendra): add a PRIVACY NOTE on map_process_launch stating cmd_line is copied verbatim into OCSF process.cmd_line with no redaction, so secrets/PII on a command line land unredacted in the durable audit trail (deliberate audit-fidelity trade-off; treat the log as sensitive). Redaction is owned by an upstream privacy layer, not this path; no general audit-output PII scrubber exists today (openshell_core::secrets [CREDENTIAL] redaction is scoped to the proxy HTTP-target logging, a separate egress path).

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc-etw): open ETW trace on caller thread so start_session reports real status

Review item #4 (Shailendra): start_session previously returned Ok(EtwSession) as soon as the pump thread was spawned, but OpenTraceW ran later inside that thread; if it failed we still handed back a live-looking session and logged 'consumer started' (silent failure = false audit coverage).

Split the two Win32 calls instead of adding a channel handshake (avoids any lost-wakeup/hang risk): the quick, synchronous OpenTraceW now runs on the caller thread (open_trace), and only the blocking ProcessTrace runs on the pump thread (run_trace). start_session returns Err if OpenTraceW fails (reclaiming the boxed Sender so the consumer disconnects, stopping the session, joining the consumer) and returns Ok/logs 'started' only once capture is genuinely open. Opened handle + LoggerName buffer + boxed Sender are carried to the pump via a Send OpenedTrace so they outlive ProcessTrace.

Box-verified on 7F203-MXC-001: consumer started=True, failed-to-start=False, 50 OCSF rows / 50 JSONL, BuffersLost=0 (no regression to capture/emit).

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc-etw): guard pending-event replay against PID recycling

CodeRabbit flagged that drain_resolved() re-resolved buffered events
against the live by_pid map, so if Windows recycled a wxc-exec PID within
PENDING_TTL a stale event from the dead sandbox could be emitted under the
new owner.

Stamp each by_pid registration with its Instant and add resolve_replay(),
used only on the buffered/replay path. It (a) never falls back to the
recycle-/ambiguity-prone by_cmd or last_pid_sid keys, and (b) trusts a PID
match only when the registration is not newer than the buffered event by
more than REPLAY_PID_GRACE (2s) - a recycled PID's registration lands well
outside that window, so the stale event ages out instead of misattributing.
The legitimate #2 seed race (registration lands ~immediately) still replays.

Adds unit tests for the recycle-refusal, in-grace acceptance, and
weak-fallback exclusion.

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc-etw): surface unexpected ProcessTrace termination (review #4)

start_session already returns Err on OpenTraceW failure (runs on the
caller thread since e41a7701), closing the first half of Shailendra's #4.
This closes the second half: ProcessTrace's result was discarded, so if
capture died mid-run the backend had no way to know.

Add a shared CaptureHealth (stopped/stopping/exit_code) between the pump
thread and EtwSession. run_trace now records ProcessTrace's WIN32_ERROR
and, when the pump returns without a deliberate stop, logs at ERROR that
MXC OCSF capture is no longer running. EtwSession::stop() sets `stopping`
before teardown so a normal shutdown isn't misreported, and
EtwSession::is_capture_alive() exposes the state for status/diagnostics.

Box-verified on 7F203-MXC-001: 5 sandboxes, 50 attributed OCSF rows,
JSONL parity 50/50, BuffersLost=0, clean start/stop (no false failure).

Signed-off-by: Akber Raza <akberr@nvidia.com>

* feat(mxc-ocsf): add ETW->OCSF audit-trail example kit; fix proxy-configured message

Add a runnable OCSF audit-trail example under examples/ (run-ocsf-audit.ps1,
mxc-ocsf-audit.toml, ocsf-audit.yaml, README) that spins up sandboxes with the
in-process ETW consumer and egress proxy on, emitting a full OCSF JSONL audit
trail across all four classes (6002/5019/1007/2004).

Fix SandboxProxyConfigured mapping to log "MXC sandbox proxy configured" instead
of a misleading "(no proxy)" when the provider reports proxyPort=0; the event's
presence already indicates proxy configuration. Verified on-box: 26 events, all
mapped ETW event types present.

Signed-off-by: Akber Raza <akberr@nvidia.com>

* feat(mxc-ocsf): clearer audit report + client-safe run-ocsf-audit.ps1

Improve the ETW to OCSF audit-trail example output and make it safe to ship.

Report:
- Add an event-type coverage count ("N of M expected event types fired");
  the denominator auto-adjusts (8 with proxy on, 7 with -NoProxy).
- Split the checklist into expected event types vs anomaly findings
  (ActivityError/FallbackError), which are reported separately and not
  counted toward coverage (a clean run may emit none).
- Verdict is now coverage-based (all expected types must fire) instead of
  the looser "at least 3 OCSF classes".
- Call out the absolute path to the durable OCSF JSONL log prominently.

Client-safety:
- Default -ShareOut to empty (no auto-copy); pass -ShareOut a UNC path to
  opt in. Removes a hardcoded internal share path from a published example.
- Drop internal-team wording ("Hand that zip back for evaluation", "BUNDLE:")
  in favor of neutral "Results bundle:".
- Update README-ocsf-audit.txt to match the opt-in -ShareOut behavior.

Verified on both MXC boxes: 7F203-MXC-001 (base-container) -> PASS, 8 of 8
event types, 26 OCSF events across 4 classes; 7F203-MXC-003 (AppContainer
fallback) -> reduced set as expected, clean output.

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc): configure OCSF audit workloads per sandbox

- remove unsupported gateway-scoped workload fields from the shipped MXC audit example.
- build the command, working directory, and filesystem grant from each run's ShareDir
- pass the workload through --driver-config-json.
- preserve the host CONNECT proxy configuration and conditional audit coverage for the future host_connect_proxy merge
- require the workload output when determining the audit verdict.

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc): omit command arguments from OCSF audit logs

- record only the executable basename for MXC CreateProcessInSandbox audit events
- leave process.cmd_line unset so workload arguments cannot reach shorthand or JSONL logs
- cover tokens, passwords, signed URLs, and PII with a secret-leak regression test
- update the audit example, architecture guidance, and published logging documentation
- preserve ETW attribution and future host_connect_proxy enforcement behavior

Signed-off-by: Akber Raza <akberr@nvidia.com>

* feat(etw): enhance ETW session management with distinct naming for concurrent gateways

* fix(etw): bound the audit queue during overload

- replace the unbounded ETW callback channel with count- and byte-bounded buffering

- keep the ETW callback non-blocking and count records rejected during overload

- emit immediate, rate-limited warnings that identify resulting audit coverage gaps

- make the audit example fail when queue overload causes dropped ETW records

- cover stalled consumers, oversized events, and warning throttling with unit tests

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(etw): harden sandbox audit attribution

- remove command-line and persistent per-PID fallback keys from live and replay resolution

- retire the driver-owned wxc-exec PID before publishing child completion

- retain established identity, activity, and correlation-vector links only for the five-second late-event window

- prevent buffered records from crossing rapid PID retirement and reuse boundaries

- add resolver and lifecycle coverage and document the attribution trust boundary

Signed-off-by: Akber Raza <akberr@nvidia.com>

* chore(mxc): address rebase follow-ups

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(mxc): align OCSF audit example with driver config

- remove unsupported egress proxy settings

- stop requiring the unavailable proxy audit event

- update example documentation for supported event coverage

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(etw): redact command-line secrets in DecodedEtwEvent summary

* fix(etw): enhance PID resolution and event attribution logic for ETW records

* fix(ocsf): restrict gateway-local JSONL sink to Windows/MXC path with opt-in configuration

* address rebase issues

* fix(mxc): address ETW audit review feedback

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(mxc): fail closed across ambiguous PID reuse

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(mxc): bind ETW attribution to process generation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Akber Raza <akberr@nvidia.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Jamie King <jamiek@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
2026-09-11 02:04:26 +00:00
krishicks 67374efdf8 fix(ocsf): emit schema-valid event identities (#3247)
Previously, every event from a sandbox reused the sandbox ID as its event
ID. Consumers deduplicating security records could mistake separate events
for the same record, and missing device types or empty image objects could
prevent schema validation.

Give each event its own ID, retain the sandbox association separately, and
classify the environment as Other/Sandbox while keeping the OS separate.
Omit unknown container details instead of emitting empty objects.

Security tooling can now distinguish events from the same sandbox and read
their identity consistently after serialization.

Refs #1055

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-09-10 16:35:24 +00:00
John T. Myers f4dc6be4b2 refactor(inference): remove managed inference routes (#3195)
* refactor(inference): remove managed inference routes

Closes #3172

Remove the inference route control plane, inference.local data path, built-in router crate, and SDK surface. Move inference workloads to explicitly imported provider profiles and native endpoints, with migration cleanup and updated tests and documentation.

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(policy): preserve alternate upstream isolation

Restore the provider policy activation guard so legacy OpenAI and Anthropic providers configured for alternate base URLs do not grant egress to the built-in public vendor endpoints.

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-09-09 18:47:22 +00:00
Adel ZaaloukandJohn Myers 48a8a4bf09 feat(ocsf): configurable schema version for SIEM backward compatibility (#2717)
* feat(ocsf): configurable schema version for SIEM backward compatibility

Add a gateway-configurable OCSF schema version target that downgrades
JSONL output for SIEMs that only support older schema versions. AWS
Security Lake requires v1.1.0, Splunk CIM Add-On targets v1.1-v1.3.

The downgrade filter strips profile-gated fields (ai_model, container,
observation_point_id), removes unknown profiles from metadata.profiles,
rewrites metadata.version, and adds an unmapped.downgraded_from
breadcrumb so auditors can distinguish "no model involved" from "model
attribution stripped."

Supported target versions (1.1, 1.3) are enforced by an allow-list in
the settings registry. Invalid values are rejected with a clear error.

The setting flows to sandboxes via the settings bundle and takes effect
on the next poll cycle. The shorthand log output is unaffected.

Closes #2662

Signed-off-by: Adel Zaalouk <zanetworker@gmail.com>
Signed-off-by: Adel Zaalouk <azaalouk@redhat.com>

* fix(ocsf): align downgrade with schema version

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: Adel Zaalouk <zanetworker@gmail.com>
Signed-off-by: Adel Zaalouk <azaalouk@redhat.com>
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-09-04 18:43:58 +00:00
Drew Newberry ef296806f5 feat(sandbox): add canonical main process (#2726)
* feat(sandbox): add canonical main process

Closes #2710

Persist and supervise one canonical workload per sandbox, attach sandbox connect to its retained session, and make every unexpected main-process exit terminal.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(sandbox): simplify canonical main process contract

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): preserve legacy VM main compatibility

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): preserve main status across driver updates

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): satisfy macOS process lint

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): gate Linux exit acknowledgement publisher

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(sandbox): simplify main process plumbing

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(supervisor): make controlling tty ioctl portable

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(supervisor): initialize canonical process environment

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* perf(supervisor): optimize retained main session

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(sandbox): detach main session on ctrl-c

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): use explicit main detach keys

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(dev): atomically stage Docker supervisor

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(test): align Docker main environment assertion

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(sdk): expose canonical main process fields

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-08-20 22:18:23 +00:00
Adel ZaaloukandClaude Opus 4.6 600bbae845 feat(ocsf): emit AI inference events via ai_operation profile on ApiActivity (#2664)
Apply the official OCSF ai_operation profile (introduced in v1.8.0) to
ApiActivity [6003] events when the inference proxy routes a model call
through inference.local. Attaches an ai_model object (name, ai_provider)
and puts token counts and latency in unmapped fields.

ApiActivity [6003] is the schema-correct class for the ai_operation
profile in v1.8.0 (HttpActivity only gets it in v1.9.0). In Splunk CIM,
ApiActivity maps to the "Change" data model, naturally separating
inference events from regular HTTP proxy traffic.

Changes:
- Add AiModel object and ai_model field on BaseEventData
- Add ApiActivityEvent struct and ApiActivityBuilder
- Add emit_ai_inference in proxy.rs using ApiActivity with ai_operation
- Vendor OCSF v1.8.0 schemas including api_activity class, ai_model
  object, and ai_operation profile definitions
- Bump OCSF_VERSION to 1.8.0
- Update schema validation to skip profile-gated required fields

Shorthand: API:INFERENCE [INFO] claude-3-haiku via anthropic 701ms [POST /v1/messages]

Splunk/SIEM backward compatibility (v1.1/v1.3 CIM mapping) is tracked
separately in #2662 as a configurable serialization concern.

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-08-18 16:30:03 +00:00
Piotr Mlocek 44bf0df485 feat(middleware): inspect WebSocket text messages (#2477)
* feat(middleware): inspect websocket text messages

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): address websocket review feedback

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(network): bound websocket message assembly

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(network): harden websocket upgrade lifecycle

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* refactor(middleware): unify in-process and remote transports

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* feat(middleware): support regex websocket redaction

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): bound persistent streaming sessions

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): accept websocket sequence gaps

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* refactor(middleware): refine websocket introspection contract

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): clarify websocket preflight lifecycle

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): clarify websocket coverage semantics

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(network): type websocket frame failures

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(network): return 503 when middleware admission is exhausted

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): align streaming API contract

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): clarify WebSocket event result scope

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(rfc): simplify middleware revision history

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* feat(examples): add WebSocket content guard support

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): unify binding payload limits

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* refactor(middleware): align payload limit terminology

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): address websocket review feedback

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(network): address websocket review findings

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* test(network): allow Linux handler setup in preflight regression

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(network): harden websocket relay finalization

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(network): inspect compressed websocket messages

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* test(network): stabilize compressed websocket regressions

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(go-sdk): regenerate middleware protobuf binding

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): clarify websocket skip lifecycle

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): address WebSocket review feedback

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

---------

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-08-14 21:51:42 +00:00
Adel Zaalouk 472e23f96a fix(proxy): include OPA deny reason in CONNECT 403 response (#2363)
* fix(proxy): include OPA deny reason in CONNECT 403 response

When a CONNECT request was denied by OPA policy, the 403 response
used a generic "not permitted by policy" message for both "endpoint
not in policy" and "endpoint matched but binary didn't match." Users
had no way to distinguish the two without reading supervisor logs.

The OPA policy already computes a detailed deny_reason (e.g.,
"binary '/usr/bin/node' not allowed in policy 'X'") but the proxy
was not including it in the HTTP response.

Now the CONNECT deny response includes a "reason" field with the
OPA deny reason when available. When the reason is empty, the field
is omitted for backward compatibility.

Fixes #2355

Signed-off-by: Adel Zaalouk <azaalouk@redhat.com>

* docs(observability): document optional reason field in CONNECT 403 response

The proxy now includes a reason field in the JSON body of denied
CONNECT responses when the policy engine provides a specific denial
cause. Update the Proxy Error Responses section to show the field
and describe when it is present vs omitted.

Signed-off-by: Adel Zaalouk <azaalouk@redhat.com>

---------

Signed-off-by: Adel Zaalouk <azaalouk@redhat.com>
2026-07-21 16:14:32 +00:00
Piotr Mlocek d556748771 feat(supervisor-middleware): add network egress middleware (#2027)
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-07-16 17:47:49 -07:00
Alexander Watson e98ea3ee93 feat(policy): add agentic approval loop (#1528) 2026-05-29 21:11:31 -07:00
Miyoung Choi 8322e4fd00 docs: style fixes (#1341)
* docs: style fixes

* docs: drop observability section overview page and rename a section title

* docs: title updates
2026-05-12 16:36:15 -07:00
John T. Myers 0914f3f4f7 fix(sandbox): log L7 parse denials (#1072)
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-04-29 16:15:37 -07:00
Piotr Mlocek a6d45528c1 feat(server,sandbox): supervisor-initiated SSH connect and exec over gRPC-multiplexed relay (#867) 2026-04-21 08:38:18 -07:00
John T. Myers 40e9bf6feb feat(policy): add incremental sandbox policy updates (#860)
* feat(policy): add incremental sandbox policy updates

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(policies): expand incremental update guidance

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* feat(policy): audit incremental updates in gateway logs

* docs(policy): quote glob specs in shell examples

---------

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-04-20 08:00:02 -07:00
Miyoung Choi b7c763204d docs: refresh user-facing docs for recent sandbox and inference changes (#868)
* docs: refresh user-facing docs for recent sandbox and inference changes

- architecture: document system CA loading for upstream TLS, `tls: skip`
  as the opt-out, gateway state persistence across restarts, and OCSF
  structured logging surface.
- inference: document per-provider header allowlist, Authorization
  stripping, 120s streaming idle tolerance, and extended-thinking
  timeout guidance.
- manage-sandboxes: add "Execute a Command in a Sandbox" section for
  `openshell sandbox exec` with flag reference.
- security best practices: expand seccomp denylist (unconditional and
  conditional blocks), document two-phase Landlock probe, High-severity
  `landlock-unavailable` finding, and inference keep-alive closure.
- observability logging: document port in HTTP log URLs, `[reason:...]`
  denial suffixes, proxy 403/502 JSON error bodies, and Landlock
  CONFIG:ENABLED/CONFIG:OTHER events.

Signed-off-by: Miyoung Choi <miyoungc@nvidia.com>

* docs(architecture): tighten high-level architecture page

Trim implementation detail (CA bundle paths, deprecated TLS keys, SSH
handshake secret) from the high-level architecture page, fix accuracy
issues surfaced during deep audit, and expand uncommon acronyms on first
mention.

- Drop unsupported "cost-based routing" claim from Privacy Router row.
- Replace "brokers requests across the platform" with auth-boundary
  description.
- Add "inference" to Policy Engine constraint list per AGENTS.md.
- Expand Deny rule to include SSRF, blocked control-plane port, and L7
  deny paths in addition to deny-by-default.
- Switch Allow/Deny labels from hyphen to colon; remove em dashes and a
  double space.
- Expand LLM, SSRF, L7, TLS, CA, PEM, SSH, OCSF, and JSONL on first use.

Signed-off-by: Miyoung Choi <miyoungc@nvidia.com>
Made-with: Cursor

* docs: address audit feedback on refresh PR

- observability/logging: rewrite the allowed_ips paragraph after the
  Denial Reasons table; the previous wording said authors "can use"
  invalid entries while also stating they were rejected, which was
  contradictory and conflated load-time validation with the runtime
  per-CONNECT denial phrases the section documents.
- about/architecture: split compound sentences in the new Gateway
  Lifecycle and Observability sections so each clause stands alone.
- inference/about: drop the streaming-tolerance sentence from the prose
  paragraph since the dedicated Streaming reliability table row already
  covers it.

Signed-off-by: Miyoung Choi <miyoungc@nvidia.com>
Made-with: Cursor

* docs(architecture): address review feedback on architecture page

- Drop the Privacy Router row from the Components table; it is not
  yet a separately exposed customer-facing component.
- Update the page description and intro count to match the remaining
  three components (gateway, sandbox, policy engine).
- Split the policy decision into the three modes that the engine
  actually implements: Explicit Deny (deny rules and hardening rules,
  takes precedence), Allow, and Implicit Deny (no rule matched).

Signed-off-by: Miyoung Choi <miyoungc@nvidia.com>
Made-with: Cursor

* docs(architecture): convert policy decisions to a table

Promote the three policy decisions (Explicit Deny, Allow, Implicit
Deny) to a top-level table with Decision, When it applies, and
Outcome columns instead of a nested bulleted list under list item 5.
Top-level tables render reliably across markdown renderers, where
nested-in-list tables do not.

Signed-off-by: Miyoung Choi <miyoungc@nvidia.com>
Made-with: Cursor

* docs: repharse a bit

---------

Signed-off-by: Miyoung Choi <miyoungc@nvidia.com>
2026-04-17 13:15:59 -07:00
John T. Myers 2ca553a4a0 fix(sandbox): validate always-blocked IPs at load time, enrich denial logs, and filter un-fixable proposals (#814) (#815)
Policies with allowed_ips entries targeting loopback, link-local, or
unspecified ranges now fail at connection time instead of being silently
blocked at runtime. The shorthand log format for DENIED events includes
a [reason:...] suffix so operators can distinguish 'allowlist miss' from
'structurally un-allowable'. The mechanistic mapper skips proposals for
always-blocked destinations, preventing the infinite TUI notification
loop. The gateway validates proposed rules on approval as defense-in-depth.

- Extract shared IP helpers (is_always_blocked_ip, is_always_blocked_net,
  is_internal_ip) to openshell_core::net
- Reject always-blocked entries in parse_allowed_ips with hard error
- Skip implicit allowed_ips synthesis for always-blocked literal IP hosts
- Add status_detail to HttpActivityBuilder for denial reason propagation
- Enrich NET and HTTP shorthand with [reason:...] for DENIED events
- Add engine: tag to HTTP shorthand (consistency with NET shorthand)
- Filter always-blocked proposals in mechanistic mapper generate_proposals
- Add validate_rule_not_always_blocked server-side defense-in-depth
- Update architecture docs, published docs, and E2E test assertions
2026-04-13 09:47:50 -07:00
Piotr Mlocek 8b15ef772e docs(fern): move published docs into docs tree (#796)
Remove the legacy Sphinx pipeline and make docs/ the single source of truth so the published site matches the repository layout.
2026-04-09 15:02:27 -07:00
John T. Myers b7779bdefa feat(sandbox): integrate OCSF structured logging for sandbox events (#720)
* feat(sandbox): integrate OCSF structured logging for all sandbox events

WIP: Replace ad-hoc tracing calls with OCSF event builders across all
sandbox subsystems (network, SSH, process, filesystem, config, lifecycle).

- Register ocsf_logging_enabled setting (defaults false)
- Replace stdout/file fmt layers with OcsfShorthandLayer
- Add conditional OcsfJsonlLayer for /var/log/openshell-ocsf.log
- Update LogPushLayer to extract OCSF shorthand for gRPC push
- Migrate ~106 log sites to OCSF builders (NetworkActivity, HttpActivity,
  SshActivity, ProcessActivity, DetectionFinding, ConfigStateChange,
  AppLifecycle)
- Add openshell-ocsf to all Docker build contexts

* fix(scripts): attach provider to all smoke test phases to avoid rate limits

GitHub's unauthenticated API rate limit (60/hour) causes flaky 403s for
Phases 1, 2, and 4. Fix by attaching the provider to all sandboxes and
upgrading the Phase 1 policy to L7 so credential injection works.

Phase 4 (tls:skip) cannot inject credentials by design, so relax the
assertion to accept either 200 or 403 from upstream -- both prove the
proxy forwarded the request.

* fix(ocsf): remove timestamp from shorthand format to avoid double-timestamp

The display layer (gateway logs, TUI, sandbox logs CLI) already prepends
a timestamp. Having one in the shorthand output too produces redundant
double-timestamps like:

  15:49:11 sandbox INFO  15:49:11.649 I NET:OPEN ALLOWED ...

Now the shorthand is just the severity + structured content:

  15:49:11 sandbox INFO  I NET:OPEN ALLOWED ...

* refactor(ocsf): replace single-char severity with bracketed labels

Replace cryptic single-character severity codes (I/L/M/H/C/F) with
readable bracketed labels: [LOW], [MED], [HIGH], [CRIT], [FATAL].

Informational severity (the happy-path default) is omitted entirely to
keep normal log output clean and avoid redundancy with the tracing-level
INFO that the display layer already provides.

Before: sandbox INFO  I NET:OPEN ALLOWED ...
After:  sandbox INFO  NET:OPEN ALLOWED ...

Before: sandbox INFO  M NET:OPEN DENIED ...
After:  sandbox INFO  [MED] NET:OPEN DENIED ...

* feat(sandbox): use OCSF level label for structured events in log push

Set the level field to 'OCSF' instead of 'INFO' for OCSF events in the
gRPC log push. This visually distinguishes structured OCSF events from
plain tracing output in the TUI and CLI sandbox logs:

  sandbox OCSF  NET:OPEN [INFO] ALLOWED python3(42) -> api.example.com:443
  sandbox OCSF  NET:OPEN [MED] DENIED python3(42) -> blocked.com:443
  sandbox INFO  Fetching sandbox policy via gRPC

* fix(sandbox): convert new Landlock path-skip warning to OCSF

PR #677 added a warn!() for inaccessible Landlock paths in best-effort
mode. Convert to ConfigStateChangeBuilder with degraded state so it
flows through the OCSF shorthand format consistently.

* fix(sandbox): use rolling appender for OCSF JSONL file

Match the main openshell.log rotation mechanics (daily, 3 files max)
instead of a single unbounded append-only file. Prevents disk exhaustion
when ocsf_logging_enabled is left on in long-running sandboxes.

* fix(sandbox): address reviewer warnings for OCSF integration

W1: Remove redundant 'OCSF' prefix from shorthand file layer — the
    class name (NET:OPEN, HTTP:GET) already identifies structured events
    and the LogPushLayer separately sets the level field.

W2: Log a debug message when OCSF_CTX.set() is called a second time
    instead of silently discarding via let _.

W3: Document the boundary between OCSF-migrated events and intentionally
    plain tracing calls (DEBUG/TRACE, transient, internal plumbing).

W4: Migrate remaining iptables LOG rule failure warnings in netns.rs
    (IPv4 TCP/UDP, IPv6 TCP/UDP) to ConfigStateChangeBuilder for
    consistency with the IPv4 bypass rule failure already migrated.

W5: Migrate malformed inference request warn to NetworkActivity with
    ActivityId::Refuse and SeverityId::Medium.

W6: Use Medium severity for L7 deny decisions (both CONNECT tunnel and
    FORWARD proxy paths) to match the CONNECT deny severity pattern.
    Allows and audits remain Informational.

* refactor(sandbox): rename ocsf_logging_enabled to ocsf_json_enabled

The shorthand logs are already OCSF-structured events. The setting
specifically controls the JSONL file export, so the name should reflect
that: ocsf_json_enabled.

* fix(ocsf): add timestamps to shorthand file layer output

The OcsfShorthandLayer writes directly to the log file with no outer
display layer to supply timestamps. Add a UTC timestamp prefix to every
line so the file output matches what tracing::fmt used to provide.

Before: CONFIG:VALIDATED [INFO] Validated 'sandbox' user exists in image
After:  2026-04-01T15:49:11.649Z CONFIG:VALIDATED [INFO] Validated ...

* fix(docker): touch openshell-ocsf source to invalidate cargo cache

The supervisor-workspace stage touches sandbox and core sources to force
recompilation over the rust-deps dummy stubs, but openshell-ocsf was
missing. This caused the Docker cargo cache to use stale ocsf objects
from the deps stage, preventing changes to the ocsf crate (like the
timestamp fix) from appearing in the final binary.

Also adds a shorthand layer test verifying timestamp output, and drafts
the observability docs section.

* fix(ocsf): add OCSF level prefix to file layer shorthand output

Without a level prefix, OCSF events in the log file have no visual
anchor at the position where standard tracing lines show INFO/WARN.
This makes scanning the file harder since the eye has nothing consistent
to lock onto after the timestamp.

Before: 2026-04-01T04:04:13.065Z CONFIG:DISCOVERY [INFO] ...
After:  2026-04-01T04:04:13.065Z OCSF CONFIG:DISCOVERY [INFO] ...

* fix(ocsf): clean up shorthand formatting for listen and SSH events

- Fix double space in NET:LISTEN, SSH:LISTEN, and other events where
  action is empty (e.g., 'NET:LISTEN [INFO]  10.200.0.1' -> 'NET:LISTEN [INFO] 10.200.0.1')
- Add listen address to SSH:LISTEN event (was empty)
- Downgrade SSH handshake intermediate steps (reading preface, verifying)
  from OCSF events to debug!() traces. Only the final verdict
  (accepted/denied) is an OCSF event now, reducing noise from 3 events
  to 1 per SSH connection.
- Apply same spacing fix to HTTP shorthand for consistency.

* docs(observability): update examples with OCSF prefix and formatting fixes

Align doc examples with the deployed output:
- Add OCSF level prefix to all shorthand examples in the log file
- Show mixed OCSF + standard tracing in the file format section
- Update listen events (no double space, SSH includes address)
- Show one SSH:OPEN per connection instead of three
- Update grep patterns to use 'OCSF NET:' etc.

* docs(agents): add OCSF logging guidance to AGENTS.md

Add a Sandbox Logging (OCSF) section to AGENTS.md so agents have
in-context guidance for deciding whether new log emissions should use
OCSF structured logging or plain tracing. Covers event class selection,
severity guidelines, builder API usage, dual-emit pattern for security
findings, and the no-secrets rule.

Also adds openshell-ocsf to the Architecture Overview table.

* fix: remove workflow files accidentally included during rebase

These files were already merged to main in separate PRs. They got
pulled into our branch during rebase conflict resolution for the
deleted docs-preview-pr.yml file.

* docs(observability): use sandbox connect instead of raw SSH

Users access sandboxes via 'openshell sandbox connect', not direct SSH.

* fix(docs): correct settings CLI syntax in OCSF JSON export page

The settings CLI requires --key and --value named flags, not positional
arguments. Also fix the per-sandbox form: the sandbox name is a
positional argument, not a --sandbox flag.

* fix(e2e): update log assertions for OCSF shorthand format

The E2E tests asserted on the old tracing::fmt key=value format
(action=allow, l7_decision=audit, FORWARD, L7_REQUEST, always-blocked).
Update to match the new OCSF shorthand (ALLOWED/DENIED, HTTP:, NET:,
engine:ssrf, policy:).

* feat(sandbox): convert WebSocket upgrade log calls to OCSF

PR #718 added two log calls for WebSocket upgrade handling:

- 101 Switching Protocols info → NetworkActivity with Upgrade activity.
  This is a significant state change (L7 enforcement drops to raw relay).

- Unsolicited 101 without client Upgrade header → DetectionFinding with
  High severity. A non-compliant upstream sending 101 without a client
  Upgrade request could be attempting to bypass L7 inspection.
2026-04-07 13:01:13 -07:00
Miyoung Choi 574ef18dfc docs: consolidate information architecture and author content (#124)
* initial doc filling

* improvements

* stage provided get started, and add clean tutorials

* pull in Kirit's content and polish

* improve observability

* moving pieces

* move TOC around

* drop support matrix from concepts

* fix links

* minor fixes and fix badges

* minor fixes

* incorporate missed content

* minor improvements

* clean up

* run dori style guide review

* clean up

* updates impacting docs

* incorporate feedback

* minor fix

* some edits

* enterprise structure

* update cards

* improve

* add some emojis

* improve landing page with animated getting started code

* fix the animated code

* small improvements

* refresh content based on PR 156 and 158

* README as the source of truth for quickstart

* update README

* run edits

* change to the new prod name, text only, code swipe later

* add Home

* incorporate dev feedback on README

* krit's edits

* edit and improve index pages

* fix build

* revert README

* revert quickstart to not pull from README
2026-03-09 11:33:56 -07:00
Miyoung ChoiandDrew Newberry 11f795a462 docs: setup initial docs/ infrastructure and scaffolding (#94)
* set up docs

* rm nv sphinx theme version

* rm v* trigger

* incorporate feedback with cursor

* minor updates

* doc docs

* more small tweaks

* clean contributing

---------

Co-authored-by: Drew Newberry <anewberry@nvidia.com>
2026-03-04 22:31:45 -08:00