Since #2726 the canonical main process's stdout and stderr are captured
in pipes that feed only the in-memory replay buffer used by sandbox
connect. Agent output therefore never reaches the container's own stdout
and stderr, so it is missing from kubectl logs, docker logs, and podman
logs and from anything that collects container logs. Before #2726 the
entrypoint inherited the container's descriptors and its output appeared
there.
Copy the main process's output to the launcher's stdout and stderr in
addition to the replay buffer, restoring the earlier behavior:
- Output is copied byte for byte to the matching stream from a
forwarder thread per stream, after it is published to the replay
buffer. When the container runtime falls behind on a stream, that
stream's reader waits instead of dropping output, so backpressure
reaches the agent as it did with inherited descriptors, while the
other stream and attachments keep receiving output.
- Before the main process's exit is published, the output readers
finish and queued output is drained to the container log, so an
agent's final lines are not lost at shutdown. A 30 second deadline
covers both; when it expires, readers waiting on the container log
are released and drain the pipes into the replay buffer only, so a
stalled container log cannot block exit reporting.
- PTY-mode processes are not copied. The terminal stream carries escape
sequences and echoed input, and terminal commands never reached the
container log before #2726.
- Exec, SSH, and SFTP sessions are not copied.
Launcher log lines keep their existing format and remain in the
container's stderr. They are written as whole lines, and a newline is
inserted first when the agent left stderr mid-line, so launcher and
agent lines do not merge.
The Docker and VM drivers appended the tail of the workload's output to
failure messages: Docker the workload container's log, and the VM driver
the guest console, which carries the launcher's stdout and stderr. Those
messages land in the sandbox's Ready condition and in platform events
that the gateway republishes to the sandbox event stream. With agent
output in that log, those messages would carry arbitrary agent output,
including anything sensitive the agent prints, into gateway status and
events. The supervisor starts its health endpoint only after the agent
starts, so every Docker failure path could include agent output, and the
VM driver reports one whenever the VM or host supervisor exits. Forward
only the supervisor's log tail, matching the Podman driver, which reads
the workload log solely to match fixed launcher markers and never
forwards raw workload output. The workload's output remains available
through docker logs and the VM's rootfs-console.log.
Document where main process output appears in the logging docs and the
cluster debugging skill.
Closes#3928
Signed-off-by: Kris Hicks <khicks@nvidia.com>
* fix(policy): refresh pending proposals when the sandbox policy changes
Approving, removing, or undoing a rule, or updating the sandbox policy,
changes the inputs every other pending proposal was evaluated against.
Only proposals the new policy covered were reconciled; the rest kept
their old prover result and review token. The review surface
(GetDraftPolicy) therefore showed a stale evaluation, and the first
approval of the next proposal refreshed it and failed with
FAILED_PRECONDITION, so approving proposals one after another always
failed once.
Re-evaluate the remaining pending proposals at each policy change,
reusing the cached prover result unless the proposal's inputs changed.
Approval still rejects a review token that does not match the stored
evaluation, so a reviewer holding a pre-refresh evaluation must still
refetch it.
When a refresh does happen at approval time (inputs changed between
fetch and approve), the CLI now explains that the rule was re-evaluated
and how to review it, instead of printing the raw gRPC status.
Closes#3884
Signed-off-by: fede-kamel <fkamelhar@gmail.com>
* fix(policy): make pending proposal refresh race-safe and bounded
Store refreshed evaluations with a compare-and-swap: the store re-reads
the proposal, refuses when its rule name, proposed rule, or review token
changed since the evaluation read it, copies only the evaluation fields
onto the stored record, and updates only if the payload is still the one
it read. A refresh can no longer revert a concurrent edit or observation,
and the edit path uses the same guard against a concurrent refresh.
Bound each refresh to the 32 newest pending proposals; the rest keep the
approval-time recheck, which still refuses a stale review token. Operator
decisions (approve, approve-all, remove, undo) refresh before responding.
UpdateConfig, which holds the gateway-wide sandbox sync guard, and
agent-driven auto-approval refresh in a background task instead, one per
sandbox with later changes coalesced into a single rerun.
Refs #3884
Signed-off-by: fede-kamel <fkamelhar@gmail.com>
* docs(policy): describe proposal rechecks after approvals and approve-all
Explain that approving, removing, or undoing a rule rechecks the other
pending proposals so they can be approved one after another, when the
recheck is deferred or bounded, and what rule approve reports when a
proposal changed after it was listed. Show rule approve-all in Run Your
First Agent with its security-flag behavior.
Refs #3884
Signed-off-by: fede-kamel <fkamelhar@gmail.com>
* fix(policy): refresh pending proposals after a full policy replacement
A full policy UpdateConfig (openshell policy set) re-reads the latest
revision after its atomic write, finds the revision it just committed,
and returns before reaching the pending-proposal refresh at the end of
the handler. Pending proposals kept their stale evaluation, so rule get
showed the old candidate and the next approval failed with the refresh
precondition. Schedule the background refresh right after the commit.
Refs #3884
Signed-off-by: fede-kamel <fkamelhar@gmail.com>
---------
Signed-off-by: fede-kamel <fkamelhar@gmail.com>
With a non-terminal stdin, sandbox exec read stdin to EOF before it sent
the exec request. A pipe that never closes (CI runners, supervisors, agent
harnesses) blocked the CLI forever in read(2) without the gateway ever
seeing the request, and a slow producer delayed the command until EOF.
Collect piped stdin on a detached reader thread for at most 200 ms. Input
that reaches EOF within that window still travels in the single request
that older gateways need. If the pipe is still open, start the command
through the streaming RPC and forward the collected prefix plus the rest of
stdin as it arrives, closing remote stdin at EOF. The 4 MiB cap covers the
prefix and the streamed remainder together.
Closes#3993
Signed-off-by: Federico Kamelhar <federico.kamelhar@oracle.com>
* fix(supervisor): wait for repair when the gateway refuses a startup policy write
Startup writes the sandbox policy to the gateway in two cases: it
uploads a discovered image policy when the gateway has none, and it
writes the policy back after adding the proxy baseline filesystem paths.
When the gateway refused either write with FAILED_PRECONDITION or
INVALID_ARGUMENT, for example because the policy binds a provider that
is not attached, startup treated the refusal as a permanent error and
the supervisor exited. The sandbox never reached the ConfigurationInvalid
repair state that other startup rejections use.
Report such a refusal as a configuration rejection carrying the
gateway's message, log it once per write and error code, and keep
polling, so attaching the provider or replacing the policy completes
startup. Other error codes keep their current handling: transient codes
are retried, and permission, not-found and authentication failures
still end startup.
Skip the baseline-path write-back while a global policy is active. The
gateway refuses every sandbox policy write in that state, so startup
exited whenever a global policy lacked a baseline path. The supervisor
now adds the paths to its own copy of the policy without saving a
revision.
Signed-off-by: Shiju <shiju@nvidia.com>
* test(supervisor): stabilize startup refusal log capture
Keep a second tracing dispatcher alive while capturing startup refusal
logs. With only one dispatcher, a parallel test thread without a default
subscriber can cache Interest::never for the shared OCSF callsite after
the capture thread rebuilds the cache.
Preserve the exact log-count, diagnostic, configuration-generation and
repair assertions. Production startup behavior is unchanged.
Signed-off-by: Shiju <shiju@nvidia.com>
* fix(supervisor): reconcile stale startup rejection reports
Refetch desired configuration immediately when a rejection report is aborted because its generation changed. Preserve acknowledged rejection pacing and all other report error handling.
Signed-off-by: Shiju <shiju@nvidia.com>
* test(supervisor): box startup repair race futures
Keep the repair regressions below the large-future lint threshold without changing their inputs, scheduling, or assertions.
Signed-off-by: Shiju <shiju@nvidia.com>
---------
Signed-off-by: Shiju <shiju@nvidia.com>
* fix(snap): simplify snap hooks
The `post-refresh` hook runs after initial snap installation as well, so
there is no need to call the `install` hook from within the
`post-refresh` hook; instead, the logic can simply be moved into the
`post-refresh` hook directly, and the `install` hook removed.
Also, the existing `install` hook logic looked for an insecure
configuration, and if found, replaced the entire configuration file with
a minimal default in the current format. But OpenShell does that default
behavior without any config file, so we may as well simply remove the
configuration file entirely to keep up-to-date with the current default
behavior. Let OpenShell create a configuration file if it needs to,
rather than auto-create one via the packaging scripts.
Signed-off-by: Oliver Calder <oliver.calder@canonical.com>
* fix(snap): remove the connect-plug-docker hook
The `openshell:docker` is auto-connected to the system `:docker` slot,
so there should not be a need to separately restart the gateway service
when the interface is connected.
For locally-built test snaps which were not published to the store, the
autoconnection is not made, but when the snap is installed, the gateway
will attempt to start anyway and fail to find any available compute
driver, so quickly restart until it hits the systemd start-limit, after
which systemd prevents the service from being started again. If a user
tries to manually connect their locally-built `openshell` snap to the
`:docker` slot, then the `connect-plug-docker` hook runs and triggers a
restart of the gateway, which will usually fail because the start limit
has already been hit. An error in the hook will thus cause the interface
connection to be undone, which is undesirable.
Thus, we can remove this hook entirely, and instead allow interface
connections to succeed as intended. The user still needs to manually
restart the gateway service after making a manual connection (as was the
case previously) and probably needs to `systemctl reset-failed` first,
but at least connection will succeed beforehand so they can proceed with
these steps.
Signed-off-by: Oliver Calder <oliver.calder@canonical.com>
* fix(snap): set refresh-mode: endure again, with manual restart
Return to the previous behavior before commit a67567e58, where the
gateway is not stopped before refreshes. The `post-refresh` hook
now restarts the gateway if the TLS configuration was corrected, so we
don't have to enforce restarting the gateway on every refresh even when
not necessary. Thus, set `refresh-mode: endure`, and let the hook decide
when the gateway needs to be restarted.
Signed-off-by: Oliver Calder <oliver.calder@canonical.com>
* fix(snap): update docs and tests to reflect snap hook changes
Signed-off-by: Oliver Calder <oliver.calder@canonical.com>
* docs(snap): remove verbose explanation of snap gateway refresh behavior
Signed-off-by: Oliver Calder <oliver.calder@canonical.com>
---------
Signed-off-by: Oliver Calder <oliver.calder@canonical.com>
* fix(network): refuse protocol upgrades on GraphQL endpoints
Refuse Upgrade headers before forwarding GraphQL-over-HTTP requests.
Share the protocol refusal table with JSON-RPC and MCP, and close
unexpected protocol switches before relaying frames.
Keep GraphQL-over-WebSocket inspection on separate WebSocket endpoints.
Cover upgrade refusal, audit mode, subscription handshakes, and ordinary
HTTP and WebSocket controls. Update the current policy documentation.
Signed-off-by: Shiju <shiju@nvidia.com>
* fix(network): refuse GraphQL upgrades before reading bodies
Validate the HTTP head and endpoint authority before upgrade refusal, then inspect ordinary GraphQL bodies. Preserve missing-authority credential rejection after body inspection.
Signed-off-by: Shiju <shiju@nvidia.com>
---------
Signed-off-by: Shiju <shiju@nvidia.com>
Store operation spans and request spans for supervisor-polled RPCs
(GetSandboxConfig, ReportProviderReadiness) use DEBUG level, so the
default INFO filter no longer exports them. The provider credential
refresh worker opens its span only when a state has work.
Refs #2698
Signed-off-by: Kris Hicks <khicks@nvidia.com>
* test(policy): reproduce raw OPA loading gaps against the typed schema
The supervisor loads a sandbox policy in two ways: through the typed
schema (parse_sandbox_policy, then from_proto) or directly into OPA
(from_strings and from_files). The raw path fills in defaults where the
typed schema is strict, so the same policy text can produce a different
sandbox configuration, or load when it should be rejected.
Add two regression tests that fail on the current code:
- An empty filesystem_policy loads with include_workdir true through raw
OPA and false through the typed schema. An absent stanza gives true on
both paths and must keep doing so.
- Raw OPA accepts a string include_workdir, a non-string read_only entry,
an unknown Landlock compatibility and an explicit null json_rpc, with
or without a version key. The typed schema rejects each. Every case has
a valid twin that both paths must accept.
A follow-up change makes raw loading apply the typed schema's rules.
Refs #3092.
Signed-off-by: Shiju <shiju@nvidia.com>
* fix(policy): align raw OPA loading with typed settings
Validate raw filesystem, Landlock, and process settings with the canonical
authored schema before normalization. Preserve the absent filesystem
default while applying the present-stanza default, and canonicalize valid
Landlock enum representations before runtime evaluation.
Reject explicit null JSON-RPC options through the shared parser. Preserve
versionless and runtime OPA data, and keep rejected reloads from replacing
the active policy or advancing its generation.
Add raw-versus-typed, file-loader, and rejected-reload regressions and
document the local loading contract.
Refs #3092.
Signed-off-by: Shiju <shiju@nvidia.com>
* fix(policy): validate raw OPA settings and redact startup errors
Validate raw network fields through the authored schema before
normalization. Preserve custom Rego data and supported runtime forms.
Apply the shared filesystem path checks and non-root identity predicate
to raw static settings.
Discard authored Rego source and nested errors from static configuration
evaluation. Cover malformed inputs, valid controls, file loading, and
rejected reloads retaining active decisions and generation.
Refs #3092.
Signed-off-by: Shiju <shiju@nvidia.com>
* test(policy): satisfy unit-returning assertion lint
Terminate the two error-assertion match arms with semicolons, as required
by Clippy. Preserve the existing checks and runtime behavior.
Refs #3092.
Signed-off-by: Shiju <shiju@nvidia.com>
---------
Signed-off-by: Shiju <shiju@nvidia.com>
* fix(mcp): explain revision-scoped policy and rejections
Explain the selected-revision method set in profile output and policy docs.
Distinguish protocol and policy rejection causes and give a next step while
preserving authorization, response statuses, error codes and YAML keys.
Cover CLI serialization, revision selection, exact extension rules, deny
precedence and rejection before forwarding with focused regressions.
Signed-off-by: Shiju <shiju@nvidia.com>
* docs(mcp): correct HTTP cancellation revision support
Limit notifications/cancelled to the three 2025 revisions in the core
method matrix. State that MCP 2026-07-28 HTTP cancellation closes the
response stream, matching the runtime rejection and sessionless docs.
Signed-off-by: Shiju <shiju@nvidia.com>
---------
Signed-off-by: Shiju <shiju@nvidia.com>
* fix(supervisor): restore canonical stdin after connection loss
Probe idle SSH peers and enforce a receive deadline during transport I/O,
including writes blocked by a stalled relay. Release the dead attachment's
stdin lease through existing handler cleanup.
Retry denied write intent on later ordinary input without displacing a
healthy owner. Preserve explicit read-only, EOF and detach behavior, and
discard control bytes retained while input ownership was denied.
Cover half-open forwarding, blocked writes, healthy idle peers and competing
reconnects through the production supervisor frame bridge and real SSH.
Fixes#3648
Signed-off-by: Shiju <shiju@nvidia.com>
* docs(skills): describe read-only reconnect input retry
Explain what an openshell-cli user sees when automatic recovery reattaches before the supervisor closes the dead connection: the attachment reports read-only, later ordinary input retries stdin acquisition and prints `input enabled`, input typed while read-only is discarded, exit keys still detach, and an explicitly read-only viewer or a healthy owner is never affected.
Signed-off-by: Shiju <shiju@nvidia.com>
---------
Signed-off-by: Shiju <shiju@nvidia.com>
Previously, Helm installations could not enable the gateway OCSF JSONL
destination through chart values because generated `gateway.toml` omitted the
`openshell.gateway.ocsf_log` table.
Now, setting `server.ocsfLog.enabled` renders the path, optional schema
version, rotation, retention, and queue limits into gateway configuration.
Output is disabled by default. The default path, `/tmp/gateway-ocsf.jsonl`,
is writable in the gateway container with either the StatefulSet or
Deployment workload, so enabling output does not require persistent storage.
Invalid schema versions, rotation values, non-positive limits, or an empty
path while enabled fail chart rendering.
Additionally, `server.extraVolumes` and `server.extraVolumeMounts` add
operator-supplied volumes to the gateway pod, so operators who want records
to survive restarts can place the OCSF path on persistent storage without
replacing chart-generated configuration.
The gateway pod's default termination grace period rises from 5 to 30
seconds. Gateway shutdown can spend up to 10 seconds on supervisor session
cleanup before allowing 5 seconds to drain queued OCSF records, so the
5-second default risked a SIGKILL before the final records were written.
The grace period is only an upper bound: the gateway exits as soon as its
shutdown completes.
Refs #2762
Signed-off-by: Kris Hicks <khicks@nvidia.com>
* feat(server): write gateway OCSF events to JSONL
Previously, gateway security activity was available only in diagnostic output,
and events not associated with a sandbox, such as TLS certificate reloads, had
no independent structured record.
Now, configuring `openshell.gateway.ocsf_log` writes every gateway-produced
OCSF record to a bounded JSONL destination independently of `RUST_LOG`. The
destination supports daily or disabled rotation, retention limits, queue bounds,
and optional schema downgrade to OCSF 1.1 or 1.3.
Additionally, existing gateway emitters (TLS reloads, service routing, and
policy approval and auto-approval audits) emit structured events, so they reach
the JSONL destination, console shorthand, and the affected sandbox's log
stream. Records identify the gateway by its configured name in `device.uid` and
`device.name`, shared across replicas, with `device.hostname` identifying the
replica and `device.os` the gateway's operating system. Metrics and warnings
expose known best-effort losses.
Refs #2762
Signed-off-by: Kris Hicks <khicks@nvidia.com>
* fix(mxc): attribute ETW events to gateway
Signed-off-by: Evan Lezar <elezar@nvidia.com>
---------
Signed-off-by: Kris Hicks <khicks@nvidia.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>
Co-authored-by: Evan Lezar <elezar@nvidia.com>
- Kubernetes now checks supervisor readiness by connecting to TCP port 5501
- Stop starting a supervisor process in every sandbox each second
- The supervisor opens the port only while its gateway session is up
- Accept IPv4 and IPv6 probes, even when net.ipv6.bindv6only is set
- Keep the health socket for Docker, Podman, and debugging
- Add tests and update the docs
Signed-off-by: divesh <dgude@nvidia.com>
* fix(sandbox-backend): sort boundary request objects before hashing
Sort boundary request objects recursively before hashing so serde_json's
preserve_order feature cannot change digest identity. Cover canonical
bytes, envelope round trips, and rejection of modified provider values
and operations.
Signed-off-by: Shiju <shiju@nvidia.com>
* feat(mcp): upgrade tower-mcp-types to 0.22.2
Upgrade tower-mcp-types from 0.12.0 to an exact-pinned 0.22.2 and use its
inspection APIs to validate MCP requests against the selected revision.
Carry inspection metadata into policy evaluation and validate requests
after header rewriting, before forwarding.
Add explicit support for the sessionless 2026-07-28 revision while keeping
2025-11-25 as the default. Validate per-request metadata and standard HTTP
header mirrors, and support discovery, tools, and subscription requests.
Delegate batch availability and parameter schemas to Tower. Share typed
request names between policy and HTTP checks, retain the local batch
resource cap, and centralize MCP policy version parsing and ordering.
Keep supported MCP revisions and shared allowlist parsing in the canonical
policy schema; core re-exports those types. Tower owns wire-profile
semantics, and every supported policy revision must map to the matching
inspector profile.
Reject duplicate JSON keys, invalid known-method parameters, unavailable
methods, and unsupported batches. Keep exact extension allow rules and
deny precedence. Document request inspection boundaries and add unit,
forwarding, and sandbox coverage.
Refs #2174.
Signed-off-by: Shiju <shiju@nvidia.com>
* test(mcp): prove authorization at the forwarding boundary
Cover March batch denial in both member orders, valid and malformed
controls, and audit behavior across both relay entry paths. Exercise real
middleware tool rewrites with matching metadata and assert the exact
upstream representation or zero forwarded bytes.
Verify legacy bodyless SSE GET remains usable while GET tool bodies and
unsupported DELETE cleanup are rejected. Clarify request-selected profile
and middleware mutation comments without changing production behavior.
Signed-off-by: Shiju <shiju@nvidia.com>
* test(mcp): exercise permitted profiles through the sandbox proxy
Cover March and June singleton policies and select November and July
separately under one endpoint allowlist. Capture upstream tool receipts
to distinguish proxy policy denial from an upstream rejection.
Extend middleware rewrite coverage to June and multi-version policies,
and preserve the sessionless discovery and subscription checks through
the shared fixture helpers.
Signed-off-by: Shiju <shiju@nvidia.com>
* test(kubernetes): box the admission check future
Keep the admission test future below Clippy's size limit when the
workspace dependency features are unified.
Signed-off-by: Shiju <shiju@nvidia.com>
* test(mcp): reuse the forwarding fixture identity cache
Share the binary identity cache across protocol-profile cases, matching
the proxy lifecycle and avoiding repeated hashes of the test executable.
Keep procfs authorization and all forwarding assertions intact.
Signed-off-by: Shiju <shiju@nvidia.com>
---------
Signed-off-by: Shiju <shiju@nvidia.com>
* fix(network): refuse protocol upgrades on JSON-RPC and MCP endpoints
JSON-RPC and MCP rules apply to each HTTP request, but the proxy could
forward a request that also carried upgrade headers. After an upstream
answered 101, route selection and the forward proxy relayed the
connection without inspection.
Refuse any request that carries an Upgrade header on JSON-RPC-family
endpoints before the L7 policy decision, in every enforcement mode.
Share the check with the existing h2c refusal and call it from
relay_jsonrpc as well. Record the refusal as a policy denial and answer
with the unsupported_l7_protocol error, because no policy rule can
allow the request.
If a JSON-RPC-family endpoint still receives 101, close the connection
instead of relaying raw bytes. Document the refusal and the WebSocket
alternative.
Signed-off-by: Shiju <shiju@nvidia.com>
* docs(observability): remove duplicate protocol error definition
Keep unsupported_l7_protocol in the response error-code list and retain its explanation in the policy troubleshooting table.
Signed-off-by: Shiju <shiju@nvidia.com>
---------
Signed-off-by: Shiju <shiju@nvidia.com>
The gateway chart always rendered the ClusterRole and ClusterRoleBinding, so
every install and upgrade required cluster-admin even when only namespaced
objects were needed. Installers that are namespace-admin GitOps or platform
controllers could not run the release at all, and clusters where cluster-scoped
RBAC is owned by a separate team had no supported way to split the install.
Add an rbac values block so a cluster-admin can apply the cluster-scoped objects
once and a namespace-admin can install and upgrade the release without
cluster-scoped permissions:
rbac:
create: true
clusterScoped:
create: true
clusterRoleName: ""
clusterRoleBindingName: ""
rbac.clusterScoped.create gates the ClusterRole and ClusterRoleBinding, and is
independent of the workspace mode. rbac.create additionally gates the namespaced
sandbox Role and RoleBinding, which matters because Kubernetes escalation
prevention stops an installer holding only the built-in admin role from creating
a Role that grants agents.x-k8s.io verbs it does not itself hold. The certgen
hook and credential driver RBAC keep their existing flags.
Both flags default to true, so current installs are unchanged. The helpers treat
a missing rbac block as enabled so upgrades with --reuse-values do not drop RBAC,
matching the existing workspaceResources pattern. The ClusterRoleBinding roleRef
follows clusterRoleName so a separately applied ClusterRole can carry a name the
cluster-admin chooses.
Document the migration for a release that already owns the cluster-scoped
objects: Helm deletes objects that leave the manifest, so annotate them with
helm.sh/resource-policy=keep before setting the flag, otherwise the gateway
loses TokenReview until a cluster-admin re-applies them.
Signed-off-by: ansjindal <ansjindal@nvidia.com>
* docs: align page file names and nav labels with published URLs
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs: redirect moved dev pages and fix agent guide redirects
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* ci(docs): check that page URLs match their file paths
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* fix(docs): align navigation checks with repo conventions
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs: omit historical redirect aliases
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* fix(docs): preserve published overview and TypeScript URLs
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* fix(docs): redirect unversioned overview and TypeScript URLs
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
---------
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(readme): show Rust SDK installation command
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* docs(architecture): lead with policy enforcement
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* docs(architecture): reuse README how it works text
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* docs(architecture): refresh architecture page and diagrams
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* docs(extensibility): simplify extension points diagram and overview
Redraw the extension points diagram as two left-to-right rows for the
data plane and control plane, describe which layers each extension point
extends, and move authentication after Building Extensions as
Authenticating Extensions with updated cross-page links.
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
---------
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* fix(snap): require mTLS for the snap gateway
Replace the installer opt-in with an authenticated snap gateway. The wrapper
no longer forces plaintext, so the gateway serves TLS from the bundle it
already generates in $SNAP_COMMON/tls. The install hook writes a config that
enables mTLS user auth instead of unauthenticated access, and a new
post-refresh hook migrates the exact legacy default on existing installs.
install.sh waits for the gateway, detects whether it serves TLS, copies the
client bundle into the target user's snap state directory, and registers the
gateway over HTTPS. Older plaintext snap revisions still register over HTTP
with a warning. The release canary asserts mTLS auth and HTTPS registration.
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* fix(snap): pass config preflight and detect the mTLS gateway reliably
An explicit [openshell.gateway.mtls_auth] table fails config preflight,
which validates mTLS auth before the local TLS bundle supplies the client
CA. Write a default that pins the Docker driver instead; with the wrapper's
TLS bundle the gateway requires client certificates and enables mTLS user
auth automatically, as the native packages do.
The mTLS gateway rejects TLS handshakes without a client certificate, and
it still answers plaintext loopback HTTP for sandbox service routing, so the
installer could misdetect it as a legacy plaintext gateway. Probe HTTPS with
the root-owned client bundle, and treat a gateway as legacy only when a
plaintext gRPC Health call succeeds.
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* chore(snap): simplify install hook comment
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* fix(snap): migrate insecure gateway configs on refresh
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* refactor(snap): simplify mTLS detection and config migration
Detect the mTLS snap from the installed revision's post-refresh hook instead
of probing plaintext gRPC, and drop the scheme global. Remove the installer's
pre-hook config fallback, which is dead now that every channel ships the
install hook and which wrote the insecure default. Give the install hook a
single write path with a simple backup name, and shorten the manual client
certificate steps.
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* fix(snap): stop keeping a copy of replaced insecure configs
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* feat(install): make the snap an opt-in install method
Stop selecting the OpenShell snap just because the snap command exists. Linux
installs default to the Debian or RPM package; OPENSHELL_INSTALL_METHOD=snap
(or deb, rpm) selects the package explicitly. Hosts that already have the
OpenShell snap keep refreshing it rather than gaining a second gateway on the
same port. The release canary and snap repro script opt in explicitly.
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* fix(snap): let the gateway auto-detect its compute driver
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* fix(snap): restart the gateway after refresh
Published revisions use refresh-mode: endure, and snapd honors the old
revision's setting during a refresh, so the plaintext gateway kept running
with the migrated config unused until a manual restart. Restart the gateway
from the post-refresh hook so the mTLS config takes effect immediately.
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* docs(snap): drop refresh notes from the snap description
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* docs(snap): trim snap refresh notes from installation docs
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
---------
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* docs(tutorials): run Pi with OpenRouter
Switch the Pi tutorial from Anthropic to OpenRouter, add fd to the Pi image
so Pi does not try to download it from GitHub inside the sandbox, and rename
the page to Run Pi with OpenRouter with a dev redirect from the old URL.
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(tutorials): drop redirect for dev-only Pi page
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(tutorials): make the Pi tutorial easier to follow
Add prerequisites and steps, explain providers and the profile fields in
plain terms, inspect the sandbox while Pi is still running, and add clean-up
and troubleshooting sections.
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(tutorials): clarify Pi's providers and sandbox additions
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(tutorials): make the Pi tutorial more conversational
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
---------
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
The /_ws_tunnel endpoint pipes a WebSocket into the full gRPC service. On a
plaintext loopback gateway any web page could open it, since browsers do not
apply CORS to WebSocket upgrades. Mount the tunnel only when
enable_websocket_tunnel is set (config file, --enable-websocket-tunnel, or
OPENSHELL_ENABLE_WEBSOCKET_TUNNEL; server.enableWebsocketTunnel in Helm).
BREAKING CHANGE: gateways behind an authenticating edge proxy must set
enable_websocket_tunnel = true for CLI edge-tunnel connections.
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* docs(readme): streamline README and move reference detail to docs
Restructure the README as a short path from overview to quickstart to
further reading. Move prerelease install steps into the installation
guide and telemetry build flags into a new observability page. Fix
broken docs links and outdated runtime and credential descriptions.
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* docs(readme): describe 0.1.0 as adding new isolation primitives
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* docs(architecture): add policy prover as a gateway component
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* docs: describe OpenShell as a runtime for fleets of autonomous AI agents
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* docs(architecture): describe policy prover as formal verification
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* docs(architecture): name OpenShell Sandbox in component table
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* docs(architecture): fold isolation backend into supervisor row
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* docs(architecture): mention formal verification in overview
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* docs(run-agent): use the published OpenCode image in the first-agent guide
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* docs(inference): correct provider examples and readiness
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
* docs(providers): correct Google binding and provider selection
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
* docs(readme): sharpen value prop, how it works, and explore further
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* docs(run-agent): use example Anthropic profile and add policy advisor step
Import the example Anthropic profile, which now allows OpenCode, instead
of editing it with sed. Add a step that shows how to review and approve
mechanistic policy proposals as the agent needs more access.
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* feat(providers): add OpenRouter example for OpenCode
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* docs(run-agent): run OpenCode against OpenRouter with a free model
Add an example OpenRouter provider profile scoped to OpenCode and switch
the first-agent guide to it, using a free Nemotron model so readers do
not need OpenRouter credits. Revert the OpenCode binary added to the
example Anthropic profile.
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* docs(readme): link first-agent guide and add agent skills section
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
---------
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Co-authored-by: Piotr Mlocek <pmlocek@nvidia.com>
* docs(policy): correct schema and default policy guidance
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): add network recipes and update command reference
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): organize lifecycle guidance and troubleshooting
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): split policy overview into concepts and management tasks
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): reorganize network recipes as a cookbook
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): restructure schema reference by field group and protocol
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): align troubleshooting, advisor, and reference pages
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): fix first policy tutorial and security guidance
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): keep overview high level and move network rules to their own page
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): focus policy management on CLI workflows and remove command reference
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): clarify policy views and sandbox deletion in management guide
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): streamline network rule concepts and examples
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): correct request path wildcard semantics
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): rewrite policy advisor guide for clarity
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): clarify policy advisor scope, setup, and review
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): rewrite policy prover guide for clarity
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): explain the two uses of the policy prover
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): describe policy prover uses, boundaries, and coverage
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): place prover before advisor and troubleshooting last
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): remove unsupported CI guidance from prover page
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): tighten policy prover introduction
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): move policy change behavior into management guide
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): name prover check types and note expanding coverage
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): prefix prover and advisor sidebar labels
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): streamline policy schema reference
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): place default policy before schema reference
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): fold troubleshooting into policy management guide
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): correct tutorial log samples and GitHub push policy steps
The first policy tutorial said the 403 body begins with error, policy, and
rule, but the proxy serializes the body with sorted keys. Its log samples also
showed the wrong CONNECT deny reason for a sandbox without network rules, and
the L7 deny sample omitted the :443 authority, the `l7` engine, and the reason
tag that the shorthand formatter emits.
The GitHub tutorial filtered denials with `--level warn`, which hides the INFO
level OCSF policy events, and showed the retired key=value log format. Its
hand-written policy also omitted /bin from the restrictive default, so
`policy set` would reject the file for removing a filesystem path on a live
sandbox. Start from `policy get --base` and add only the network rules.
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): improve flow and terminology across policy pages
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): correct network rule matching and protocol details
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): align policy management steps with CLI behavior
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): correct policy advisor proposal and approval details
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): correct policy section, default, and schema details
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): correct prover installation and coverage limits
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): recommend tls skip for server-first protocols
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): fix stale baseline path and interpreter examples
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): move policy pages under how-it-works and fix links
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): align native TCP guidance in security best practices
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): restore policy.local and policy DNS details from main
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): state exact glob matching rules
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
---------
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* fix(policy): propose rules for unknown DNS hosts
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
* docs(policy): clarify synthetic DNS use across protocols
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
* fix(policy): harden unknown-host DNS observations
- Emit the policy_dns_ineligible denial for every unknown name and
report observation staging failures as DNS failure events.
- Refuse unknown names during fail-closed quarantine and after the
observation budget, now a quarter of each address family's pool.
- Pin transparent TCP to the mapping of the deciding policy generation
so a reload between DNS and authorization fails closed.
- Stop Docker workloads from inheriting host DNS search domains, which
let the first expanded short name claim an observation address.
- Share mechanistic draft polling in conformance, register
new-hostname-proposal in the installed suite, and update docs.
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
* test(policy): build policy DNS proxy tests on every target
The proxy tests name PolicyEndpointId, which proxy.rs imported only on
Linux, so the macOS test build failed. Import it for test builds too.
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
* fix(policy): name DNS queries and mapped hosts in OCSF denials
DNS denial and failure events attached port 53 to the queried name,
which read as a connection to that host. They now carry only the name.
Transparent TCP denials for a policy DNS address show the mapped
hostname and keep the synthetic address in dst_endpoint.ip.
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
---------
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
* docs(extensibility): reorganize extensibility and middleware guides
Add an extensibility overview that introduces the extension protocol and
links every extension point, and move extension caller authentication to
a shared page used by middleware and gateway interceptors.
Split the supervisor middleware guide into an overview, a Configure and
Operate guide, and a Middleware Operations reference for service authors.
The overview explains when to use middleware, shows where it runs, lists
current limitations, and defines the service contract. Configure and
Operate covers policy attachment, service registration, failure behavior,
and observability. Middleware Operations describes HTTP request, HTTP
response, and WebSocket message operations with shared inputs and results,
per-operation diagrams, and detail accordions.
Pin page slugs so links resolve, redirect the replaced dev middleware
URL, and update reference and architecture links.
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
* docs(middleware): clarify navigation and service contracts
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
* docs(middleware): drop obsolete dev URL redirect
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
---------
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
* fix(install): avoid installing incompatible docker snap
The work to land RFC-0012 added new restrictions when interacting with
Docker by setting `NoNewPrivs`. This prevents the `docker` snap from
transitioning its AppArmor profile from `snap.docker.dockerd` to
`docker-default` when it tries to launch a container. Thus, the `docker`
snap is currently incompatible with OpenShell.
This commit prevents `install.sh` from installing the `docker` snap
before installing the `openshell` snap, and instead requires the user to
install a non-snap Docker daemon before proceeding with installing the
snap. Systems without the `snap` command are unaffected, since they
install native packages without checking for the presence of Docker or
other compute providers.
Signed-off-by: Oliver Calder <oliver.calder@canonical.com>
* fix(snap): only snapd 2.76 for openshell snap since store installs work
Signed-off-by: Oliver Calder <oliver.calder@canonical.com>
* fixup! fix(install): avoid installing incompatible docker snap
Signed-off-by: Oliver Calder <oliver.calder@canonical.com>
* fixup! fix(snap): only snapd 2.76 for openshell snap since store installs work
Signed-off-by: Oliver Calder <oliver.calder@canonical.com>
---------
Signed-off-by: Oliver Calder <oliver.calder@canonical.com>
* perf(server): enable WAL and NORMAL sync for the SQLite store
On-disk SQLite stores ran with sqlx defaults: rollback journal
(`journal_mode=delete`) and `synchronous=FULL`. Every autocommit write paid
several fsyncs and blocked readers while it held the lock, so gateway hot
paths made of many small writes serialized on disk latency. The clearest
case is `openshell forward service`, which mints and revokes an SSH session
token around every forwarded TCP connection: two commits per connection,
tens of milliseconds each on a virtual disk, wall clock linear in the
number of concurrent connections, and enough queueing that bursts hit the
per-sandbox connection cap and get refused.
Switch on-disk databases to WAL with `synchronous=NORMAL`. The mode change
runs once on a single connection before the pool opens: entering WAL needs
exclusive access to the file, so doing it up front means pool connections
only ever re-apply the pragma to a file already in WAL mode, and a failure
surfaces as one clear connect error. The first start after upgrading an
existing database therefore needs the file to be otherwise unopened.
`synchronous` is applied through the connect options on every pooled
connection. In-memory databases keep their defaults. A crash can now roll
back the most recent transactions without corrupting the database, which
is the standard WAL trade-off and fits the single-node scope of the SQLite
backend.
Tests cover a fresh store, an existing rollback-journal file that must be
switched on connect, sidecar permissions, and concurrent readers under a
burst of insert-then-update writes. Architecture, configuration and Helm
docs describe the durability trade-off, the sidecar files, and the local
filesystem requirement.
Signed-off-by: Jason T. Greene <jason.greene@redhat.com>
* fix(server): keep SQLite commits durable except SSH session issuance
WAL with synchronous=NORMAL can roll back acknowledged commits after a
power loss or kernel crash, including SSH session revocations and other
authorization-tightening writes. Run the main pool with synchronous=FULL
so every acknowledged write is durable; in WAL mode that is a single
fsync of the WAL per commit.
Add Store::create_relaxed for inserts that are safe to lose, and use it
only for SSH session issuance: a dropped token just fails validation.
On file-backed SQLite it runs on a dedicated single-connection pool with
synchronous=NORMAL. Both pools share one WAL, so the next FULL commit
also makes earlier relaxed commits durable. Postgres treats it as an
ordinary durable MustCreate insert.
Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
---------
Signed-off-by: Jason T. Greene <jason.greene@redhat.com>
Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
Co-authored-by: Mrunal Patel <mrunalp@gmail.com>
* fix(podman): restore host gateway alias mediation
Signed-off-by: Gordon Sim <gsim@redhat.com>
* fix(podman-e2e-tests): enable broader test podman e2e coverage
Signed-off-by: Gordon Sim <gsim@redhat.com>
* fix(tests): make test more reliable
Signed-off-by: Gordon Sim <gsim@redhat.com>
* fix(podman): fix macos linting error
Signed-off-by: Gordon Sim <gsim@redhat.com>
---------
Signed-off-by: Gordon Sim <gsim@redhat.com>