Commit Graph
1574 Commits
Author SHA1 Message Date
Adrien Langou 0b4b71da7f fix(cli): stop uploads when Git filtering fails or selects no files
Signed-off-by: Adrien Langou <alangou@nvidia.com>
2026-10-01 19:17:37 +02:00
Evan Lezar 5acaaba192 test(conformance): verify deletion through sandbox list (#3792)
Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-09-30 09:15:45 +00:00
Polite_realismandEvan Lezar b8ffe5244c test(podman): move podman_preflight into driver-podman integration tests (#3783)
* test(podman): move podman_preflight into driver-podman integration tests

podman_preflight verifies that openshell-driver-podman fails fast when
its Podman socket is unreachable. It only needs the standalone driver
binary, not a gateway, so it never fit the gateway-backed e2e-podman
harness it lived under and never ran anywhere in CI.

Move it into crates/openshell-driver-podman/tests/ as a plain Cargo
integration test. It now runs via the existing required workspace test
job with no special mise task, workflow step, or coverage exception.

Signed-off-by: politerealism <burdcat17@gmail.com>

* test(podman): make preflight diagnostics portable

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: politerealism <burdcat17@gmail.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>
Co-authored-by: Evan Lezar <elezar@nvidia.com>
2026-09-30 06:05:18 +00:00
Eric Busto b8932d43be Fix/startup provider readiness (#3819)
* fix(supervisor): preserve startup provider readiness

Signed-off-by: Eric Busto <ebusto@nvidia.com>

* test(supervisor): cover startup provider polling

Signed-off-by: Eric Busto <ebusto@nvidia.com>

---------

Signed-off-by: Eric Busto <ebusto@nvidia.com>
2026-09-30 06:03:38 +00:00
Shiju 798500ccdb fix(policy): validate raw OPA settings and redact startup errors (#3788)
* test(policy): reproduce raw OPA loading gaps against the typed schema

The supervisor loads a sandbox policy in two ways: through the typed
schema (parse_sandbox_policy, then from_proto) or directly into OPA
(from_strings and from_files). The raw path fills in defaults where the
typed schema is strict, so the same policy text can produce a different
sandbox configuration, or load when it should be rejected.

Add two regression tests that fail on the current code:

- An empty filesystem_policy loads with include_workdir true through raw
  OPA and false through the typed schema. An absent stanza gives true on
  both paths and must keep doing so.
- Raw OPA accepts a string include_workdir, a non-string read_only entry,
  an unknown Landlock compatibility and an explicit null json_rpc, with
  or without a version key. The typed schema rejects each. Every case has
  a valid twin that both paths must accept.

A follow-up change makes raw loading apply the typed schema's rules.

Refs #3092.

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(policy): align raw OPA loading with typed settings

Validate raw filesystem, Landlock, and process settings with the canonical
authored schema before normalization. Preserve the absent filesystem
default while applying the present-stanza default, and canonicalize valid
Landlock enum representations before runtime evaluation.

Reject explicit null JSON-RPC options through the shared parser. Preserve
versionless and runtime OPA data, and keep rejected reloads from replacing
the active policy or advancing its generation.

Add raw-versus-typed, file-loader, and rejected-reload regressions and
document the local loading contract.

Refs #3092.

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(policy): validate raw OPA settings and redact startup errors

Validate raw network fields through the authored schema before
normalization. Preserve custom Rego data and supported runtime forms.
Apply the shared filesystem path checks and non-root identity predicate
to raw static settings.

Discard authored Rego source and nested errors from static configuration
evaluation. Cover malformed inputs, valid controls, file loading, and
rejected reloads retaining active decisions and generation.

Refs #3092.

Signed-off-by: Shiju <shiju@nvidia.com>

* test(policy): satisfy unit-returning assertion lint

Terminate the two error-assertion match arms with semicolons, as required
by Clippy. Preserve the existing checks and runtime behavior.

Refs #3092.

Signed-off-by: Shiju <shiju@nvidia.com>

---------

Signed-off-by: Shiju <shiju@nvidia.com>
2026-09-30 05:59:31 +00:00
Drew Newberry 252882f37f feat(providers): serve sandbox config files on demand (#3832)
* feat(providers): serve sandbox config files on demand

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(providers): defer managed file documentation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(providers): mark managed file api experimental

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* style(go): format provider profile fields

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(providers): preserve legacy environment with managed files

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-30 01:41:32 +00:00
Fede Kamelhar 5c0c9e446e feat(providers): add OCI Generative AI example provider profile (#3904)
Add an example oci-genai inference profile for Oracle Cloud Infrastructure
Generative AI through its OpenAI-compatible endpoint. The profile injects a
compartment-scoped Generative AI API key as a bearer token only at the
regional OCI inference hosts and only under /openai/v1 with GET, POST, and
DELETE, so the sandbox never holds the key and the key cannot reach any other
OCI surface.

The header comments carry the OCI-side setup (create the IAM policy before
the key, least-privilege statement, key creation and rotation), the
operations verified through the sandbox proxy with a real key (chat
completions with streaming, tool calling, vision input, embeddings, and the
Responses API), the OCI error messages operators will meet, and the realm
and signed-transport caveats.

Scoped to the profile YAML per #3906; the only code change is the entry in
the profile listing test, which enumerates providers/*.yaml.

Signed-off-by: Federico Kamelhar <federico.kamelhar@oracle.com>
2026-09-29 22:36:06 +00:00
Shiju ba16b9f2c7 fix(mcp): explain revision-scoped policy and rejections (#3850)
* fix(mcp): explain revision-scoped policy and rejections

Explain the selected-revision method set in profile output and policy docs.
Distinguish protocol and policy rejection causes and give a next step while
preserving authorization, response statuses, error codes and YAML keys.

Cover CLI serialization, revision selection, exact extension rules, deny
precedence and rejection before forwarding with focused regressions.

Signed-off-by: Shiju <shiju@nvidia.com>

* docs(mcp): correct HTTP cancellation revision support

Limit notifications/cancelled to the three 2025 revisions in the core
method matrix. State that MCP 2026-07-28 HTTP cancellation closes the
response stream, matching the runtime rejection and sessionless docs.

Signed-off-by: Shiju <shiju@nvidia.com>

---------

Signed-off-by: Shiju <shiju@nvidia.com>
2026-09-29 21:50:35 +00:00
Shiju 0ea0d31020 fix(supervisor): restore canonical stdin after connection loss (#3852)
* fix(supervisor): restore canonical stdin after connection loss

Probe idle SSH peers and enforce a receive deadline during transport I/O,
including writes blocked by a stalled relay. Release the dead attachment's
stdin lease through existing handler cleanup.

Retry denied write intent on later ordinary input without displacing a
healthy owner. Preserve explicit read-only, EOF and detach behavior, and
discard control bytes retained while input ownership was denied.

Cover half-open forwarding, blocked writes, healthy idle peers and competing
reconnects through the production supervisor frame bridge and real SSH.

Fixes #3648

Signed-off-by: Shiju <shiju@nvidia.com>

* docs(skills): describe read-only reconnect input retry

Explain what an openshell-cli user sees when automatic recovery reattaches before the supervisor closes the dead connection: the attachment reports read-only, later ordinary input retries stdin acquisition and prints `input enabled`, input typed while read-only is discarded, exit keys still detach, and an explicitly read-only viewer or a healthy owner is never affected.

Signed-off-by: Shiju <shiju@nvidia.com>

---------

Signed-off-by: Shiju <shiju@nvidia.com>
2026-09-29 20:28:55 +00:00
krishicks 33a8eac196 feat(helm): configure gateway OCSF JSONL output (#3876)
Previously, Helm installations could not enable the gateway OCSF JSONL
destination through chart values because generated `gateway.toml` omitted the
`openshell.gateway.ocsf_log` table.

Now, setting `server.ocsfLog.enabled` renders the path, optional schema
version, rotation, retention, and queue limits into gateway configuration.
Output is disabled by default. The default path, `/tmp/gateway-ocsf.jsonl`,
is writable in the gateway container with either the StatefulSet or
Deployment workload, so enabling output does not require persistent storage.
Invalid schema versions, rotation values, non-positive limits, or an empty
path while enabled fail chart rendering.

Additionally, `server.extraVolumes` and `server.extraVolumeMounts` add
operator-supplied volumes to the gateway pod, so operators who want records
to survive restarts can place the OCSF path on persistent storage without
replacing chart-generated configuration.

The gateway pod's default termination grace period rises from 5 to 30
seconds. Gateway shutdown can spend up to 10 seconds on supervisor session
cleanup before allowing 5 seconds to drain queued OCSF records, so the
5-second default risked a SIGKILL before the final records were written.
The grace period is only an upper bound: the gateway exits as soon as its
shutdown completes.

Refs #2762

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-09-29 19:23:43 +00:00
Piotr Mlocek c0eb3dbd30 fix(ci): restore repository permission vetters (#3875)
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-09-29 18:05:21 +00:00
krishicks cf1bbb965d docs: remove the architecture directory (#3799)
Remove architecture/. It was a constant source of merge conflicts, became
an effectively append-only log of the project, and was of dubious value.
Design records live in rfc/, crate details in crate READMEs, and user
documentation in docs/.

Move the git-ignored plans directory from architecture/plans to plans/,
keeping the old .gitignore entry. Remove the arch-doc-writer agents and
update AGENTS.md, CONTRIBUTING.md, skills, the feature request template,
and links in proto/, rfc/, and examples/ that pointed into architecture/.

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-09-29 17:35:22 +00:00
Shiju a6eefcf29a fix(network): preserve pipelined requests after chunked inspection (#3861)
* fix(network): preserve pipelined requests after chunked inspection

Stop chunked MCP and JSON-RPC body reads at each framing boundary so the
connection reader retains the next request for independent inspection.
Keep payload reads bounded by the remaining chunk length and scan framing
lines incrementally.

Cover buffered prefixes, fragmented framing, trailers, and malformed input.
Verify allowed and denied pipelined requests through both relay entry paths.

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(network): share bounded HTTP body inspection with GraphQL

Remove GraphQL's duplicate chunk decoder so all buffered HTTP inspectors
preserve the next request on a persistent connection. Keep GraphQL's
header checks, configured body limit and query classification.

Cover GraphQL trailers and fragmented framing, and exercise subsequent
request authorization for REST, GraphQL, MCP and JSON-RPC through both
persistent relay entry paths.

Signed-off-by: Shiju <shiju@nvidia.com>

---------

Signed-off-by: Shiju <shiju@nvidia.com>
2026-09-29 17:04:59 +00:00
krishicks 2ad77ad8a2 test(sandbox): bind ephemeral port in accepted loopback stream test (#3872)
The test reserved a loopback port, released it, and rebound it inside the
workload. A concurrent nextest process could claim the port in between,
failing the bind and surfacing only as a RecvError on the ready channel.
Bind port 0 in the workload and send the assigned address instead, and
report the workload error when the listener never becomes ready.

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-09-29 16:46:52 +00:00
krishicksandEvan Lezar a875add234 feat(server): write gateway OCSF events to JSONL (#3264)
* feat(server): write gateway OCSF events to JSONL

Previously, gateway security activity was available only in diagnostic output,
and events not associated with a sandbox, such as TLS certificate reloads, had
no independent structured record.

Now, configuring `openshell.gateway.ocsf_log` writes every gateway-produced
OCSF record to a bounded JSONL destination independently of `RUST_LOG`. The
destination supports daily or disabled rotation, retention limits, queue bounds,
and optional schema downgrade to OCSF 1.1 or 1.3.

Additionally, existing gateway emitters (TLS reloads, service routing, and
policy approval and auto-approval audits) emit structured events, so they reach
the JSONL destination, console shorthand, and the affected sandbox's log
stream. Records identify the gateway by its configured name in `device.uid` and
`device.name`, shared across replicas, with `device.hostname` identifying the
replica and `device.os` the gateway's operating system. Metrics and warnings
expose known best-effort losses.

Refs #2762

Signed-off-by: Kris Hicks <khicks@nvidia.com>

* fix(mxc): attribute ETW events to gateway

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Kris Hicks <khicks@nvidia.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>
Co-authored-by: Evan Lezar <elezar@nvidia.com>
2026-09-29 16:19:53 +00:00
Shiju 12ef86c285 fix(cli): keep policy and provider diagnostics readable (#3444)
Reuse character-safe truncation for policy history errors so a multibyte character cannot panic the table renderer.

Distinguish unavailable provider-profile YAML from an absent profile, display a bounded diagnostic, and preserve strict serialization and redacted object navigation.

Cover the actual CLI renderer and TUI display/navigation paths, including invalid and absent profiles, Unicode input, and redacted errors.

Signed-off-by: Shiju <shiju@nvidia.com>
2026-09-29 15:35:56 +00:00
Eric Curtin cfcc3733bd fix(e2e): stop sandbox leaks from async Drop cleanup (#3750)
* fix(e2e): stop sandbox leaks from async Drop cleanup

Closes #2922

SandboxGuard::Drop spawned a detached thread to delete the sandbox.
The thread got killed with the test process before the delete
finished. Switch to a blocking command in Drop, like ManagedCleanup
already does. Also wrap two tests' manual cleanup in RAII guards so
a panic does not leak a sandbox.

Signed-off-by: Eric Curtin <eric.curtin@docker.com>

* test(e2e): arm sandbox guards before create

Address review: install guards with explicit names first.

Signed-off-by: Eric Curtin <eric.curtin@docker.com>

---------

Signed-off-by: Eric Curtin <eric.curtin@docker.com>
2026-09-29 10:10:29 +00:00
Philippe MartinandJohn Myers 9cb72baa2e feat(docker): support corporate proxy CA bundles (#3549)
* feat(docker): support corporate proxy CA bundles

Closes #3545

Validate and stage operator-owned proxy CA bundles for Docker supervisors, add corporate proxy E2E coverage, and document the trust contract.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(docker): validate proxy config on startup

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(docker): use the E2E workload image for proxy tests

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(docker): generate strict corporate proxy certificates

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(docker): surface intercepted TLS fixture errors

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(docker): drain buffered TLS proxy data

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(docker): relay intercepted HTTP deterministically

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: Philippe Martin <phmartin@redhat.com>
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-09-29 05:15:02 +00:00
Divesh 2fe5a0e19c perf(kubernetes): use a TCP readiness probe for the supervisor (#3700)
- Kubernetes now checks supervisor readiness by connecting to TCP port 5501
- Stop starting a supervisor process in every sandbox each second
- The supervisor opens the port only while its gateway session is up
- Accept IPv4 and IPv6 probes, even when net.ipv6.bindv6only is set
- Keep the health socket for Docker, Podman, and debugging
- Add tests and update the docs

Signed-off-by: divesh <dgude@nvidia.com>
2026-09-29 04:47:12 +00:00
Drew Newberry acbac9cb79 feat(sandbox): add main restart policy (#2798)
* feat(sandbox): add main restart policy

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): harden policy-driven restarts

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): address restart review feedback

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>

* fix(sandbox): port restart policy to current runtime

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* perf(sandbox): restart promptly after terminal delivery

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>
2026-09-29 02:18:30 +00:00
Shiju 1358941b81 feat(mcp): inspect requests with Tower-selected protocol profiles (#3335)
* fix(sandbox-backend): sort boundary request objects before hashing

Sort boundary request objects recursively before hashing so serde_json's
preserve_order feature cannot change digest identity. Cover canonical
bytes, envelope round trips, and rejection of modified provider values
and operations.

Signed-off-by: Shiju <shiju@nvidia.com>

* feat(mcp): upgrade tower-mcp-types to 0.22.2

Upgrade tower-mcp-types from 0.12.0 to an exact-pinned 0.22.2 and use its
inspection APIs to validate MCP requests against the selected revision.
Carry inspection metadata into policy evaluation and validate requests
after header rewriting, before forwarding.

Add explicit support for the sessionless 2026-07-28 revision while keeping
2025-11-25 as the default. Validate per-request metadata and standard HTTP
header mirrors, and support discovery, tools, and subscription requests.

Delegate batch availability and parameter schemas to Tower. Share typed
request names between policy and HTTP checks, retain the local batch
resource cap, and centralize MCP policy version parsing and ordering.

Keep supported MCP revisions and shared allowlist parsing in the canonical
policy schema; core re-exports those types. Tower owns wire-profile
semantics, and every supported policy revision must map to the matching
inspector profile.

Reject duplicate JSON keys, invalid known-method parameters, unavailable
methods, and unsupported batches. Keep exact extension allow rules and
deny precedence. Document request inspection boundaries and add unit,
forwarding, and sandbox coverage.

Refs #2174.

Signed-off-by: Shiju <shiju@nvidia.com>

* test(mcp): prove authorization at the forwarding boundary

Cover March batch denial in both member orders, valid and malformed
controls, and audit behavior across both relay entry paths. Exercise real
middleware tool rewrites with matching metadata and assert the exact
upstream representation or zero forwarded bytes.

Verify legacy bodyless SSE GET remains usable while GET tool bodies and
unsupported DELETE cleanup are rejected. Clarify request-selected profile
and middleware mutation comments without changing production behavior.

Signed-off-by: Shiju <shiju@nvidia.com>

* test(mcp): exercise permitted profiles through the sandbox proxy

Cover March and June singleton policies and select November and July
separately under one endpoint allowlist. Capture upstream tool receipts
to distinguish proxy policy denial from an upstream rejection.

Extend middleware rewrite coverage to June and multi-version policies,
and preserve the sessionless discovery and subscription checks through
the shared fixture helpers.

Signed-off-by: Shiju <shiju@nvidia.com>

* test(kubernetes): box the admission check future

Keep the admission test future below Clippy's size limit when the
workspace dependency features are unified.

Signed-off-by: Shiju <shiju@nvidia.com>

* test(mcp): reuse the forwarding fixture identity cache

Share the binary identity cache across protocol-profile cases, matching
the proxy lifecycle and avoiding repeated hashes of the test executable.
Keep procfs authorization and all forwarding assertions intact.

Signed-off-by: Shiju <shiju@nvidia.com>

---------

Signed-off-by: Shiju <shiju@nvidia.com>
2026-09-28 20:48:50 +00:00
Evan Lezar b77f5ddfc1 test(install): support Bash 3.2 mock capture (#3790)
Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-09-28 20:41:52 +00:00
Drew Newberry e63cfa1182 fix(cli): keep SSH forwards owned by spawned process (#3759)
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-28 19:34:13 +00:00
Drew Newberry 36b0386c92 feat(cli): detach sandbox sessions with Ctrl-D (#3744)
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-28 09:36:14 -07:00
Shiju 45e3308d39 fix(network): refuse protocol upgrades on JSON-RPC and MCP endpoints (#3753)
* fix(network): refuse protocol upgrades on JSON-RPC and MCP endpoints

JSON-RPC and MCP rules apply to each HTTP request, but the proxy could
forward a request that also carried upgrade headers. After an upstream
answered 101, route selection and the forward proxy relayed the
connection without inspection.

Refuse any request that carries an Upgrade header on JSON-RPC-family
endpoints before the L7 policy decision, in every enforcement mode.
Share the check with the existing h2c refusal and call it from
relay_jsonrpc as well. Record the refusal as a policy denial and answer
with the unsupported_l7_protocol error, because no policy rule can
allow the request.

If a JSON-RPC-family endpoint still receives 101, close the connection
instead of relaying raw bytes. Document the refusal and the WebSocket
alternative.

Signed-off-by: Shiju <shiju@nvidia.com>

* docs(observability): remove duplicate protocol error definition

Keep unsupported_l7_protocol in the response error-code list and retain its explanation in the policy troubleshooting table.

Signed-off-by: Shiju <shiju@nvidia.com>

---------

Signed-off-by: Shiju <shiju@nvidia.com>
2026-09-28 16:08:19 +00:00
Evan Lezar eef8bec0c9 test(e2e): run podman suite with tmachine (#3637)
* test(e2e): remove superseded podman userns coverage

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(tmachine): run podman e2e archive

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(tmachine): generate podman e2e archive inventory

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-09-28 11:15:33 +00:00
Drew Newberry 9f60f55c6b fix(server): retry transient sandbox CAS conflicts (#3501)
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-28 09:23:47 +00:00
Evan Lezar 0c29d8e061 test(conformance): migrate file transfer scenarios (#3597)
Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-09-28 09:02:44 +00:00
Evan Lezar c63f8ce564 test(cli): migrate gateway-free smoke coverage (#3641)
* test(cli): migrate gateway-free smoke coverage

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(e2e): remove migrated tests from podman CI

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-09-28 09:00:30 +00:00
Drew Newberry 6f00d5cacc fix(e2e): reserve distinct corporate proxy fixture ports (#3761)
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-28 07:10:23 +00:00
Drew Newberry 7a4da31249 fix(supervisor): keep session retries during startup (#3765)
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-28 07:06:19 +00:00
ansjindal d009f30121 feat(helm): make cluster-scoped RBAC optional (#3459)
The gateway chart always rendered the ClusterRole and ClusterRoleBinding, so
every install and upgrade required cluster-admin even when only namespaced
objects were needed. Installers that are namespace-admin GitOps or platform
controllers could not run the release at all, and clusters where cluster-scoped
RBAC is owned by a separate team had no supported way to split the install.

Add an rbac values block so a cluster-admin can apply the cluster-scoped objects
once and a namespace-admin can install and upgrade the release without
cluster-scoped permissions:

  rbac:
    create: true
    clusterScoped:
      create: true
      clusterRoleName: ""
      clusterRoleBindingName: ""

rbac.clusterScoped.create gates the ClusterRole and ClusterRoleBinding, and is
independent of the workspace mode. rbac.create additionally gates the namespaced
sandbox Role and RoleBinding, which matters because Kubernetes escalation
prevention stops an installer holding only the built-in admin role from creating
a Role that grants agents.x-k8s.io verbs it does not itself hold. The certgen
hook and credential driver RBAC keep their existing flags.

Both flags default to true, so current installs are unchanged. The helpers treat
a missing rbac block as enabled so upgrades with --reuse-values do not drop RBAC,
matching the existing workspaceResources pattern. The ClusterRoleBinding roleRef
follows clusterRoleName so a separately applied ClusterRole can carry a name the
cluster-admin chooses.

Document the migration for a release that already owns the cluster-scoped
objects: Helm deletes objects that leave the manifest, so annotate them with
helm.sh/resource-policy=keep before setting the flag, otherwise the gateway
loses TokenReview until a cluster-admin re-applies them.

Signed-off-by: ansjindal <ansjindal@nvidia.com>
2026-09-28 05:50:07 +00:00
Drew Newberry 6648bd0c29 perf(sandbox): wake orphan reaper on child exit (#3757)
Closes #3756

Replace the idle 50 ms procfs scan with SIGCHLD notifications and a 30 second recovery sweep. Preserve managed-child wait ownership and cover idle and exit behavior in isolated tests.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
v0.1.2
2026-09-28 03:08:16 +00:00
Johnny Greco 87929ad130 docs: keep page URLs aligned with file paths (#3713)
* docs: align page file names and nav labels with published URLs

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs: redirect moved dev pages and fix agent guide redirects

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* ci(docs): check that page URLs match their file paths

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(docs): align navigation checks with repo conventions

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs: omit historical redirect aliases

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(docs): preserve published overview and TypeScript URLs

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(docs): redirect unversioned overview and TypeScript URLs

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

---------

Signed-off-by: Johnny Greco <jogreco@nvidia.com>
2026-09-28 02:08:54 +00:00
Johnny Greco 8934b74a85 fix(docs): sync redirects with versioned snapshots (#3754)
* fix(docs): sync redirects with versioned snapshots

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(docs): make redirect validation types explicit

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

---------

Signed-off-by: Johnny Greco <jogreco@nvidia.com>
2026-09-28 00:27:56 +00:00
Drew Newberry c9da461a58 fix(ci): keep snap canary sandbox name within limit (#3751)
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-27 19:20:30 +00:00
Drew Newberry d1a19c70ee fix(homebrew): overwrite generated gateway config during migration (#3752)
Closes #3746

Use Homebrew atomic_write for exact legacy config migrations and keep the first-install write distinct in the formula test.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-27 18:56:08 +00:00
Davanum Srinivas f37d89b584 fix(supervisor-network): keep workload bytes read with a mediated CONNECT header (#3745)
The proxy read the synthesized CONNECT header of a mediated open with a
full-width read into an 8192-byte buffer through a BufReader of the same
size, so the workload's request could arrive in the same read. The
CONNECT path never looked past the header, and the request was lost
until the workload timed out. Take the header out of the BufReader with
fill_buf and consume only through the terminator, so the bytes behind it
stay buffered for the relay.

Signed-off-by: Davanum Srinivas <dsrinivas@nvidia.com>
2026-09-27 17:49:05 +00:00
Drew Newberry a3ed8c79cf docs: refresh architecture and extensibility pages (#3741)
* docs(readme): show Rust SDK installation command

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(architecture): lead with policy enforcement

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(architecture): reuse README how it works text

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(architecture): refresh architecture page and diagrams

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(extensibility): simplify extension points diagram and overview

Redraw the extension points diagram as two left-to-right rows for the
data plane and control plane, describe which layers each extension point
extends, and move authentication after Building Extensions as
Authenticating Extensions with updated cross-page links.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-27 02:18:29 +00:00
Drew Newberry 2f1ea658fe docs(policy): clarify sandbox-local loopback access (#3740)
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-26 18:18:28 +00:00
Drew Newberry 4ce767fc0c docs: describe 0.1 series in release callouts (#3732)
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
v0.1.1
2026-09-26 03:06:48 +00:00
Drew Newberry a67567e583 fix(snap): require mTLS for the snap gateway (#3726)
* fix(snap): require mTLS for the snap gateway

Replace the installer opt-in with an authenticated snap gateway. The wrapper
no longer forces plaintext, so the gateway serves TLS from the bundle it
already generates in $SNAP_COMMON/tls. The install hook writes a config that
enables mTLS user auth instead of unauthenticated access, and a new
post-refresh hook migrates the exact legacy default on existing installs.

install.sh waits for the gateway, detects whether it serves TLS, copies the
client bundle into the target user's snap state directory, and registers the
gateway over HTTPS. Older plaintext snap revisions still register over HTTP
with a warning. The release canary asserts mTLS auth and HTTPS registration.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(snap): pass config preflight and detect the mTLS gateway reliably

An explicit [openshell.gateway.mtls_auth] table fails config preflight,
which validates mTLS auth before the local TLS bundle supplies the client
CA. Write a default that pins the Docker driver instead; with the wrapper's
TLS bundle the gateway requires client certificates and enables mTLS user
auth automatically, as the native packages do.

The mTLS gateway rejects TLS handshakes without a client certificate, and
it still answers plaintext loopback HTTP for sandbox service routing, so the
installer could misdetect it as a legacy plaintext gateway. Probe HTTPS with
the root-owned client bundle, and treat a gateway as legacy only when a
plaintext gRPC Health call succeeds.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* chore(snap): simplify install hook comment

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(snap): migrate insecure gateway configs on refresh

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(snap): simplify mTLS detection and config migration

Detect the mTLS snap from the installed revision's post-refresh hook instead
of probing plaintext gRPC, and drop the scheme global. Remove the installer's
pre-hook config fallback, which is dead now that every channel ships the
install hook and which wrote the insecure default. Give the install hook a
single write path with a simple backup name, and shorten the manual client
certificate steps.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(snap): stop keeping a copy of replaced insecure configs

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(install): make the snap an opt-in install method

Stop selecting the OpenShell snap just because the snap command exists. Linux
installs default to the Debian or RPM package; OPENSHELL_INSTALL_METHOD=snap
(or deb, rpm) selects the package explicitly. Hosts that already have the
OpenShell snap keep refreshing it rather than gaining a second gateway on the
same port. The release canary and snap repro script opt in explicitly.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(snap): let the gateway auto-detect its compute driver

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(snap): restart the gateway after refresh

Published revisions use refresh-mode: endure, and snapd honors the old
revision's setting during a refresh, so the plaintext gateway kept running
with the migrated config unused until a manual restart. Restart the gateway
from the post-refresh hook so the mTLS config takes effect immediately.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(snap): drop refresh notes from the snap description

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(snap): trim snap refresh notes from installation docs

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-26 01:50:02 +00:00
Johnny Greco f155899527 docs(tutorials): run Pi with OpenRouter (#3722)
* docs(tutorials): run Pi with OpenRouter

Switch the Pi tutorial from Anthropic to OpenRouter, add fd to the Pi image
so Pi does not try to download it from GitHub inside the sandbox, and rename
the page to Run Pi with OpenRouter with a dev redirect from the old URL.

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(tutorials): drop redirect for dev-only Pi page

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(tutorials): make the Pi tutorial easier to follow

Add prerequisites and steps, explain providers and the profile fields in
plain terms, inspect the sandbox while Pi is still running, and add clean-up
and troubleshooting sections.

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(tutorials): clarify Pi's providers and sandbox additions

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(tutorials): make the Pi tutorial more conversational

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

---------

Signed-off-by: Johnny Greco <jogreco@nvidia.com>
2026-09-26 01:38:59 +00:00
grs b52eed72e3 test(binary-identity): stabilize procfs identity fixtures (#3708)
* test(binary-identity): stabilize procfs identity fixtures

Signed-off-by: Gordon Sim <gsim@redhat.com>

* test(binary-identity): wait for child fixture readiness

Signed-off-by: Gordon Sim <gsim@redhat.com>

---------

Signed-off-by: Gordon Sim <gsim@redhat.com>
2026-09-26 00:01:04 +00:00
Drew Newberry d4f5034d7f fix(gateway)!: make the WebSocket tunnel opt-in (#3727)
The /_ws_tunnel endpoint pipes a WebSocket into the full gRPC service. On a
plaintext loopback gateway any web page could open it, since browsers do not
apply CORS to WebSocket upgrades. Mount the tunnel only when
enable_websocket_tunnel is set (config file, --enable-websocket-tunnel, or
OPENSHELL_ENABLE_WEBSOCKET_TUNNEL; server.enableWebsocketTunnel in Helm).

BREAKING CHANGE: gateways behind an authenticating edge proxy must set
enable_websocket_tunnel = true for CLI edge-tunnel connections.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-25 23:43:18 +00:00
Drew Newberry 496ebba293 docs(readme): remove alpha status badge (#3723)
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
v0.1.0
2026-09-25 21:04:39 +00:00
Drew NewberryandPiotr Mlocek 73a181d32f docs: streamline README, add policy prover to architecture docs (#3718)
* docs(readme): streamline README and move reference detail to docs

Restructure the README as a short path from overview to quickstart to
further reading. Move prerelease install steps into the installation
guide and telemetry build flags into a new observability page. Fix
broken docs links and outdated runtime and credential descriptions.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(readme): describe 0.1.0 as adding new isolation primitives

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(architecture): add policy prover as a gateway component

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs: describe OpenShell as a runtime for fleets of autonomous AI agents

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(architecture): describe policy prover as formal verification

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(architecture): name OpenShell Sandbox in component table

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(architecture): fold isolation backend into supervisor row

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(architecture): mention formal verification in overview

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(run-agent): use the published OpenCode image in the first-agent guide

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(inference): correct provider examples and readiness

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(providers): correct Google binding and provider selection

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(readme): sharpen value prop, how it works, and explore further

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(run-agent): use example Anthropic profile and add policy advisor step

Import the example Anthropic profile, which now allows OpenCode, instead
of editing it with sed. Add a step that shows how to review and approve
mechanistic policy proposals as the agent needs more access.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(providers): add OpenRouter example for OpenCode

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(run-agent): run OpenCode against OpenRouter with a free model

Add an example OpenRouter provider profile scoped to OpenCode and switch
the first-agent guide to it, using a free Nemotron model so readers do
not need OpenRouter credits. Revert the OpenCode binary added to the
example Anthropic profile.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(readme): link first-agent guide and add agent skills section

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Co-authored-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-09-25 13:53:45 -07:00
Piotr Mlocek 6f596828ad docs(fern): publish ordered version snapshots (#3721)
* docs(fern): publish and order versioned release docs

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(fern): remove version availability badges

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(fern): keep version badges optional

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

---------

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-09-25 19:56:05 +00:00
Johnny Greco d7f921190b docs(policy): refresh policy documentation and references (#3563)
* docs(policy): correct schema and default policy guidance

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): add network recipes and update command reference

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): organize lifecycle guidance and troubleshooting

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): split policy overview into concepts and management tasks

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): reorganize network recipes as a cookbook

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): restructure schema reference by field group and protocol

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): align troubleshooting, advisor, and reference pages

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): fix first policy tutorial and security guidance

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): keep overview high level and move network rules to their own page

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): focus policy management on CLI workflows and remove command reference

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): clarify policy views and sandbox deletion in management guide

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): streamline network rule concepts and examples

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): correct request path wildcard semantics

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): rewrite policy advisor guide for clarity

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): clarify policy advisor scope, setup, and review

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): rewrite policy prover guide for clarity

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): explain the two uses of the policy prover

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): describe policy prover uses, boundaries, and coverage

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): place prover before advisor and troubleshooting last

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): remove unsupported CI guidance from prover page

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): tighten policy prover introduction

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): move policy change behavior into management guide

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): name prover check types and note expanding coverage

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): prefix prover and advisor sidebar labels

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): streamline policy schema reference

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): place default policy before schema reference

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): fold troubleshooting into policy management guide

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): correct tutorial log samples and GitHub push policy steps

The first policy tutorial said the 403 body begins with error, policy, and
rule, but the proxy serializes the body with sorted keys. Its log samples also
showed the wrong CONNECT deny reason for a sandbox without network rules, and
the L7 deny sample omitted the :443 authority, the `l7` engine, and the reason
tag that the shorthand formatter emits.

The GitHub tutorial filtered denials with `--level warn`, which hides the INFO
level OCSF policy events, and showed the retired key=value log format. Its
hand-written policy also omitted /bin from the restrictive default, so
`policy set` would reject the file for removing a filesystem path on a live
sandbox. Start from `policy get --base` and add only the network rules.

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): improve flow and terminology across policy pages

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): correct network rule matching and protocol details

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): align policy management steps with CLI behavior

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): correct policy advisor proposal and approval details

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): correct policy section, default, and schema details

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): correct prover installation and coverage limits

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): recommend tls skip for server-first protocols

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): fix stale baseline path and interpreter examples

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): move policy pages under how-it-works and fix links

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): align native TCP guidance in security best practices

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): restore policy.local and policy DNS details from main

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): state exact glob matching rules

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

---------

Signed-off-by: Johnny Greco <jogreco@nvidia.com>
2026-09-25 18:19:16 +00:00
Divesh Chowdary 585bcf0b82 (Fix) ha sandbox resilience with k8s (#3644)
* fix(server): retry internal store updates on version conflict

- Re-read and reapply internal CAS updates (expected version 0) up to 5 times on conflict
- Client-supplied versions still fail on conflict
- Add concurrent-writer test

Signed-off-by: divesh <dgude@nvidia.com>

* fix(kubernetes): keep sandboxes running while the supervisor reconnects

- Treat a running but not-Ready supervisor Pod as degraded, not unavailable
- Report degraded sandboxes as not ready without suspending them
- Suspend only when the supervisor Pod is missing, terminated, or deleting
- Add availability test

Signed-off-by: divesh <dgude@nvidia.com>

* fix(kubernetes): skip fence generation check for suspended sandboxes

- Check the fence generation only for running or bootstrapping sandboxes
- Stop re-suspending stopped sandboxes and logging a warning every reconcile

Signed-off-by: divesh <dgude@nvidia.com>

* fix(kubernetes): rank degraded supervisor above unknown dependencies

- Report a degraded supervisor as unavailable even when another dependency read is unknown
- Extract dependency aggregation and readiness mapping into pure helpers
- Add mixed degraded and unknown regression test

Signed-off-by: divesh <dgude@nvidia.com>

---------

Signed-off-by: divesh <dgude@nvidia.com>
2026-09-25 18:10:55 +00:00