Resource admission now accepts the managed workspace volume when its options
match the workload container's final identity, or are empty for volumes created
by older gateways. The channel volume still requires empty options.
Also remove the unused rootless driver field, keep user-mount control-path
checks as on main, restore the merge-upload coverage from the deleted Docker
custom_image e2e, and drop formatting-only churn.
Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>
Podman now creates the managed /sandbox volume with uid/gid options for the
resolved workload identity, so the workload starts directly as that identity.
This removes the root-then-drop workspace chown and fixes rootful sandboxes whose
image USER or policy run_as_user could not write to a root-owned /sandbox.
Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>
Add an example oci-genai inference profile for Oracle Cloud Infrastructure
Generative AI through its OpenAI-compatible endpoint. The profile injects a
compartment-scoped Generative AI API key as a bearer token only at the
regional OCI inference hosts and only under /openai/v1 with GET, POST, and
DELETE, so the sandbox never holds the key and the key cannot reach any other
OCI surface.
The header comments carry the OCI-side setup (create the IAM policy before
the key, least-privilege statement, key creation and rotation), the
operations verified through the sandbox proxy with a real key (chat
completions with streaming, tool calling, vision input, embeddings, and the
Responses API), the OCI error messages operators will meet, and the realm
and signed-transport caveats.
Scoped to the profile YAML per #3906; the only code change is the entry in
the profile listing test, which enumerates providers/*.yaml.
Signed-off-by: Federico Kamelhar <federico.kamelhar@oracle.com>
* fix(mcp): explain revision-scoped policy and rejections
Explain the selected-revision method set in profile output and policy docs.
Distinguish protocol and policy rejection causes and give a next step while
preserving authorization, response statuses, error codes and YAML keys.
Cover CLI serialization, revision selection, exact extension rules, deny
precedence and rejection before forwarding with focused regressions.
Signed-off-by: Shiju <shiju@nvidia.com>
* docs(mcp): correct HTTP cancellation revision support
Limit notifications/cancelled to the three 2025 revisions in the core
method matrix. State that MCP 2026-07-28 HTTP cancellation closes the
response stream, matching the runtime rejection and sessionless docs.
Signed-off-by: Shiju <shiju@nvidia.com>
---------
Signed-off-by: Shiju <shiju@nvidia.com>
Add an oci-image feature testsuite that runs against installed artifacts
on Docker rootful, Podman rootful, and Podman rootless tmachine guests.
It covers custom WORKDIR placement, image content and ownership,
workspace writes from the main process and exec, file transfer, OCI
user identity, the managed /sandbox fallback, and rejection of an
unwritable WORKDIR.
Replace the Docker-only custom_image e2e with the shared suite and keep
podman_oci_identity focused on Podman image-ID pinning and identity.
Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>
* fix(supervisor): restore canonical stdin after connection loss
Probe idle SSH peers and enforce a receive deadline during transport I/O,
including writes blocked by a stalled relay. Release the dead attachment's
stdin lease through existing handler cleanup.
Retry denied write intent on later ordinary input without displacing a
healthy owner. Preserve explicit read-only, EOF and detach behavior, and
discard control bytes retained while input ownership was denied.
Cover half-open forwarding, blocked writes, healthy idle peers and competing
reconnects through the production supervisor frame bridge and real SSH.
Fixes#3648
Signed-off-by: Shiju <shiju@nvidia.com>
* docs(skills): describe read-only reconnect input retry
Explain what an openshell-cli user sees when automatic recovery reattaches before the supervisor closes the dead connection: the attachment reports read-only, later ordinary input retries stdin acquisition and prints `input enabled`, input typed while read-only is discarded, exit keys still detach, and an explicitly read-only viewer or a healthy owner is never affected.
Signed-off-by: Shiju <shiju@nvidia.com>
---------
Signed-off-by: Shiju <shiju@nvidia.com>
Drop image-volume admission (split to a follow-up), the Docker and core
mount-validation refactor, redundant per-arm Podman mount checks, and
unrelated doc and test churn. Keep image VOLUME masking validation for
custom workdirs.
Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>
Previously, Helm installations could not enable the gateway OCSF JSONL
destination through chart values because generated `gateway.toml` omitted the
`openshell.gateway.ocsf_log` table.
Now, setting `server.ocsfLog.enabled` renders the path, optional schema
version, rotation, retention, and queue limits into gateway configuration.
Output is disabled by default. The default path, `/tmp/gateway-ocsf.jsonl`,
is writable in the gateway container with either the StatefulSet or
Deployment workload, so enabling output does not require persistent storage.
Invalid schema versions, rotation values, non-positive limits, or an empty
path while enabled fail chart rendering.
Additionally, `server.extraVolumes` and `server.extraVolumeMounts` add
operator-supplied volumes to the gateway pod, so operators who want records
to survive restarts can place the OCSF path on persistent storage without
replacing chart-generated configuration.
The gateway pod's default termination grace period rises from 5 to 30
seconds. Gateway shutdown can spend up to 10 seconds on supervisor session
cleanup before allowing 5 seconds to drain queued OCSF records, so the
5-second default risked a SIGKILL before the final records were written.
The grace period is only an upper bound: the gateway exits as soon as its
shutdown completes.
Refs #2762
Signed-off-by: Kris Hicks <khicks@nvidia.com>
Remove architecture/. It was a constant source of merge conflicts, became
an effectively append-only log of the project, and was of dubious value.
Design records live in rfc/, crate details in crate READMEs, and user
documentation in docs/.
Move the git-ignored plans directory from architecture/plans to plans/,
keeping the old .gitignore entry. Remove the arch-doc-writer agents and
update AGENTS.md, CONTRIBUTING.md, skills, the feature request template,
and links in proto/, rfc/, and examples/ that pointed into architecture/.
Signed-off-by: Kris Hicks <khicks@nvidia.com>
* fix(network): preserve pipelined requests after chunked inspection
Stop chunked MCP and JSON-RPC body reads at each framing boundary so the
connection reader retains the next request for independent inspection.
Keep payload reads bounded by the remaining chunk length and scan framing
lines incrementally.
Cover buffered prefixes, fragmented framing, trailers, and malformed input.
Verify allowed and denied pipelined requests through both relay entry paths.
Signed-off-by: Shiju <shiju@nvidia.com>
* fix(network): share bounded HTTP body inspection with GraphQL
Remove GraphQL's duplicate chunk decoder so all buffered HTTP inspectors
preserve the next request on a persistent connection. Keep GraphQL's
header checks, configured body limit and query classification.
Cover GraphQL trailers and fragmented framing, and exercise subsequent
request authorization for REST, GraphQL, MCP and JSON-RPC through both
persistent relay entry paths.
Signed-off-by: Shiju <shiju@nvidia.com>
---------
Signed-off-by: Shiju <shiju@nvidia.com>
The test reserved a loopback port, released it, and rebound it inside the
workload. A concurrent nextest process could claim the port in between,
failing the bind and surfacing only as a RecvError on the ready channel.
Bind port 0 in the workload and send the assigned address instead, and
report the workload error when the listener never becomes ready.
Signed-off-by: Kris Hicks <khicks@nvidia.com>
* feat(server): write gateway OCSF events to JSONL
Previously, gateway security activity was available only in diagnostic output,
and events not associated with a sandbox, such as TLS certificate reloads, had
no independent structured record.
Now, configuring `openshell.gateway.ocsf_log` writes every gateway-produced
OCSF record to a bounded JSONL destination independently of `RUST_LOG`. The
destination supports daily or disabled rotation, retention limits, queue bounds,
and optional schema downgrade to OCSF 1.1 or 1.3.
Additionally, existing gateway emitters (TLS reloads, service routing, and
policy approval and auto-approval audits) emit structured events, so they reach
the JSONL destination, console shorthand, and the affected sandbox's log
stream. Records identify the gateway by its configured name in `device.uid` and
`device.name`, shared across replicas, with `device.hostname` identifying the
replica and `device.os` the gateway's operating system. Metrics and warnings
expose known best-effort losses.
Refs #2762
Signed-off-by: Kris Hicks <khicks@nvidia.com>
* fix(mxc): attribute ETW events to gateway
Signed-off-by: Evan Lezar <elezar@nvidia.com>
---------
Signed-off-by: Kris Hicks <khicks@nvidia.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>
Co-authored-by: Evan Lezar <elezar@nvidia.com>
Reuse character-safe truncation for policy history errors so a multibyte character cannot panic the table renderer.
Distinguish unavailable provider-profile YAML from an absent profile, display a bounded diagnostic, and preserve strict serialization and redacted object navigation.
Cover the actual CLI renderer and TUI display/navigation paths, including invalid and absent profiles, Unicode input, and redacted errors.
Signed-off-by: Shiju <shiju@nvidia.com>
* fix(e2e): stop sandbox leaks from async Drop cleanup
Closes#2922
SandboxGuard::Drop spawned a detached thread to delete the sandbox.
The thread got killed with the test process before the delete
finished. Switch to a blocking command in Drop, like ManagedCleanup
already does. Also wrap two tests' manual cleanup in RAII guards so
a panic does not leak a sandbox.
Signed-off-by: Eric Curtin <eric.curtin@docker.com>
* test(e2e): arm sandbox guards before create
Address review: install guards with explicit names first.
Signed-off-by: Eric Curtin <eric.curtin@docker.com>
---------
Signed-off-by: Eric Curtin <eric.curtin@docker.com>
- Kubernetes now checks supervisor readiness by connecting to TCP port 5501
- Stop starting a supervisor process in every sandbox each second
- The supervisor opens the port only while its gateway session is up
- Accept IPv4 and IPv6 probes, even when net.ipv6.bindv6only is set
- Keep the health socket for Docker, Podman, and debugging
- Add tests and update the docs
Signed-off-by: divesh <dgude@nvidia.com>
* fix(sandbox-backend): sort boundary request objects before hashing
Sort boundary request objects recursively before hashing so serde_json's
preserve_order feature cannot change digest identity. Cover canonical
bytes, envelope round trips, and rejection of modified provider values
and operations.
Signed-off-by: Shiju <shiju@nvidia.com>
* feat(mcp): upgrade tower-mcp-types to 0.22.2
Upgrade tower-mcp-types from 0.12.0 to an exact-pinned 0.22.2 and use its
inspection APIs to validate MCP requests against the selected revision.
Carry inspection metadata into policy evaluation and validate requests
after header rewriting, before forwarding.
Add explicit support for the sessionless 2026-07-28 revision while keeping
2025-11-25 as the default. Validate per-request metadata and standard HTTP
header mirrors, and support discovery, tools, and subscription requests.
Delegate batch availability and parameter schemas to Tower. Share typed
request names between policy and HTTP checks, retain the local batch
resource cap, and centralize MCP policy version parsing and ordering.
Keep supported MCP revisions and shared allowlist parsing in the canonical
policy schema; core re-exports those types. Tower owns wire-profile
semantics, and every supported policy revision must map to the matching
inspector profile.
Reject duplicate JSON keys, invalid known-method parameters, unavailable
methods, and unsupported batches. Keep exact extension allow rules and
deny precedence. Document request inspection boundaries and add unit,
forwarding, and sandbox coverage.
Refs #2174.
Signed-off-by: Shiju <shiju@nvidia.com>
* test(mcp): prove authorization at the forwarding boundary
Cover March batch denial in both member orders, valid and malformed
controls, and audit behavior across both relay entry paths. Exercise real
middleware tool rewrites with matching metadata and assert the exact
upstream representation or zero forwarded bytes.
Verify legacy bodyless SSE GET remains usable while GET tool bodies and
unsupported DELETE cleanup are rejected. Clarify request-selected profile
and middleware mutation comments without changing production behavior.
Signed-off-by: Shiju <shiju@nvidia.com>
* test(mcp): exercise permitted profiles through the sandbox proxy
Cover March and June singleton policies and select November and July
separately under one endpoint allowlist. Capture upstream tool receipts
to distinguish proxy policy denial from an upstream rejection.
Extend middleware rewrite coverage to June and multi-version policies,
and preserve the sessionless discovery and subscription checks through
the shared fixture helpers.
Signed-off-by: Shiju <shiju@nvidia.com>
* test(kubernetes): box the admission check future
Keep the admission test future below Clippy's size limit when the
workspace dependency features are unified.
Signed-off-by: Shiju <shiju@nvidia.com>
* test(mcp): reuse the forwarding fixture identity cache
Share the binary identity cache across protocol-profile cases, matching
the proxy lifecycle and avoiding repeated hashes of the test executable.
Keep procfs authorization and all forwarding assertions intact.
Signed-off-by: Shiju <shiju@nvidia.com>
---------
Signed-off-by: Shiju <shiju@nvidia.com>
* fix(network): refuse protocol upgrades on JSON-RPC and MCP endpoints
JSON-RPC and MCP rules apply to each HTTP request, but the proxy could
forward a request that also carried upgrade headers. After an upstream
answered 101, route selection and the forward proxy relayed the
connection without inspection.
Refuse any request that carries an Upgrade header on JSON-RPC-family
endpoints before the L7 policy decision, in every enforcement mode.
Share the check with the existing h2c refusal and call it from
relay_jsonrpc as well. Record the refusal as a policy denial and answer
with the unsupported_l7_protocol error, because no policy rule can
allow the request.
If a JSON-RPC-family endpoint still receives 101, close the connection
instead of relaying raw bytes. Document the refusal and the WebSocket
alternative.
Signed-off-by: Shiju <shiju@nvidia.com>
* docs(observability): remove duplicate protocol error definition
Keep unsupported_l7_protocol in the response error-code list and retain its explanation in the policy troubleshooting table.
Signed-off-by: Shiju <shiju@nvidia.com>
---------
Signed-off-by: Shiju <shiju@nvidia.com>
The gateway chart always rendered the ClusterRole and ClusterRoleBinding, so
every install and upgrade required cluster-admin even when only namespaced
objects were needed. Installers that are namespace-admin GitOps or platform
controllers could not run the release at all, and clusters where cluster-scoped
RBAC is owned by a separate team had no supported way to split the install.
Add an rbac values block so a cluster-admin can apply the cluster-scoped objects
once and a namespace-admin can install and upgrade the release without
cluster-scoped permissions:
rbac:
create: true
clusterScoped:
create: true
clusterRoleName: ""
clusterRoleBindingName: ""
rbac.clusterScoped.create gates the ClusterRole and ClusterRoleBinding, and is
independent of the workspace mode. rbac.create additionally gates the namespaced
sandbox Role and RoleBinding, which matters because Kubernetes escalation
prevention stops an installer holding only the built-in admin role from creating
a Role that grants agents.x-k8s.io verbs it does not itself hold. The certgen
hook and credential driver RBAC keep their existing flags.
Both flags default to true, so current installs are unchanged. The helpers treat
a missing rbac block as enabled so upgrades with --reuse-values do not drop RBAC,
matching the existing workspaceResources pattern. The ClusterRoleBinding roleRef
follows clusterRoleName so a separately applied ClusterRole can carry a name the
cluster-admin chooses.
Document the migration for a release that already owns the cluster-scoped
objects: Helm deletes objects that leave the manifest, so annotate them with
helm.sh/resource-policy=keep before setting the flag, otherwise the gateway
loses TokenReview until a cluster-admin re-applies them.
Signed-off-by: ansjindal <ansjindal@nvidia.com>
Closes#3756
Replace the idle 50 ms procfs scan with SIGCHLD notifications and a 30 second recovery sweep. Preserve managed-child wait ownership and cover idle and exit behavior in isolated tests.
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* docs: align page file names and nav labels with published URLs
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs: redirect moved dev pages and fix agent guide redirects
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* ci(docs): check that page URLs match their file paths
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* fix(docs): align navigation checks with repo conventions
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs: omit historical redirect aliases
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* fix(docs): preserve published overview and TypeScript URLs
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* fix(docs): redirect unversioned overview and TypeScript URLs
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
---------
Signed-off-by: Johnny Greco <jogreco@nvidia.com>