docs(website): sync v0.0.116 docs from d00f1fdf0c

This commit is contained in:
github-actions[bot]
2026-09-25 21:55:23 +00:00
parent 357560355d
commit b3896aff4d
79 changed files with 12213 additions and 0 deletions
+4
View File
@@ -11,3 +11,7 @@ snapshots:
source-ref: 496ebba293f5cc2bb2753444dddd534f0b4aeb6a
source-sha: 496ebba293f5cc2bb2753444dddd534f0b4aeb6a
version: 0.1.0
v0.0.116:
source-ref: d00f1fdf0c616a4a0debc9e5c0464ffd93175021
source-sha: d00f1fdf0c616a4a0debc9e5c0464ffd93175021
version: 0.0.116
+9
View File
@@ -38,6 +38,7 @@ experimental:
mdx-components:
- ./pages-latest/_components
- ./pages-dev/_components
- ./pages-v0.0.116/_components
- ./pages-v0.1.0/_components
- ./components
versions:
@@ -65,6 +66,14 @@ versions:
in OpenShell 0.1.0:</strong> a stable release cadence, new isolation primitives,
an expanded extension surface, and new APIs. <a href="https://docs.nvidia.com/openshell/dev/upgrade/0-1-0"
target="_blank" rel="noreferrer">Read the 0.1.0 upgrade guide</a>.</span>'
- display-name: v0.0.116
path: ./versions/v0.0.116.yml
slug: v0.0.116
announcement:
message: '<span style="display: block; padding: 0.375rem 0; text-align: left;"><strong>OpenShell
0.0.x has reached end of support.</strong> We recommend migrating to 0.1.0.
<a href="https://docs.nvidia.com/openshell/dev/upgrade/0-1-0" target="_blank"
rel="noreferrer">Read the 0.1.0 upgrade guide</a>.</span>'
redirects:
- source: /openshell/index.html
destination: /openshell/latest
@@ -0,0 +1,10 @@
{
"config": {
// MDX pages get their title from Fern frontmatter, not a top-level H1.
"MD041": false,
// MDX uses JSX components (<Callout>, <CodeBlock>, ...) that look like HTML.
"MD033": false,
// Published docs should label fenced code blocks for rendering and copy UX.
"MD040": true
}
}
+183
View File
@@ -0,0 +1,183 @@
---
# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
title: "Contributing to NVIDIA OpenShell Documentation"
description: ""
---
This guide covers how to write, edit, and review documentation for NVIDIA OpenShell. If you change code that affects user-facing behavior, update the relevant docs in the same PR.
The published docs live in `docs/`, navigation is defined in `docs/index.yml`, and `fern/` contains the site config, components, and theme assets.
## Use the Agent Skills
If you use an AI coding agent (Cursor, Claude Code, Codex, etc.), the repo includes skills that automate doc work. Use them before writing from scratch.
| Skill | What it does | When to use |
|---|---|---|
| `update-docs` | Scans recent commits for user-facing changes and drafts doc updates. | After landing features, before a release, or to find doc gaps. |
| `build-from-issue` | Plans and implements work from a GitHub issue, including doc updates. | When working from an issue that has doc impact. |
The skills live in `.agents/skills/` and follow the style guide below automatically. To use one, ask your agent to run it (e.g., "catch up the docs for everything merged since v0.2.0").
## When to Update Docs
Update documentation when your change:
- Adds, removes, or renames a CLI command or flag.
- Changes default behavior or configuration.
- Adds, removes, renames, or changes defaults for gateway TOML fields or driver-specific config options. Update `docs/reference/gateway-config.mdx` for these changes.
- Adds a new feature that users interact with.
- Fixes a bug that the docs describe incorrectly.
- Changes an API, protocol, or policy schema.
## Building Docs Locally
Use the local `mise` tasks for preview and validation, or run the Fern CLI directly from `fern/` if you already have it installed.
To preview Fern docs locally, run:
```shell
mise run docs:serve
```
To run non-interactive validation, run:
```shell
mise run docs
```
If you already have the Fern CLI installed, the equivalent commands from `fern/` are `fern docs dev` and `fern check`.
PRs that touch `docs/**` or `fern/**` are validated by `.github/workflows/branch-docs.yml`, and they also get a preview when `FERN_TOKEN` is available to the workflow.
## Writing Conventions
### Format
- Published docs use Fern MDX under `docs/`.
- Every page starts with YAML frontmatter. Use `title` and `description` on every page, then add page-level metadata like `sidebar-title`, `keywords`, and `position` when the page needs them.
- Include the SPDX license header as YAML comments inside frontmatter:
```text
---
# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
title: "Page Title"
---
```
- Do not repeat the page title as a body H1. Fern renders the title from frontmatter.
### Frontmatter Template
```yaml
---
# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
title: "Page Title"
sidebar-title: "Short Nav Title"
description: "One-sentence summary of the page."
keywords: "Generative AI, Cybersecurity, AI Agents, Sandboxing"
---
```
- `title` sets the page heading and browser title.
- `sidebar-title` sets the shorter label in the sidebar when the full page title is too long.
- `keywords` is a comma-separated string for page metadata.
- `position` controls ordering for pages discovered through a `folder:` entry.
- `slug` optionally overrides the page URL with a full path from the docs root.
For explicit entries in `docs/index.yml`, keep `page:`. Fern still requires it. If the page defines `sidebar-title`, set `page:` to that value. Otherwise set `page:` to the frontmatter `title`.
### Page Structure
1. Frontmatter `title` and `description`, plus any relevant page metadata.
2. A one- or two-sentence introduction stating what the page covers.
3. Sections organized by task or concept, using H2 and H3. Start each section with an introductory sentence that orients the reader.
4. A "Next Steps" section at the bottom linking to related pages when it helps the reader continue.
## Style Guide
Write like you are explaining something to a colleague. Be direct, specific, and concise.
### Voice and Tone
- Use active voice. "The CLI creates a gateway" not "A gateway is created by the CLI."
- Use second person ("you") when addressing the reader.
- Use present tense. "The command returns an error" not "The command will return an error."
- State facts. Do not hedge with "simply," "just," "easily," or "of course."
### Things to Avoid
These patterns are common in LLM-generated text and erode trust with technical readers. Remove them during review.
| Pattern | Problem | Fix |
|---|---|---|
| Unnecessary bold | "This is a **critical** step" on routine instructions. | Reserve bold for UI labels, parameter names, and genuine warnings. |
| Em dashes everywhere | "The gateway — which runs in Docker — creates sandboxes." | Use commas or split into two sentences. Em dashes are fine sparingly but should not appear multiple times per paragraph. |
| Superlatives | "OpenShell provides a powerful, robust, seamless experience." | Say what it does, not how great it is. |
| Hedge words | "Simply run the command" or "You can easily configure..." | Drop the adverb. "Run the command." |
| Emoji in prose | "🚀 Let's get started!" | No emoji in documentation prose. |
| Rhetorical questions | "Want to secure your agents? Look no further!" | State the purpose directly. |
### Formatting Rules
- End every sentence with a period.
- Use `code` formatting for CLI commands, file paths, flags, parameter names, and values.
- Use `shell` code blocks for copyable CLI examples. Do not prefix commands with `$`:
```shell
openshell gateway add http://127.0.0.1:18080 --local --name local
```
- Use `text` code blocks for transcripts, log output, and examples that should not be copied verbatim.
- Use tables for structured comparisons. Keep tables simple (no nested formatting).
- Use Fern components like `<Note>`, `<Tip>`, and `<Warning>` for callouts, not bold text.
- Use Fern components like `<Steps>` and `<Tabs>` when the page clearly benefits from them.
- Do not number section titles. Write "Deploy a Gateway" not "Section 1: Deploy a Gateway" or "Step 3: Verify."
- Do not use colons in titles. Write "Deploy and Manage Gateways" not "Gateways: Deploy and Manage."
- Use colons only to introduce a list. Do not use colons as general-purpose punctuation between clauses.
### Word List
Use these consistently:
| Use | Do not use |
|---|---|
| gateway | Gateway (unless starting a sentence) |
| sandbox | Sandbox (unless starting a sentence) |
| CLI | cli, Cli |
| API key | api key, API Key |
| NVIDIA | Nvidia, nvidia |
| OpenShell | Open Shell, openShell, Openshell, openshell |
| mTLS | MTLS, mtls |
| YAML | yaml, Yaml |
## Submitting Doc Changes
1. Create a branch following the project convention: `docs/<issue-id>-<description>/<username>`.
2. Make your changes.
3. Preview locally with `mise run docs:serve`.
4. Run `mise run docs`.
5. Run `mise run pre-commit` to catch formatting issues.
6. Open a PR with `docs:` as the conventional commit type.
```text
docs: update gateway deployment instructions
```
If your doc change accompanies a code change, include both in the same PR and use the code change's commit type:
```text
feat(cli): add gateway registration flag
```
## Reviewing Doc PRs
When reviewing documentation:
- Check that the style guide rules above are followed.
- Watch for LLM-generated patterns (excessive bold, em dashes, filler).
- Verify code examples are accurate and runnable.
- Confirm cross-references and links are not broken.
- Preview the page with `fern docs dev`, run `fern check`, and, if available, review the PR preview from `branch-docs.yml`.
@@ -0,0 +1,60 @@
/*
* SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
* SPDX-License-Identifier: Apache-2.0
*/
/**
* Badge links for GitHub, License, project status, Discord, etc.
* Uses a flex wrapper to display badges horizontally and hides Fern's
* external-link icon that otherwise stacks under each badge image.
* Requires the `.badge-links` CSS rule from main.css.
*/
declare const React: unknown;
export type BadgeItem = {
href: string;
src: string;
alt: string;
};
export function BadgeLinks({ badges = [] }: { badges?: BadgeItem[] }) {
if (badges.length === 0) {
return null;
}
return (
<div
className="badge-links"
style={{
alignItems: "center",
display: "flex",
flexWrap: "wrap",
gap: "8px",
lineHeight: 0,
margin: "0.25rem 0 0.75rem",
}}
>
{badges.map((b) => (
<a
key={b.href}
href={b.href}
target="_blank"
rel="noreferrer"
style={{
alignItems: "center",
display: "inline-flex",
width: "auto",
}}
>
<img
src={b.src}
alt={b.alt}
style={{
display: "block",
margin: 0,
}}
/>
</a>
))}
</div>
);
}
@@ -0,0 +1,130 @@
/*
* SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
* SPDX-License-Identifier: Apache-2.0
*/
declare const React: unknown;
const rotatingAgents = ["", "claude", "opencode", "codex"];
export function CommandTerminal({ command }: { command: string }) {
return (
<div
style={{
background: "#1a1a2e",
borderRadius: "8px",
boxShadow: "0 4px 16px rgb(0 0 0 / 25%)",
fontFamily:
'"SFMono-Regular", Menlo, Monaco, Consolas, "Liberation Mono", monospace',
fontSize: "0.875rem",
lineHeight: 1.8,
margin: "1.5rem 0",
overflow: "hidden",
}}
>
<style>{`
@keyframes nc-cycle {
0%,
20% {
opacity: 1;
}
25%,
100% {
opacity: 0;
}
}
@keyframes nc-blink {
50% {
opacity: 0;
}
}
`}</style>
<div
style={{
alignItems: "center",
background: "#252545",
display: "flex",
gap: "7px",
padding: "10px 14px",
}}
>
<span style={dotStyle("#ff5f56")} />
<span style={dotStyle("#ffbd2e")} />
<span style={dotStyle("#27c93f")} />
</div>
<div
style={{
color: "#d4d4d8",
display: "grid",
gridTemplateRows: "repeat(2, 1.8em)",
overflowX: "auto",
padding: "16px 20px",
}}
>
<div style={{ minWidth: "max-content", whiteSpace: "nowrap" }}>
<span style={{ color: "#76B900", userSelect: "none" }}>$ </span>
<span>{command}</span>
</div>
<div style={{ minWidth: "max-content", whiteSpace: "nowrap" }}>
<span style={{ color: "#76B900", userSelect: "none" }}>$ </span>
<span>{"openshell sandbox create "}</span>
<span
style={{
display: "inline-block",
height: "1.8em",
minWidth: "12ch",
overflow: "hidden",
position: "relative",
verticalAlign: "top",
}}
>
{rotatingAgents.map((agent, index) => (
<span
key={agent}
style={{
animation: "nc-cycle 12s ease-in-out infinite",
animationDelay: `${index * 3}s`,
inset: "0 auto auto 0",
opacity: 0,
position: "absolute",
whiteSpace: "nowrap",
}}
>
{agent !== "" && (
<span>
{"-- "}
<span style={{ color: "#76B900", fontWeight: 600 }}>
{agent}
</span>
<span
style={{
animation: "nc-blink 1s step-end infinite",
background: "#d4d4d8",
display: "inline-block",
height: "1.1em",
marginLeft: "1px",
verticalAlign: "text-bottom",
width: "2px",
}}
/>
</span>
)}
</span>
))}
</span>
</div>
</div>
</div>
);
}
function dotStyle(background: string) {
return {
background,
borderRadius: "50%",
display: "inline-block",
height: "12px",
width: "12px",
};
}
+4
View File
@@ -0,0 +1,4 @@
// SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
// SPDX-License-Identifier: Apache-2.0
declare const React: unknown;
@@ -0,0 +1,202 @@
---
# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
title: "Running the Gateway as a Container"
sidebar-title: "Container Gateway"
description: "Run the OpenShell gateway using docker run or docker-compose without the installer."
keywords: "Generative AI, Cybersecurity, AI Agents, Sandboxing, Docker, Podman, docker-compose, container, immutable OS, bootc, rpm-ostree"
position: 4
---
Use this approach when you want to run the OpenShell gateway as a container instead of installing it with the system package manager. This is useful on immutable OS distributions (Fedora CoreOS, bootc-based images, Silverblue) where the standard installer is not appropriate, or anywhere you prefer a container-first workflow.
The gateway image is published at `ghcr.io/nvidia/openshell/gateway`.
## Prerequisites for the Docker Driver
When the gateway runs as a container and creates Docker-backed sandboxes, the gateway container
communicates with the host Docker daemon via the mounted socket. This requires three things beyond
a basic `docker run`:
1. **Docker socket access.** The gateway process must be able to read and write the Docker socket.
Add the `docker` group (or the GID of `/var/run/docker.sock`) so the socket is accessible
without running as root.
2. **gRPC endpoint.** Sandbox containers call back to the gateway over the `OPENSHELL_GRPC_ENDPOINT`
address. The Docker driver substitutes `host.openshell.internal` as the host and the gateway's
own bind port as the port — only the **scheme** (`http` or `https`) is preserved. Use
`http://host.openshell.internal:8080` when TLS is disabled and `https://host.openshell.internal:8080`
when mTLS is enabled. The docker driver automatically binds the gateway to the bridge network
interface so sandbox containers can reach it — you do not need to expose the port on `0.0.0.0`.
3. **Supervisor binary on the host.** The gateway bind-mounts the `openshell-sandbox` supervisor
binary into each sandbox container. Because bind-mount paths are resolved by the host Docker
daemon (not inside the gateway container), the binary must exist at a path on the **host**
filesystem and be mounted at the **same absolute path** inside the gateway container. That way
the path the gateway records internally matches what Docker can find on the host when it
creates sandbox containers.
## Quick Start
Extract the supervisor binary to the host once, then start the gateway:
```shell
mkdir -p ~/openshell/supervisor
docker create --name tmp-supervisor ghcr.io/nvidia/openshell/supervisor:latest
docker cp tmp-supervisor:/openshell-sandbox ~/openshell/supervisor/openshell-sandbox
docker rm tmp-supervisor
chmod +x ~/openshell/supervisor/openshell-sandbox
```
Start the gateway:
```shell
docker run -d \
--name openshell-gateway \
--restart unless-stopped \
--group-add docker \
-p 127.0.0.1:8080:8080 \
-v openshell-state:/var/openshell \
-v /var/run/docker.sock:/var/run/docker.sock \
-v ~/openshell/supervisor/openshell-sandbox:~/openshell/supervisor/openshell-sandbox:ro \
-e OPENSHELL_DRIVERS=docker \
-e OPENSHELL_GRPC_ENDPOINT=http://host.openshell.internal:8080 \
-e OPENSHELL_DOCKER_SUPERVISOR_BIN=~/openshell/supervisor/openshell-sandbox \
-e OPENSHELL_DB_URL=sqlite:/var/openshell/openshell.db \
-e OPENSHELL_DISABLE_TLS=true \
ghcr.io/nvidia/openshell/gateway:latest
```
The volume mount uses `~/openshell/supervisor/openshell-sandbox` for both the host and container
paths. The shell expands `~` in both halves before passing the argument to Docker, so both sides
resolve to the same absolute path (e.g., `/home/user/openshell/supervisor/openshell-sandbox`).
This satisfies the same-path requirement so the host Docker daemon can find the binary when
creating sandbox containers.
Register the gateway with the CLI. If running on the same machine, use `--local`:
```shell
openshell gateway add http://127.0.0.1:8080 --local --name local
```
If registering from a different machine on the same network, use the host IP and `--remote`:
```shell
openshell gateway add http://HOST_IP:8080 --remote --name remote
```
Confirm the CLI can reach the gateway:
```shell
openshell status
```
<Warning>
Disabling TLS removes authentication. This example binds to `127.0.0.1` so only local
connections are accepted. To accept remote connections, enable mTLS or restrict access with
a firewall rule.
</Warning>
## Full mTLS Setup
To run the gateway with mutual TLS, generate the PKI bundle first, then start the gateway with the cert paths configured.
Bootstrap the PKI into a local state directory:
```shell
mkdir -p ~/.local/state/openshell/tls
docker run --rm \
-v "$HOME/.local/state/openshell:/home/openshell/.local/state/openshell" \
-v "$HOME/.config/openshell:/home/openshell/.config/openshell" \
ghcr.io/nvidia/openshell/gateway:latest \
generate-certs \
--output-dir /home/openshell/.local/state/openshell/tls \
--server-san host.openshell.internal
```
This writes the server and client certificates under `~/.local/state/openshell/tls/`, writes sandbox JWT signing keys under `~/.local/state/openshell/tls/jwt/`, and copies the client bundle to `~/.config/openshell/gateways/openshell/mtls/` so the CLI picks it up automatically.
Start the gateway with mTLS enabled:
```shell
docker run -d \
--name openshell-gateway \
--restart unless-stopped \
--group-add docker \
-p 127.0.0.1:8080:8080 \
-v "$HOME/.local/state/openshell:/home/openshell/.local/state/openshell" \
-v /var/run/docker.sock:/var/run/docker.sock \
-v ~/openshell/supervisor/openshell-sandbox:~/openshell/supervisor/openshell-sandbox:ro \
-e OPENSHELL_DRIVERS=docker \
-e OPENSHELL_GRPC_ENDPOINT=https://127.0.0.1:8080 \
-e OPENSHELL_DOCKER_SUPERVISOR_BIN=~/openshell/supervisor/openshell-sandbox \
-e OPENSHELL_DB_URL=sqlite:/home/openshell/.local/state/openshell/openshell.db \
-e OPENSHELL_LOCAL_TLS_DIR=/home/openshell/.local/state/openshell/tls \
-e OPENSHELL_TLS_CERT=/home/openshell/.local/state/openshell/tls/server/tls.crt \
-e OPENSHELL_TLS_KEY=/home/openshell/.local/state/openshell/tls/server/tls.key \
-e OPENSHELL_TLS_CLIENT_CA=/home/openshell/.local/state/openshell/tls/ca.crt \
-e OPENSHELL_ENABLE_MTLS_AUTH=true \
-e OPENSHELL_DOCKER_TLS_CA=/home/openshell/.local/state/openshell/tls/ca.crt \
-e OPENSHELL_DOCKER_TLS_CERT=/home/openshell/.local/state/openshell/tls/client/tls.crt \
-e OPENSHELL_DOCKER_TLS_KEY=/home/openshell/.local/state/openshell/tls/client/tls.key \
ghcr.io/nvidia/openshell/gateway:latest
```
Register the gateway with mTLS:
```shell
openshell gateway add https://127.0.0.1:8080 --local --name local
```
## Docker Compose
The [`deploy/docker/`](https://github.com/NVIDIA/OpenShell/tree/main/deploy/docker) directory in
the repository contains a production-ready Compose setup with full inline documentation:
| File | Purpose |
|---|---|
| `docker-compose.yml` | Gateway service, volumes, and environment variables |
| `gateway.toml` | TOML configuration mounted into the container |
Clone or copy those files, then start the gateway:
```shell
docker compose -f deploy/docker/docker-compose.yml up -d
```
Register the gateway with the CLI. If registering from the same machine:
```shell
openshell gateway add http://127.0.0.1:8080 --local --name local
```
If registering from a different machine on the same network, replace `HOST_IP` with the
machine's LAN address:
```shell
openshell gateway add http://HOST_IP:8080 --remote --name remote
```
## Using Podman
Replace `docker` with `podman` in the commands above. Mount the Podman socket instead of the Docker socket and set the driver to `podman`:
```shell
podman run -d \
--name openshell-gateway \
-p 127.0.0.1:8080:8080 \
-v openshell-state:/var/openshell \
-v "$XDG_RUNTIME_DIR/podman/podman.sock:/var/run/podman.sock" \
-e OPENSHELL_DRIVERS=podman \
-e OPENSHELL_PODMAN_SOCKET=/var/run/podman.sock \
-e OPENSHELL_DB_URL=sqlite:/var/openshell/openshell.db \
-e OPENSHELL_DISABLE_TLS=true \
ghcr.io/nvidia/openshell/gateway:latest
```
## Next Steps
- To create your first sandbox, refer to the [Quickstart](/get-started/quickstart).
- To control what the agent can access, refer to [Policies](/sandboxes/policies).
- For environment variable reference, refer to [Sandbox Compute Drivers](/reference/sandbox-compute-drivers).
+131
View File
@@ -0,0 +1,131 @@
---
# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
title: "How OpenShell Works"
sidebar-title: "How It Works"
description: "Understand the OpenShell architecture, runtime boundaries, gateways, sandboxes, and ecosystem integration points."
keywords: "Generative AI, Cybersecurity, AI Agents, Architecture, Gateway, Sandbox, Inference Routing"
position: 2
---
OpenShell is built around three stable runtime components: the **CLI**, the **Gateway**, and the **Supervisor**.
The CLI, SDK, and TUI provide user-facing access. The gateway is the
control plane: it owns API access, state, policy and settings delivery, provider and inference configuration, and relay coordination. The supervisor runs inside every sandbox workload and is the local security boundary. It launches the agent as a restricted child process and enforces policy where process identity, filesystem access, network egress, and
runtime credentials are visible.
Infrastructure-specific work sits behind integration boundaries. Compute,
credentials, control-plane identity, and sandbox identity each have a driver or
adapter boundary so OpenShell can integrate with native runtimes, secret stores,
identity providers, and workload identity systems without moving those concerns
into the core gateway or sandbox model.
```mermaid
flowchart TB
subgraph UI["User interfaces"]
CLI["CLI"]
SDK["SDK"]
TUI["TUI"]
end
subgraph CP["Control plane"]
GW["Gateway"]
DB[("Entity persistence")]
DRIVERS["Compute, credentials, and identity drivers"]
end
subgraph INFRA["Integrated infrastructure"]
RUNTIME["Docker, Podman, Kubernetes, or VM"]
SECRETSTORE["Secret stores"]
IDP["Identity providers"]
WORKLOADID["Workload identity"]
end
subgraph DP["Sandbox data plane"]
SUP["Supervisor"]
AGENT["Restricted agent process"]
PROXY["Policy proxy"]
POLICY["OPA policy engine"]
ROUTER["Inference router"]
end
CLI -->|"gRPC / HTTP"| GW
SDK -->|"gRPC / HTTP"| GW
TUI -->|"gRPC / HTTP"| GW
GW --> DB
GW --> DRIVERS
DRIVERS --> RUNTIME
DRIVERS --> SECRETSTORE
DRIVERS --> IDP
DRIVERS --> WORKLOADID
RUNTIME -->|"provisions workload"| SUP
SUP -->|"control, config, logs, relay"| GW
SUP -->|"spawn and restrict"| AGENT
AGENT -->|"ordinary egress"| PROXY
PROXY -->|"evaluate"| POLICY
PROXY -->|"allowed traffic"| EXT["External services"]
PROXY -->|"inference.local"| ROUTER
ROUTER -->|"managed inference"| MODEL["Inference backends"]
```
## Deployment Models
OpenShell can run on a single local machine or in a remote Kubernetes cluster.
The CLI workflow stays the same: users point the CLI, SDK, or TUI at a gateway,
and the gateway provisions sandboxes through its configured compute driver.
| Deployment | How it works | Best for |
|---|---|---|
| Local machine | The gateway runs on the user's workstation or a nearby development host and creates sandboxes with Docker, Podman, or a VM runtime. The supervisor inside each sandbox connects back to that local gateway. | Individual development, local agent experiments, and private workstation workflows. |
| Remote Kubernetes cluster | The gateway runs as a cluster service and creates sandbox pods in the configured namespace. Supervisors connect outbound to the gateway endpoint, so clients do not need direct pod access. | Shared teams, centrally managed policy, remote compute, GPUs, and production-like environments. |
This deployment split keeps the runtime model consistent. Local deployments use
the host's container or VM runtime as the integrated infrastructure. Kubernetes
deployments use the cluster scheduler, networking, secrets, identity, and GPU
device plugins without changing the gateway and sandbox contract.
## Core Components
| Component | Boundary |
|---|---|
| [Sandboxes](/sandboxes/manage-sandboxes) | Data-plane workloads that run the supervisor, launch restricted agent processes, apply local isolation, push logs, and maintain the gateway session. |
| [Gateways](/sandboxes/manage-gateways) | Authenticated control plane that owns API access, durable state, sandbox lifecycle, settings delivery, authorization, and relay coordination. |
| [Providers](/sandboxes/manage-providers) | Credential and provider records that map logical agent needs to platform or user-managed secrets without exposing raw credentials to the agent process. |
| [Policies](/sandboxes/policies) | Declarative controls for filesystem access, process identity, network egress, L7 rules, credential injection, and runtime policy updates. |
| [Inference Routing](/sandboxes/inference-routing) | Managed `https://inference.local` path that routes model traffic to configured backends while keeping provider credentials outside the sandbox. |
## Gateways and Sandboxes
The gateway and sandbox split control-plane authority from runtime enforcement. The gateway owns durable platform state: sandboxes, policy revisions, runtime settings, provider records, inference configuration, session records, and authorization decisions. A sandbox owns the local execution boundary: process identity, filesystem access, network egress, credential injection, local logs, and the agent child process.
The relationship is supervisor initiated. Each sandbox supervisor connects outbound to a known gateway endpoint, authenticates as a sandbox workload, and keeps a live session open for control traffic and relays. This avoids requiring every compute driver to solve gateway-to-sandbox reachability through pod IPs, bridge networks, port mappings, NAT traversal, or custom tunnels.
The gateway delivers desired state. The supervisor applies it locally, keeps last-known-good config when refresh fails, and leaves static isolation controls in place until the sandbox is recreated. Live operations such as config refresh, policy updates, credential delivery, log push, connect, exec, file sync, and relay setup use the same authenticated gateway-supervisor relationship.
## Supervisor Protection Layers
The supervisor is the sandbox-local enforcement component. It starts before the
agent process, prepares the sandbox runtime, fetches gateway configuration, and
then launches the agent under the active policy.
| Protection layer | Supervisor responsibility |
|---|---|
| Process | Drops privileges, applies process identity rules, disables privilege escalation paths, and starts the agent as a restricted child process. |
| Filesystem | Applies filesystem policy before the agent starts so undeclared paths are inaccessible and declared paths are read-only or read-write as configured. |
| Network | Routes ordinary egress through the policy proxy so destination, port, binary identity, and L7 request rules can be evaluated before traffic leaves the sandbox. |
| Credentials | Receives credential material from the gateway and injects it only through configured policy paths or request-time proxy rules. |
| Inference | Intercepts `https://inference.local` and forwards model traffic through the configured inference route instead of exposing provider credentials to the agent. |
| Observability | Emits local security and lifecycle logs, pushes sandbox logs to the gateway, and keeps relay endpoints available for connect, exec, and file transfer operations. |
Static controls such as filesystem and process isolation are established at
sandbox start and require sandbox recreation to change. Dynamic controls such as
network policy, credential delivery, and inference routing can refresh over the
live gateway-supervisor session.
## Ecosystem Integration
OpenShell integrates with infrastructure ecosystems instead of replacing them. Runtimes, schedulers, secret stores, identity providers, workload identity systems, image pipelines, storage, and GPU or device exposure remain owned by the platforms that provide them.
The gateway owns OpenShell control-plane semantics: sandbox state, lifecycle ordering, policy and settings resolution, credential mapping, authorization, inference configuration, and relay coordination. Drivers translate those semantics into platform-native operations.
The supervisor owns OpenShell sandbox semantics. Filesystem policy, process privilege reduction, network proxying, inference interception, credential injection, security logging, and gateway relay behavior stay consistent across Docker, Podman, Kubernetes, VM-backed sandboxes, and future integrations.
+157
View File
@@ -0,0 +1,157 @@
---
# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
title: "Installation"
sidebar-title: "Installation"
description: "Install OpenShell, choose a compute driver, and connect to a gateway."
keywords: "Generative AI, Cybersecurity, AI Agents, Sandboxing, Installation, Setup, Gateway, Docker, Podman, MicroVM, Kubernetes"
position: 3
---
## Install OpenShell
Install OpenShell with a single command:
```shell
curl -LsSf https://raw.githubusercontent.com/NVIDIA/OpenShell/main/install.sh | sh
```
The script detects your operating system and installs the OpenShell CLI and gateway with your native package manager. It then starts the local gateway server so you can begin creating sandboxes.
You can also download release artifacts directly from the [OpenShell GitHub Releases](https://github.com/NVIDIA/OpenShell/releases) page.
The `openshell` package on PyPI provides the Python SDK only. It does not install the `openshell` CLI. Add the SDK to a Python project with:
```shell
uv add openshell
```
Use `openshell status` to confirm the CLI can reach the gateway.
## Supported Compute Drivers
OpenShell supports several local compute drivers. Package-managed gateways leave the driver unset by default so the gateway can auto-detect an available driver. Set `compute_drivers` in the gateway TOML when you need to pin a specific driver.
| Compute Driver | How It Is Configured | System Requirements |
|---|---|---|
| Podman | The gateway is configured to create rootless Podman containers through the Podman API socket. | Linux with Podman 5.x, cgroups v2, rootless networking, and an active Podman user socket. |
| Docker | The gateway is configured to create containers through Docker Desktop or Docker Engine. | Docker Desktop or Docker Engine 28.0 or later on the gateway host. |
| MicroVM | The gateway is configured to create VM-backed sandboxes. | Host virtualization support. MicroVM uses Hypervisor.framework on macOS, KVM on Linux, and QEMU for GPU-backed sandboxes on Linux. |
For detailed driver behavior, refer to [Sandbox Compute Drivers](/reference/sandbox-compute-drivers). For gateway and sandbox operations, refer to [Gateways](/sandboxes/manage-gateways) and [Sandboxes](/sandboxes/manage-sandboxes).
## macOS
On macOS, the install script uses Homebrew. The Homebrew package installs the `openshell` CLI, the gateway binary, and a Homebrew-managed gateway service.
The Homebrew service uses the gateway's built-in `127.0.0.1:17670` listener and generates a local mTLS bundle on install. The installer registers `https://localhost:17670` with the CLI so TLS uses a DNS name covered by the generated certificate. The formula creates a Homebrew prefix config, such as `/opt/homebrew/var/openshell/gateway.toml`, without overriding `bind_address`. Docker Desktop and Podman Machine reuse the primary listener for sandbox callbacks when they can reach it. The gateway reads `~/.config/openshell/gateway.toml` instead when that file exists. Homebrew preserves user-edited prefix and user configs during upgrades; it removes the IPv6 bind only from an unchanged config generated by the affected formula.
The CLI reads the client bundle from `~/.config/openshell/gateways/openshell/mtls/`.
The installer starts the service for you. Use Homebrew service commands when you need to inspect, restart, or stop the gateway service:
```shell
brew services list
brew services restart openshell
```
## Linux
On Fedora and RHEL, the install script uses RPM packages. The RPM installs the `openshell` CLI, the `openshell-gateway` daemon, and a systemd user service.
On Debian and Ubuntu, the install script uses a Debian package. The Debian package installs the `openshell` CLI, the `openshell-gateway` daemon, VM sandbox support, and a systemd user service.
Linux packages require glibc 2.28 or newer. The installer checks libc before downloading packages and exits with an error on older glibc versions, Alpine, musl-based distributions, or unknown libc environments.
The Linux user service listens on `https://127.0.0.1:17670`, starts from built-in defaults, and generates a local mTLS bundle before the gateway starts. Create `~/.config/openshell/gateway.toml` only when you need to override those defaults.
The CLI reads the client bundle from `~/.config/openshell/gateways/openshell/mtls/`.
The installer starts the service for you. Use systemd user commands when you need to inspect, restart, or stop the gateway service:
```shell
systemctl --user status openshell-gateway
systemctl --user restart openshell-gateway
journalctl --user -u openshell-gateway -f
```
To keep the user service running after logout, enable linger:
```shell
sudo loginctl enable-linger $USER
```
## Snap
Install the OpenShell snap from the Snap Store:
```shell
sudo snap install openshell
```
The snap defines two apps: the `openshell` CLI and the `openshell.gateway`
systemd service. The gateway listens on `https://127.0.0.1:17670` and
stores its database at `$SNAP_COMMON/gateway.db` (typically
`/var/snap/openshell/common/gateway.db`). Create `$SNAP_COMMON/gateway.toml`
when you need to override gateway settings.
The snap CLI stores per-user config, data, and state under `$SNAP_USER_COMMON`,
typically `~/snap/openshell/common`. Gateway registrations live under
`$SNAP_USER_COMMON/.config/openshell/gateways/` instead of
`~/.config/openshell/gateways/`.
### Snap store installs
When installing from the Snap Store, snapd automatically connects the `home`,
`network`, and `network-bind` plugs. The `docker` plug still
requires manual connection:
```shell
sudo snap connect openshell:docker docker:docker-daemon
```
The snap declares `default-provider: docker` on the Docker plug so snapd will
offer to install the Docker snap, but the connection itself must be made
manually.
### Locally built snap packages
When installing a locally built `.snap` file, no plugs are connected by default:
```shell
sudo snap install ./openshell_*.snap --dangerous
sudo snap connect openshell:home
sudo snap connect openshell:network
sudo snap connect openshell:network-bind
sudo snap connect openshell:docker docker:docker-daemon
sudo snap connect openshell:log-observe
sudo snap connect openshell:system-observe
```
The `log-observe` and `system-observe` plugs are needed for the gateway service
to read logs and inspect system processes. The `docker` plug requires the
`docker:docker-daemon` slot from the Docker snap and does not work with
system-installed Docker.
### Gateway service
The gateway runs as a snap daemon with `refresh-mode: endure`, meaning snapd
will not restart it during snap refreshes. This prevents the gateway from
killing active sandbox sessions mid-refresh. Restart the service manually after
a snap refresh when you need the updated binary:
```shell
sudo systemctl restart snap.openshell.gateway
```
## Kubernetes
Kubernetes deployments use the OpenShell Helm chart. For step-by-step installation, refer to [Kubernetes Setup](/kubernetes/setup). For chart values and packaging details, refer to the [Helm chart README](https://github.com/NVIDIA/OpenShell/blob/main/deploy/helm/openshell/README.md).
## Next Steps
- To create your first sandbox, refer to the [Quickstart](/get-started/quickstart).
- To run the gateway as a container without the installer, refer to [Running the Gateway as a Container](/about/container-gateway).
- To register, select, and inspect gateways, refer to [Gateways](/sandboxes/manage-gateways).
- To supply API keys or tokens, refer to [Manage Providers](/sandboxes/manage-providers).
- To control what the agent can access, refer to [Policies](/sandboxes/policies).
+58
View File
@@ -0,0 +1,58 @@
---
# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
title: "Overview of NVIDIA OpenShell"
sidebar-title: "Overview"
description: "OpenShell is the safe, private runtime for autonomous AI agents. Run agents in sandboxed environments that protect your data, credentials, and infrastructure."
keywords: "Generative AI, Cybersecurity, AI Agents, Sandboxing, Security, Privacy, Inference Routing"
position: 1
---
NVIDIA OpenShell is an open-source runtime for executing autonomous AI agents in sandboxed environments with kernel-level isolation. It combines sandbox runtime controls and a declarative YAML policy so teams can run agents without giving them unrestricted access to local files, credentials, and external networks.
## Why OpenShell Exists
AI agents are most useful when they can read files, install packages, call APIs, and use credentials. That same access can create material risk. OpenShell is designed for this tradeoff: preserve agent capability while enforcing explicit controls over what the agent can access.
## Common Risks and Controls
The table below summarizes common failure modes and how OpenShell mitigates them.
| Threat | Without controls | With OpenShell |
|---|---|---|
| Data exfiltration | Agent uploads source code or internal files to unauthorized endpoints. | Network policies allow only approved destinations; other outbound traffic is denied. |
| Credential theft | Agent reads local secrets such as SSH keys or cloud credentials. | Filesystem restrictions (Landlock) confine access to declared paths only. |
| Unauthorized API usage | Agent sends prompts or data to unapproved model providers. | Privacy routing and network policies control where inference traffic can go. |
| Privilege escalation | Agent attempts `sudo`, setuid paths, or dangerous syscall behavior. | Unprivileged process identity and seccomp restrictions block escalation paths. |
## Protection Layers at a Glance
OpenShell applies defense in depth across the following policy domains.
| Layer | What it protects | When it applies |
|---|---|---|
| Filesystem | Prevents reads/writes outside allowed paths. | Locked at sandbox creation. |
| Network | Blocks unauthorized outbound connections. | Hot-reloadable at runtime. |
| Process | Blocks privilege escalation and dangerous syscalls. | Locked at sandbox creation. |
| Inference | Reroutes model API calls to controlled backends. | Hot-reloadable at runtime. |
For details, refer to [Customize Sandbox Policies](/sandboxes/policies) and [Default Policy](/reference/default-policy).
## Common Use Cases
OpenShell supports a range of agent deployment patterns.
| Use Case | Description |
|-----------------------------|----------------------------------------------------------------------------------------------------------|
| Secure coding agents | Run Claude Code, OpenCode, Codex, or GitHub Copilot CLI with constrained file and network access. |
| Private enterprise development | Route inference to self-hosted or private backends while keeping sensitive context under your control. |
| Compliance and audit | Treat policy YAML as version-controlled security controls that can be reviewed and audited. |
| Reusable environments | Use community sandbox images or bring your own containerized runtime. |
## Next Steps
Explore these topics to go deeper:
- To understand the runtime architecture, refer to [How OpenShell Works](/about/how-it-works).
- To install the CLI and create your first sandbox, refer to the [Quickstart](/get-started/quickstart).
- To learn how OpenShell enforces policy controls across protection layers, refer to [Customize Sandbox Policies](/sandboxes/policies).
@@ -0,0 +1,18 @@
---
# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
title: "NVIDIA OpenShell Release Notes"
sidebar-title: "Release Notes"
description: "Track the latest changes and improvements to NVIDIA OpenShell."
keywords: "Generative AI, Cybersecurity, Release Notes, Changelog, AI Agents"
position: 6
---
NVIDIA OpenShell follows a frequent release cadence. Use the following GitHub resources directly.
| Resource | Description |
|---|---|
| [Releases](https://github.com/NVIDIA/OpenShell/releases) | Versioned release notes and downloadable assets. |
| [Release comparison](https://github.com/NVIDIA/OpenShell/compare) | Diff between any two tags or branches. |
| [Merged pull requests](https://github.com/NVIDIA/OpenShell/pulls?q=is%3Apr+is%3Amerged) | Individual changes with review discussion. |
| [Commit history](https://github.com/NVIDIA/OpenShell/commits/main) | Full commit log on `main`. |
@@ -0,0 +1,24 @@
---
# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
title: "Supported Agents"
description: "AI agent frameworks and runtimes compatible with OpenShell sandboxes."
keywords: "Generative AI, Cybersecurity, AI Agents, Sandboxing, Claude, Codex, Cursor"
position: 5
---
The following table summarizes the agents that run in OpenShell sandboxes. Most agent sandbox images are maintained in the [OpenShell Community](https://github.com/NVIDIA/OpenShell-Community) repository. Agents in the base image are auto-configured when passed as the trailing command to `openshell sandbox create`.
| Agent | Source | Default Policy | Notes |
|---|---|---|---|
| [Claude Code](https://docs.anthropic.com/en/docs/claude-code) | [`base`](https://github.com/NVIDIA/OpenShell-Community/tree/main/sandboxes/base) | Full coverage | Works out of the box. Requires `ANTHROPIC_API_KEY` for direct Anthropic access, or use `inference.local` with a configured provider (e.g. Vertex AI). |
| [OpenCode](https://opencode.ai/) | [`base`](https://github.com/NVIDIA/OpenShell-Community/tree/main/sandboxes/base) | Partial coverage | Pre-installed. Use `ANTHROPIC_BASE_URL="https://inference.local/v1"` with a configured provider. Add `opencode.ai` endpoint and OpenCode binary paths to the policy for full functionality. |
| [Codex](https://developers.openai.com/codex) | [`base`](https://github.com/NVIDIA/OpenShell-Community/tree/main/sandboxes/base) | No coverage | Pre-installed. Requires a custom policy with OpenAI endpoints and Codex binary paths. Requires `OPENAI_API_KEY`. |
| [GitHub Copilot CLI](https://docs.github.com/en/copilot/github-copilot-in-the-cli) | [`base`](https://github.com/NVIDIA/OpenShell-Community/tree/main/sandboxes/base) | Full coverage | Pre-installed. Works out of the box. Requires `GITHUB_TOKEN` or `COPILOT_GITHUB_TOKEN`. |
| [OpenClaw](https://openclaw.ai/) | [NemoClaw](https://github.com/NVIDIA/NemoClaw) | Blueprint-managed | Run OpenClaw more securely inside NVIDIA OpenShell with managed inference using NemoClaw. |
| [Hermes Agent](https://github.com/NousResearch/hermes-agent) | [NemoClaw](https://github.com/NVIDIA/NemoClaw) | Blueprint-managed | Run Hermes Agent more securely inside NVIDIA OpenShell with managed inference using NemoClaw. |
| [Ollama](https://ollama.com/) | [`ollama`](https://github.com/NVIDIA/OpenShell-Community/tree/main/sandboxes/ollama) | Bundled | Run cloud and local models. Includes Claude Code, Codex, and OpenCode. Launch with `openshell sandbox create --from ollama`. |
| [Pi](https://pi.dev/) | [`pi`](https://github.com/NVIDIA/OpenShell-Community/tree/main/sandboxes/pi) | Bundled | Comes with Pi pre-installed. Launch with `openshell sandbox create --from pi`. |
For base image details and `--from` usage, refer to [Sandboxes](/sandboxes/manage-sandboxes#base-sandbox-container).
For a complete support matrix, refer to the [Support Matrix](/reference/support-matrix) page.
Binary file not shown.

After

Width:  |  Height:  |  Size: 351 B

Binary file not shown.

After

Width:  |  Height:  |  Size: 3.0 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 554 B

Binary file not shown.

After

Width:  |  Height:  |  Size: 796 B

Binary file not shown.

After

Width:  |  Height:  |  Size: 12 KiB

@@ -0,0 +1,7 @@
<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 100 100" width="100" height="100" role="img" aria-label="OpenShell">
<path d="M50 9 L84 23 V51 C84 72 69 86 50 92 C31 86 16 72 16 51 V23 Z" fill="#76B900"></path>
<g fill="none" stroke="#FFFFFF" stroke-width="6" stroke-linecap="round" stroke-linejoin="round">
<path d="M37 41 L49 50 L37 59"></path>
<path d="M55 59 H67"></path>
</g>
</svg>

After

Width:  |  Height:  |  Size: 406 B

Binary file not shown.

After

Width:  |  Height:  |  Size: 33 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 11 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 33 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 11 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 70 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 69 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 167 KiB

@@ -0,0 +1 @@
<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 560 164" width="560" height="164" role="img" aria-label="OpenShell"><g transform="translate(24 22) scale(1.2)"><path d="M50 9 L84 23 V51 C84 72 69 86 50 92 C31 86 16 72 16 51 V23 Z" fill="#76B900"></path><g fill="none" stroke="#FFFFFF" stroke-width="6" stroke-linecap="round" stroke-linejoin="round"><path d="M37 41 L49 50 L37 59"></path><path d="M55 59 H67"></path></g></g><path transform="translate(162 110.04) scale(0.08 -0.08)" d="M369 -12Q272 -12 203.5 30.5Q135 73 99 149.5Q63 226 63 328L63 372Q63 474 99 550.5Q135 627 203.5 670Q272 713 369 713L381 713Q478 713 546.5 670Q615 627 651 550.5Q687 474 687 372L687 328Q687 226 651 149.5Q615 73 546.5 30.5Q478 -12 381 -12L369 -12ZM369 61L381 61Q488 61 546 132Q604 203 604 326L604 375Q604 498 546 568.5Q488 639 381 639L369 639Q263 639 204.5 568.5Q146 498 146 375L146 326Q146 203 204.5 132Q263 61 369 61Z" fill="#FFFFFF"></path><path transform="translate(220.32 110.04) scale(0.08 -0.08)" d="M80 -200L80 509L153 509L153 422Q176 467 219.5 494Q263 521 325 521L332 521Q396 521 445 492.5Q494 464 522.5 408.5Q551 353 551 273L551 236Q551 156 522.5 100.5Q494 45 445 16.5Q396 -12 332 -12L325 -12Q268 -12 225.5 12Q183 36 160 77L160 -200L80 -200ZM310 57L319 57Q384 57 427.5 101Q471 145 471 237L471 273Q471 365 427.5 408.5Q384 452 319 452L310 452Q269 452 234.5 433Q200 414 180 374Q160 334 160 270L160 239Q160 144 204 100.5Q248 57 310 57Z" fill="#FFFFFF"></path><path transform="translate(266.96 110.04) scale(0.08 -0.08)" d="M279 -12Q174 -12 113 53Q52 118 52 236L52 273Q52 350 80 405.5Q108 461 159.5 491Q211 521 279 521L287 521Q354 521 402.5 493Q451 465 478 413.5Q505 362 505 293L505 232L131 232Q132 143 171.5 99Q211 55 279 55L294 55Q344 55 375.5 73Q407 91 426 125L487 88Q461 42 411.5 15Q362 -12 294 -12L279 -12ZM132 293L426 293L426 304Q426 378 388.5 416Q351 454 287 454L279 454Q217 454 176.5 413Q136 372 132 293Z" fill="#FFFFFF"></path><path transform="translate(309.52 110.04) scale(0.08 -0.08)" d="M80 0L80 509L153 509L153 418Q174 464 217 492.5Q260 521 324 521L333 521Q420 521 472.5 468Q525 415 525 313L525 0L445 0L445 300Q445 375 412 413.5Q379 452 317 452L310 452Q270 452 236 432.5Q202 413 181 374.5Q160 336 160 278L160 0L80 0Z" fill="#FFFFFF"></path><path transform="translate(355.76 110.04) scale(0.08 -0.08)" d="M307 -14Q88 -14 28 141L151 204Q186 111 307 111L320 111Q447 111 447 193L447 203Q447 235 425 258Q403 281 362 286L253 299Q153 311 101 361.5Q49 412 49 499L49 511Q49 573 81 618.5Q113 664 171.5 689Q230 714 309 714L324 714Q420 714 486 675.5Q552 637 581 564L453 510Q426 589 323 589L310 589Q256 589 225.5 568Q195 547 195 512L195 502Q195 441 279 430L388 417Q490 405 541.5 350Q593 295 593 207L593 195Q593 95 523 40.5Q453 -14 322 -14L307 -14Z" fill="#FFFFFF"></path><path transform="translate(403.92 110.04) scale(0.08 -0.08)" d="M65 0L65 730L205 730L205 469Q258 530 344 530L354 530Q442 530 493.5 475.5Q545 421 545 320L545 0L405 0L405 301Q405 353 383.5 381.5Q362 410 316 410L308 410Q283 410 259.5 397Q236 384 220.5 357Q205 330 205 286L205 0L65 0Z" fill="#FFFFFF"></path><path transform="translate(450.64 110.04) scale(0.08 -0.08)" d="M286 -14Q168 -14 103 53Q38 120 38 242L38 272Q38 394 104 462Q170 530 286 530L294 530Q404 530 465.5 463.5Q527 397 527 280L527 217L176 217Q183 98 286 98L300 98Q379 98 411 151L512 84Q484 38 428.5 12Q373 -14 301 -14L286 -14ZM177 311L391 311L391 314Q391 418 294 418L286 418Q190 418 177 311Z" fill="#FFFFFF"></path><path transform="translate(493.92 110.04) scale(0.08 -0.08)" d="M65 0L65 730L205 730L205 0L65 0Z" fill="#FFFFFF"></path><path transform="translate(513.92 110.04) scale(0.08 -0.08)" d="M65 0L65 730L205 730L205 0L65 0Z" fill="#FFFFFF"></path></svg>

After

Width:  |  Height:  |  Size: 3.6 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 166 KiB

@@ -0,0 +1 @@
<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 560 164" width="560" height="164" role="img" aria-label="OpenShell"><g transform="translate(24 22) scale(1.2)"><path d="M50 9 L84 23 V51 C84 72 69 86 50 92 C31 86 16 72 16 51 V23 Z" fill="#76B900"></path><g fill="none" stroke="#FFFFFF" stroke-width="6" stroke-linecap="round" stroke-linejoin="round"><path d="M37 41 L49 50 L37 59"></path><path d="M55 59 H67"></path></g></g><path transform="translate(162 110.04) scale(0.08 -0.08)" d="M369 -12Q272 -12 203.5 30.5Q135 73 99 149.5Q63 226 63 328L63 372Q63 474 99 550.5Q135 627 203.5 670Q272 713 369 713L381 713Q478 713 546.5 670Q615 627 651 550.5Q687 474 687 372L687 328Q687 226 651 149.5Q615 73 546.5 30.5Q478 -12 381 -12L369 -12ZM369 61L381 61Q488 61 546 132Q604 203 604 326L604 375Q604 498 546 568.5Q488 639 381 639L369 639Q263 639 204.5 568.5Q146 498 146 375L146 326Q146 203 204.5 132Q263 61 369 61Z" fill="#18181A"></path><path transform="translate(220.32 110.04) scale(0.08 -0.08)" d="M80 -200L80 509L153 509L153 422Q176 467 219.5 494Q263 521 325 521L332 521Q396 521 445 492.5Q494 464 522.5 408.5Q551 353 551 273L551 236Q551 156 522.5 100.5Q494 45 445 16.5Q396 -12 332 -12L325 -12Q268 -12 225.5 12Q183 36 160 77L160 -200L80 -200ZM310 57L319 57Q384 57 427.5 101Q471 145 471 237L471 273Q471 365 427.5 408.5Q384 452 319 452L310 452Q269 452 234.5 433Q200 414 180 374Q160 334 160 270L160 239Q160 144 204 100.5Q248 57 310 57Z" fill="#18181A"></path><path transform="translate(266.96 110.04) scale(0.08 -0.08)" d="M279 -12Q174 -12 113 53Q52 118 52 236L52 273Q52 350 80 405.5Q108 461 159.5 491Q211 521 279 521L287 521Q354 521 402.5 493Q451 465 478 413.5Q505 362 505 293L505 232L131 232Q132 143 171.5 99Q211 55 279 55L294 55Q344 55 375.5 73Q407 91 426 125L487 88Q461 42 411.5 15Q362 -12 294 -12L279 -12ZM132 293L426 293L426 304Q426 378 388.5 416Q351 454 287 454L279 454Q217 454 176.5 413Q136 372 132 293Z" fill="#18181A"></path><path transform="translate(309.52 110.04) scale(0.08 -0.08)" d="M80 0L80 509L153 509L153 418Q174 464 217 492.5Q260 521 324 521L333 521Q420 521 472.5 468Q525 415 525 313L525 0L445 0L445 300Q445 375 412 413.5Q379 452 317 452L310 452Q270 452 236 432.5Q202 413 181 374.5Q160 336 160 278L160 0L80 0Z" fill="#18181A"></path><path transform="translate(355.76 110.04) scale(0.08 -0.08)" d="M307 -14Q88 -14 28 141L151 204Q186 111 307 111L320 111Q447 111 447 193L447 203Q447 235 425 258Q403 281 362 286L253 299Q153 311 101 361.5Q49 412 49 499L49 511Q49 573 81 618.5Q113 664 171.5 689Q230 714 309 714L324 714Q420 714 486 675.5Q552 637 581 564L453 510Q426 589 323 589L310 589Q256 589 225.5 568Q195 547 195 512L195 502Q195 441 279 430L388 417Q490 405 541.5 350Q593 295 593 207L593 195Q593 95 523 40.5Q453 -14 322 -14L307 -14Z" fill="#18181A"></path><path transform="translate(403.92 110.04) scale(0.08 -0.08)" d="M65 0L65 730L205 730L205 469Q258 530 344 530L354 530Q442 530 493.5 475.5Q545 421 545 320L545 0L405 0L405 301Q405 353 383.5 381.5Q362 410 316 410L308 410Q283 410 259.5 397Q236 384 220.5 357Q205 330 205 286L205 0L65 0Z" fill="#18181A"></path><path transform="translate(450.64 110.04) scale(0.08 -0.08)" d="M286 -14Q168 -14 103 53Q38 120 38 242L38 272Q38 394 104 462Q170 530 286 530L294 530Q404 530 465.5 463.5Q527 397 527 280L527 217L176 217Q183 98 286 98L300 98Q379 98 411 151L512 84Q484 38 428.5 12Q373 -14 301 -14L286 -14ZM177 311L391 311L391 314Q391 418 294 418L286 418Q190 418 177 311Z" fill="#18181A"></path><path transform="translate(493.92 110.04) scale(0.08 -0.08)" d="M65 0L65 730L205 730L205 0L65 0Z" fill="#18181A"></path><path transform="translate(513.92 110.04) scale(0.08 -0.08)" d="M65 0L65 730L205 730L205 0L65 0Z" fill="#18181A"></path></svg>

After

Width:  |  Height:  |  Size: 3.6 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 127 KiB

@@ -0,0 +1 @@
<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 343 230" width="343" height="230" role="img" aria-label="OpenShell"><g transform="translate(106.5 22) scale(1.3)"><path d="M50 9 L84 23 V51 C84 72 69 86 50 92 C31 86 16 72 16 51 V23 Z" fill="#76B900"></path><g fill="none" stroke="#FFFFFF" stroke-width="6" stroke-linecap="round" stroke-linejoin="round"><path d="M37 41 L49 50 L37 59"></path><path d="M55 59 H67"></path></g></g><path transform="translate(22.09 215.63) scale(0.064 -0.064)" d="M369 -12Q272 -12 203.5 30.5Q135 73 99 149.5Q63 226 63 328L63 372Q63 474 99 550.5Q135 627 203.5 670Q272 713 369 713L381 713Q478 713 546.5 670Q615 627 651 550.5Q687 474 687 372L687 328Q687 226 651 149.5Q615 73 546.5 30.5Q478 -12 381 -12L369 -12ZM369 61L381 61Q488 61 546 132Q604 203 604 326L604 375Q604 498 546 568.5Q488 639 381 639L369 639Q263 639 204.5 568.5Q146 498 146 375L146 326Q146 203 204.5 132Q263 61 369 61Z" fill="#FFFFFF"></path><path transform="translate(68.75 215.63) scale(0.064 -0.064)" d="M80 -200L80 509L153 509L153 422Q176 467 219.5 494Q263 521 325 521L332 521Q396 521 445 492.5Q494 464 522.5 408.5Q551 353 551 273L551 236Q551 156 522.5 100.5Q494 45 445 16.5Q396 -12 332 -12L325 -12Q268 -12 225.5 12Q183 36 160 77L160 -200L80 -200ZM310 57L319 57Q384 57 427.5 101Q471 145 471 237L471 273Q471 365 427.5 408.5Q384 452 319 452L310 452Q269 452 234.5 433Q200 414 180 374Q160 334 160 270L160 239Q160 144 204 100.5Q248 57 310 57Z" fill="#FFFFFF"></path><path transform="translate(106.06 215.63) scale(0.064 -0.064)" d="M279 -12Q174 -12 113 53Q52 118 52 236L52 273Q52 350 80 405.5Q108 461 159.5 491Q211 521 279 521L287 521Q354 521 402.5 493Q451 465 478 413.5Q505 362 505 293L505 232L131 232Q132 143 171.5 99Q211 55 279 55L294 55Q344 55 375.5 73Q407 91 426 125L487 88Q461 42 411.5 15Q362 -12 294 -12L279 -12ZM132 293L426 293L426 304Q426 378 388.5 416Q351 454 287 454L279 454Q217 454 176.5 413Q136 372 132 293Z" fill="#FFFFFF"></path><path transform="translate(140.11 215.63) scale(0.064 -0.064)" d="M80 0L80 509L153 509L153 418Q174 464 217 492.5Q260 521 324 521L333 521Q420 521 472.5 468Q525 415 525 313L525 0L445 0L445 300Q445 375 412 413.5Q379 452 317 452L310 452Q270 452 236 432.5Q202 413 181 374.5Q160 336 160 278L160 0L80 0Z" fill="#FFFFFF"></path><path transform="translate(177.1 215.63) scale(0.064 -0.064)" d="M307 -14Q88 -14 28 141L151 204Q186 111 307 111L320 111Q447 111 447 193L447 203Q447 235 425 258Q403 281 362 286L253 299Q153 311 101 361.5Q49 412 49 499L49 511Q49 573 81 618.5Q113 664 171.5 689Q230 714 309 714L324 714Q420 714 486 675.5Q552 637 581 564L453 510Q426 589 323 589L310 589Q256 589 225.5 568Q195 547 195 512L195 502Q195 441 279 430L388 417Q490 405 541.5 350Q593 295 593 207L593 195Q593 95 523 40.5Q453 -14 322 -14L307 -14Z" fill="#FFFFFF"></path><path transform="translate(215.63 215.63) scale(0.064 -0.064)" d="M65 0L65 730L205 730L205 469Q258 530 344 530L354 530Q442 530 493.5 475.5Q545 421 545 320L545 0L405 0L405 301Q405 353 383.5 381.5Q362 410 316 410L308 410Q283 410 259.5 397Q236 384 220.5 357Q205 330 205 286L205 0L65 0Z" fill="#FFFFFF"></path><path transform="translate(253 215.63) scale(0.064 -0.064)" d="M286 -14Q168 -14 103 53Q38 120 38 242L38 272Q38 394 104 462Q170 530 286 530L294 530Q404 530 465.5 463.5Q527 397 527 280L527 217L176 217Q183 98 286 98L300 98Q379 98 411 151L512 84Q484 38 428.5 12Q373 -14 301 -14L286 -14ZM177 311L391 311L391 314Q391 418 294 418L286 418Q190 418 177 311Z" fill="#FFFFFF"></path><path transform="translate(287.63 215.63) scale(0.064 -0.064)" d="M65 0L65 730L205 730L205 0L65 0Z" fill="#FFFFFF"></path><path transform="translate(303.63 215.63) scale(0.064 -0.064)" d="M65 0L65 730L205 730L205 0L65 0Z" fill="#FFFFFF"></path></svg>

After

Width:  |  Height:  |  Size: 3.6 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 127 KiB

@@ -0,0 +1 @@
<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 343 230" width="343" height="230" role="img" aria-label="OpenShell"><g transform="translate(106.5 22) scale(1.3)"><path d="M50 9 L84 23 V51 C84 72 69 86 50 92 C31 86 16 72 16 51 V23 Z" fill="#76B900"></path><g fill="none" stroke="#FFFFFF" stroke-width="6" stroke-linecap="round" stroke-linejoin="round"><path d="M37 41 L49 50 L37 59"></path><path d="M55 59 H67"></path></g></g><path transform="translate(22.09 215.63) scale(0.064 -0.064)" d="M369 -12Q272 -12 203.5 30.5Q135 73 99 149.5Q63 226 63 328L63 372Q63 474 99 550.5Q135 627 203.5 670Q272 713 369 713L381 713Q478 713 546.5 670Q615 627 651 550.5Q687 474 687 372L687 328Q687 226 651 149.5Q615 73 546.5 30.5Q478 -12 381 -12L369 -12ZM369 61L381 61Q488 61 546 132Q604 203 604 326L604 375Q604 498 546 568.5Q488 639 381 639L369 639Q263 639 204.5 568.5Q146 498 146 375L146 326Q146 203 204.5 132Q263 61 369 61Z" fill="#18181A"></path><path transform="translate(68.75 215.63) scale(0.064 -0.064)" d="M80 -200L80 509L153 509L153 422Q176 467 219.5 494Q263 521 325 521L332 521Q396 521 445 492.5Q494 464 522.5 408.5Q551 353 551 273L551 236Q551 156 522.5 100.5Q494 45 445 16.5Q396 -12 332 -12L325 -12Q268 -12 225.5 12Q183 36 160 77L160 -200L80 -200ZM310 57L319 57Q384 57 427.5 101Q471 145 471 237L471 273Q471 365 427.5 408.5Q384 452 319 452L310 452Q269 452 234.5 433Q200 414 180 374Q160 334 160 270L160 239Q160 144 204 100.5Q248 57 310 57Z" fill="#18181A"></path><path transform="translate(106.06 215.63) scale(0.064 -0.064)" d="M279 -12Q174 -12 113 53Q52 118 52 236L52 273Q52 350 80 405.5Q108 461 159.5 491Q211 521 279 521L287 521Q354 521 402.5 493Q451 465 478 413.5Q505 362 505 293L505 232L131 232Q132 143 171.5 99Q211 55 279 55L294 55Q344 55 375.5 73Q407 91 426 125L487 88Q461 42 411.5 15Q362 -12 294 -12L279 -12ZM132 293L426 293L426 304Q426 378 388.5 416Q351 454 287 454L279 454Q217 454 176.5 413Q136 372 132 293Z" fill="#18181A"></path><path transform="translate(140.11 215.63) scale(0.064 -0.064)" d="M80 0L80 509L153 509L153 418Q174 464 217 492.5Q260 521 324 521L333 521Q420 521 472.5 468Q525 415 525 313L525 0L445 0L445 300Q445 375 412 413.5Q379 452 317 452L310 452Q270 452 236 432.5Q202 413 181 374.5Q160 336 160 278L160 0L80 0Z" fill="#18181A"></path><path transform="translate(177.1 215.63) scale(0.064 -0.064)" d="M307 -14Q88 -14 28 141L151 204Q186 111 307 111L320 111Q447 111 447 193L447 203Q447 235 425 258Q403 281 362 286L253 299Q153 311 101 361.5Q49 412 49 499L49 511Q49 573 81 618.5Q113 664 171.5 689Q230 714 309 714L324 714Q420 714 486 675.5Q552 637 581 564L453 510Q426 589 323 589L310 589Q256 589 225.5 568Q195 547 195 512L195 502Q195 441 279 430L388 417Q490 405 541.5 350Q593 295 593 207L593 195Q593 95 523 40.5Q453 -14 322 -14L307 -14Z" fill="#18181A"></path><path transform="translate(215.63 215.63) scale(0.064 -0.064)" d="M65 0L65 730L205 730L205 469Q258 530 344 530L354 530Q442 530 493.5 475.5Q545 421 545 320L545 0L405 0L405 301Q405 353 383.5 381.5Q362 410 316 410L308 410Q283 410 259.5 397Q236 384 220.5 357Q205 330 205 286L205 0L65 0Z" fill="#18181A"></path><path transform="translate(253 215.63) scale(0.064 -0.064)" d="M286 -14Q168 -14 103 53Q38 120 38 242L38 272Q38 394 104 462Q170 530 286 530L294 530Q404 530 465.5 463.5Q527 397 527 280L527 217L176 217Q183 98 286 98L300 98Q379 98 411 151L512 84Q484 38 428.5 12Q373 -14 301 -14L286 -14ZM177 311L391 311L391 314Q391 418 294 418L286 418Q190 418 177 311Z" fill="#18181A"></path><path transform="translate(287.63 215.63) scale(0.064 -0.064)" d="M65 0L65 730L205 730L205 0L65 0Z" fill="#18181A"></path><path transform="translate(303.63 215.63) scale(0.064 -0.064)" d="M65 0L65 730L205 730L205 0L65 0Z" fill="#18181A"></path></svg>

After

Width:  |  Height:  |  Size: 3.6 KiB

@@ -0,0 +1,12 @@
<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 100 100" width="100" height="100" role="img" aria-label="OpenShell">
<defs>
<mask id="prompt">
<rect width="100" height="100" fill="#fff"></rect>
<g fill="none" stroke="#000" stroke-width="6" stroke-linecap="round" stroke-linejoin="round">
<path d="M37 41 L49 50 L37 59"></path>
<path d="M55 59 H67"></path>
</g>
</mask>
</defs>
<path d="M50 9 L84 23 V51 C84 72 69 86 50 92 C31 86 16 72 16 51 V23 Z" fill="#18181A" mask="url(#prompt)"></path>
</svg>

After

Width:  |  Height:  |  Size: 550 B

@@ -0,0 +1,12 @@
<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 100 100" width="100" height="100" role="img" aria-label="OpenShell">
<defs>
<mask id="prompt">
<rect width="100" height="100" fill="#fff"></rect>
<g fill="none" stroke="#000" stroke-width="6" stroke-linecap="round" stroke-linejoin="round">
<path d="M37 41 L49 50 L37 59"></path>
<path d="M55 59 H67"></path>
</g>
</mask>
</defs>
<path d="M50 9 L84 23 V51 C84 72 69 86 50 92 C31 86 16 72 16 51 V23 Z" fill="#FFFFFF" mask="url(#prompt)"></path>
</svg>

After

Width:  |  Height:  |  Size: 550 B

@@ -0,0 +1,7 @@
<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 100 100" width="100" height="100" role="img" aria-label="OpenShell">
<path d="M50 9 L84 23 V51 C84 72 69 86 50 92 C31 86 16 72 16 51 V23 Z" fill="#FFFFFF"></path>
<g fill="none" stroke="#76B900" stroke-width="6" stroke-linecap="round" stroke-linejoin="round">
<path d="M37 41 L49 50 L37 59"></path>
<path d="M55 59 H67"></path>
</g>
</svg>

After

Width:  |  Height:  |  Size: 406 B

@@ -0,0 +1,7 @@
<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 100 100" width="100" height="100" role="img" aria-label="OpenShell">
<path d="M50 9 L84 23 V51 C84 72 69 86 50 92 C31 86 16 72 16 51 V23 Z" fill="#76B900"></path>
<g fill="none" stroke="#FFFFFF" stroke-width="6" stroke-linecap="round" stroke-linejoin="round">
<path d="M37 41 L49 50 L37 59"></path>
<path d="M55 59 H67"></path>
</g>
</svg>

After

Width:  |  Height:  |  Size: 406 B

@@ -0,0 +1,178 @@
---
# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
title: "Gateway Interceptors"
sidebar-title: "Gateway Interceptors"
description: "Extend OpenShell gateway operations with deployment-specific governance and business logic."
keywords: "Generative AI, Cybersecurity, AI Agents, Gateway Interceptors, Extensibility, Governance"
---
Gateway interceptors let operators add deployment-specific governance to OpenShell control-plane operations without modifying the gateway. An external gRPC service can modify or validate selected API writes before the gateway handles them, then observe successful responses after commit.
See the [governance interceptor example](https://github.com/NVIDIA/OpenShell/tree/main/examples/governance-interceptor) for a complete service that vends provider profiles, applies a signed policy to new sandboxes, and rejects attempts to weaken that policy.
## Choose Gateway Interceptors
Use a gateway interceptor when an external service needs to govern gateway API operations. For example, an interceptor can:
- Apply an approved policy to every new sandbox.
- Reject provider or policy changes that violate organizational rules.
- Enforce tenant quotas or naming conventions.
- Observe committed operations for an audit or inventory service.
- Vend an authoritative or composed provider profile catalog.
Gateway interceptors do not replace gateway persistence, authentication, authorization, policy safety checks, or driver validation. The gateway remains the system of record and validates an operation after interceptor modification.
## How Interception Works
The gateway runs interceptors after authentication and before dispatching a request to its handler:
`authenticate → decode and omit secrets → modify_operation → validate → gateway handler → post_commit`
| Phase | Input | Capabilities |
| ------------------ | ------------------------------------------- | ------------------------------------------------ |
| `modify_operation` | Proposed request | Allow, deny, or return RFC 6902 JSON patches. |
| `validate` | Modified request and optional current state | Allow or deny. |
| `post_commit` | Successful gateway response | Observe the response and attach log annotations. |
Only explicitly allowlisted unary mutation RPCs are interceptable. The gateway converts an operation to its protobuf JSON representation before evaluation. It applies patches atomically, validates the result against the RPC's protobuf schema, and converts the operation back to protobuf before the handler receives it.
The gateway runs built-in operation and driver validation after `modify_operation`. A patched operation cannot bypass gateway-owned invariants.
`post_commit` is observational. It cannot deny or modify an operation that the gateway has already committed.
## Implement an Interceptor Service
An interceptor implements the `openshell.gateway_interceptor.v1.GatewayInterceptor` gRPC service defined in [`proto/gateway_interceptor.proto`](https://github.com/NVIDIA/OpenShell/blob/main/proto/gateway_interceptor.proto):
- `Describe` declares the service's bindings and capabilities.
- `Evaluate` handles one selected operation phase.
- `SnapshotProviderProfiles` optionally returns a provider profile catalog.
Each `InterceptorEvaluation` identifies the configured interceptor, manifest binding, public OpenShell service and method, authenticated principal, and active phase. The phase determines whether the payload contains a proposed operation, optional current state, or committed response.
The interceptor returns an `InterceptorResult` with an allow or deny decision, an optional denial status and reason, JSON patches during `modify_operation`, and non-secret log annotations.
## Declare and Authorize Bindings
The interceptor declares bindings in its `InterceptorManifest` returned by `Describe`. Each binding selects one public OpenShell RPC and one or more phases. The gateway validates these declarations at startup against its compiled protobuf descriptors and explicit interceptable RPC allowlist.
The operator chooses how the manifest and gateway configuration combine with `binding_policy`:
| Policy | Behavior |
| ----------- | ------------------------------------------------------------------------------------------------------------------------------------------------- |
| `dynamic` | Enables valid manifest bindings. Gateway configuration may narrow or disable them. This is the compatibility default and emits a startup warning. |
| `allowlist` | Enables only operator-configured RPCs and phases. Extra manifest bindings are ignored and logged. |
| `exact` | Requires the configured RPCs and phases to match the manifest exactly. |
Use `allowlist` or `exact` when interceptor authority is part of a security boundary. These modes select bindings by public RPC rather than manifest binding ID, so renaming a binding does not change its authority.
## Register an Interceptor Service
Start the interceptor before the gateway, then register it in gateway TOML:
```toml
[[openshell.gateway.interceptors]]
name = "policy-governance"
grpc_endpoint = "https://governance.example:18081"
tls_ca_cert_path = "/etc/openshell/governance-ca.pem"
audience = "urn:example:governance"
order = 10
failure_policy = "fail_closed"
binding_policy = "allowlist"
timeout = "500ms"
max_response_bytes = 1048576
max_patches = 32
[[openshell.gateway.interceptors.bindings]]
rpc = "openshell.v1.OpenShell/CreateSandbox"
phases = ["modify_operation", "validate"]
[[openshell.gateway.interceptors.bindings]]
rpc = "openshell.v1.OpenShell/UpdateConfig"
phases = ["validate"]
```
The gateway supports `http://`, `https://`, and `unix://` interceptor endpoints. When gateway JWT signing is configured, authenticated network interceptors use `https://`; Unix sockets remain available for local integrations. HTTPS uses platform trust roots unless `tls_ca_cert_path` supplies a private CA, and normal hostname verification remains enabled. The gateway calls `Describe` and builds an immutable execution plan during startup. An unavailable service, invalid manifest, missing credential, or unauthorized configured binding prevents the gateway from starting.
The gateway attaches a short-lived EdDSA bearer token to `Describe`, `Evaluate`, and provider-profile snapshot calls. The token uses the configured `audience` (defaulting to `urn:openshell:extension:interceptor:<name>`) and `caller_kind: gateway`.
Return your expected audience in the `expected_audience` field of your `Describe` manifest. After authenticated `Describe` succeeds, the gateway compares the advertised value with its operator-configured audience and refuses to start when they differ. This is a post-authentication consistency assertion, not audience discovery: a strict verifier may reject an incorrect audience before returning the manifest, in which case startup reports an authentication failure. Leave the field empty to skip the consistency check.
Provision the trusted gateway URL, expected gateway ID, and public key or JWKS through the deployment. This operator-provisioned key material is the authoritative cold-start trust anchor. The expected issuer is exactly `openshell-gateway:<gateway_id>`; fetching JWKS does not establish that identity by itself. After initial trust is established, `GET /.well-known/openid-configuration` and its `jwks_uri` provide steady-state key refresh and operational convenience. The document is OIDC-shaped rather than OIDC-compliant because `issuer` is the gateway identity rather than the serving URL; compare `iss` against the configured value and fetch updates only over authenticated TLS at the trusted gateway URL. Pin `alg` to `EdDSA`, require `typ` to be exactly `openshell-ext+jwt`, and validate `kid`, signature, expected issuer, exact audience, positive expiry, and caller kind.
Set `allow_insecure_transport = true` on an interceptor to keep a plaintext `http://` endpoint working with no credential attached. The gateway logs a warning naming the interceptor at every startup, and the service cannot distinguish the gateway from any other client that can reach it.
Registration is static. Restart the gateway after adding, removing, or changing an interceptor. See [Gateway Configuration](/reference/gateway-config#gateway-interceptors) for the complete field reference.
## Select RPCs and Phases
Start with the invariant the interceptor must preserve, then identify every gateway RPC that can establish or weaken it. For example, an interceptor that owns sandbox policy can use `modify_operation` on `CreateSandbox` to apply an approved initial policy and `validate` on `UpdateConfig` to reject unauthorized changes.
The gateway maintains the canonical [interceptable route allowlist](https://github.com/NVIDIA/OpenShell/blob/main/crates/openshell-gateway-interceptors/src/routes.rs). Read the interceptor manifest and gateway startup diagnostics when selecting bindings. Unknown, streaming, read-only, and non-allowlisted RPCs cannot be intercepted.
Choose a phase by intent:
- Use `modify_operation` to apply defaults or controlled changes.
- Use `validate` to enforce a rule without changing the request.
- Use `post_commit` to notify or audit an external system after success.
## Mutate Operations Safely
Operations and committed responses use protobuf JSON field names and shapes. Only `modify_operation` accepts RFC 6902 JSON patches.
The gateway applies all patches returned by one binding as an atomic candidate. It then encodes that candidate as the RPC's protobuf request type and decodes it back to canonical protobuf JSON. If any patch or the resulting operation is invalid, the gateway discards the complete candidate and applies the binding's failure policy. Later bindings see only schema-valid operations that the handler can receive.
Fields marked secret in the protobuf schema are recursively omitted from interceptor requests and committed responses. An interceptor cannot patch an omitted field, use one as a patch source, or replace a containing object. Gateway handlers retain the complete operation and continue to receive secret fields that were omitted from the interceptor view.
## Configure Failure Behavior
Failure policy controls what happens when the gateway cannot obtain or apply a valid result. Failures include timeouts, transport errors, invalid responses, response-size violations, invalid phase behavior, and patch-limit violations.
| Policy | Behavior |
| ------------- | -------------------------------------------------------------------------------------------------- |
| `fail_closed` | Rejects the API operation before handler dispatch. |
| `fail_open` | Skips the failed result, continues with the previous valid operation, and emits warning telemetry. |
A valid deny result during `modify_operation` or `validate` always rejects the operation. It is not an interceptor failure and does not follow the failure policy.
Bindings that include `post_commit` must resolve to `fail_open`. The gateway rejects fail-closed post-commit configuration at startup because an observer cannot revoke an already committed response. A post-commit observation or evaluation failure is logged and counted without replacing the successful gateway response.
Use `fail_open` only when bypassing an unavailable or invalid interceptor preserves the intended governance boundary.
## Vend Provider Profiles
An interceptor can advertise `provider_profiles = true` in its manifest and implement `SnapshotProviderProfiles`. The RPC returns a `ProviderProfileSnapshot`. Add that interceptor to `provider_profile_sources` to include its profiles in the gateway's effective catalog.
Select only the interceptor to make its catalog authoritative:
```toml
[openshell.gateway]
provider_profile_sources = [
{ type = "interceptor", name = "provider-governance" },
]
```
Include `{ type = "builtin" }` or `{ type = "user" }` entries to compose interceptor profiles with the built-in or user-managed sources. Duplicate normalized profile IDs fail instead of overriding one source with another.
The gateway validates snapshot structure and provider profile semantics. It treats the configured interceptor as a trusted source and does not verify interceptor-defined signature, hash, or key annotations.
## Operate and Observe Interceptors
Plan interceptor deployment around these boundaries:
- Start every configured service before the gateway.
- Restart the gateway after changing registrations or bindings.
- Keep fail-closed services available whenever the gateway accepts writes.
- Treat the endpoint and its operator configuration as part of the gateway trust boundary.
- Do not include secrets in denial reasons or log annotations.
The gateway emits structured evaluation logs containing the interceptor name, binding ID, RPC, phase, decision, patch count, and interceptor-provided log annotations. Metrics record evaluation decisions, latency, patch application, fail-open and fail-closed outcomes, and post-commit observation failures.
## Current Limitations
- Only explicitly allowlisted unary write RPCs are interceptable. New gateway RPCs are non-interceptable until added to the allowlist.
- `current_state` is available only in the `validate` contract. The gateway does not yet populate it with method-specific state.
- Registration changes require a gateway restart.
- mTLS client authentication, service health checks, runtime registration, and overlapping signing-key rotation are not available.
- Extension tokens and sandbox-to-gateway tokens are signed by the same key, separated by audience and `typ`. The extension credential path cannot yet be rotated or revoked independently of sandbox admission.
- Interceptors cannot receive or mutate protobuf fields marked secret.
@@ -0,0 +1,234 @@
---
# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
title: "Supervisor Middleware"
sidebar-title: "Supervisor Middleware"
description: "Configure and operate built-in and operator-run middleware for sandbox HTTP requests and WebSocket messages."
keywords: "Generative AI, Cybersecurity, AI Agents, Supervisor Middleware, Extensibility, Request Filtering"
---
Supervisor middleware adds ordered processing stages to allowed HTTP and WebSocket egress. Middleware runs after network and L7 policy admit traffic and before OpenShell injects provider credentials. A stage can allow or deny an HTTP request or client WebSocket text message, replace its payload, add approved HTTP headers, and report audit-safe findings.
Middleware selection is independent of the network policy rule that admitted the request. OpenShell matches middleware by destination host, so the same middleware applies consistently across broad, specific, user-authored, and provider-derived network policies.
## Request Flow
For each inspected HTTP request, the supervisor:
1. Evaluates network and L7 policy.
2. Selects middleware whose host selectors match the admitted destination.
3. Buffers the request body using the largest body limit in the selected chain.
4. Runs matching middleware by ascending `order`. Policy validation rejects duplicate order values.
5. Re-checks body-aware protocol policy (GraphQL, JSON-RPC, MCP) after each stage that replaces the body. Every middleware receives a payload the policy admits, and a transformation cannot smuggle a denied or unparseable operation to a later stage or the upstream.
6. Applies allowed transformations, injects provider credentials, and forwards the request.
For an RFC 6455 upgrade over `ws://` or `wss://`, the supervisor first finds every host-matched attachment, then selects only implementations that advertise `WEBSOCKET_MESSAGE/PRE_CREDENTIALS`. It opens one ordered, phase-specific `EvaluateWebSocketSession` stream per selected stage. OpenShell sends `WebSocketSessionEvent` values, while the service returns `WebSocketSessionEventResult` values only for preflight and message events; session start and end are notifications. Future upstream-to-client inspection uses the same RPC with `PRE_RETURN`; an implementation that advertises both phases receives two independent streams for the WebSocket session. An attachment without the selected binding can still inspect the HTTP upgrade request when it advertises the HTTP binding, but it is not a failed WebSocket stage. OpenShell allows post-upgrade traffic and emits an informational `binding_not_selected` coverage event for that attachment.
1. A preflight before the upgrade is sent upstream. The stage chooses `INSPECT`, voluntary `SKIP`, or authoritative `DENY` and may return a bounded diagnostic reason, stable reason code, findings, and metadata. OpenShell runs selected preflights concurrently; any `DENY` rejects the upgrade regardless of `on_error`.
2. A session-start event after the upstream accepts the upgrade, including the negotiated subprotocol.
3. Complete client-to-upstream text messages in sequence order. OpenShell reassembles fragmented messages and decompresses negotiated `permessage-deflate` messages before evaluation.
4. A best-effort session-end event when the stage stream remains writable. OpenShell attempts at most one terminal event for each opened stream, including streams opened during a preflight that rejects the upgrade before session start.
The protobuf represents each logical message with a `text` or `binary` payload variant. Text uses the protobuf `string` type, so invalid UTF-8 cannot enter the middleware contract. Results use an optional matching replacement variant: absence preserves the input, while presence represents a replacement even when its content is empty. OpenShell rejects attempts to change the message type. Allowed replacements are re-framed, re-compressed when required, and forwarded. Binary messages, control frames, and upstream-to-client traffic remain uninspected. Binary messages pass through under both `on_error` modes. For each active selected stage, OpenShell emits an informational `unsupported_message_type` coverage event and advances the session-global sequence; the next text message can therefore reach the stage with a valid sequence gap.
The network supervisor reserves process-wide assembly capacity before buffering every parsed WebSocket text message, even when no middleware is selected. At most 32 assemblies run while 64 additional callers wait without buffering payload bytes. When both bounds are full, OpenShell closes the WebSocket with code `1013` before reading the new message payload. A text message may contain at most 4,096 fragments, must make input progress within 30 seconds, and must finish assembly within 2 minutes. Forwarding the completed text frame must finish within another 2 minutes. The assembly budget lasts for the supervisor process lifetime, so policy reloads do not reset its capacity.
Active middleware sessions additionally reserve shared middleware capacity before buffering WebSocket text, and HTTP middleware reserves the same capacity before buffering request bodies; at most 32 evaluations run and 64 additional unbuffered callers wait for capacity. When both middleware bounds are full, OpenShell sheds an HTTP request with `503 Service Unavailable` before reading its body. Persistent middleware streams use a separate process-wide budget of 32 sessions. WebSocket session admission does not wait: if the budget is full, OpenShell applies each selected config's `on_error` behavior before opening a stream.
Because each transformed body is re-checked before the next stage runs, a middleware hook always receives a request that satisfies the sandbox policy. A stage whose output the policy rejects stops the chain; under `enforcement: audit` the rejection is logged and the request proceeds.
If post-transformation policy evaluation itself fails, OpenShell denies the request and emits a high-severity detection finding. This failure is separate from middleware `on_error` because the middleware completed successfully; the sandbox policy could not validate its output.
Middleware receives the request before credential injection. Operator-run services cannot inspect OpenShell-managed credentials. Middleware-visible request headers are delivered in wire order and repeated header names are preserved as separate entries. OpenShell filters credential, routing, framing, and hop-by-hop headers before invoking middleware. It rejects malformed request headers and unsupported transfer-coding sequences before middleware or policy dispatch. Headers named by a request's `Connection` field are omitted from middleware input and removed before forwarding, except for the validated WebSocket upgrade pair.
The request context identifies the originating sandbox to operator-run services. It carries the sandbox ID (`sandbox_id`), the sandbox name (`sandbox_name`), and the workspace (`workspace`), letting audit and approval interfaces show a human-readable name and its workspace instead of an opaque ID. `sandbox_name` and `workspace` are for display and logging only: names are workspace-scoped and may be reused for different sandbox instances, so services must use `sandbox_id` for authorization, persistence, durable correlation, and identity. `sandbox_id` is always present on middleware requests. `sandbox_name` and `workspace` are best-effort: a supervisor that cannot resolve a value, or an older supervisor that predates a field, sends an empty string. Services should fall back to the sandbox ID when the name or workspace is empty.
## Choose a Middleware Type
| Type | Registration | Payload limit | Deployment |
| --- | --- | --- | --- |
| Built-in | None | Defined by OpenShell | Runs inside the supervisor |
| Operator-run service | Required in gateway TOML | Set by the operator, up to the service capability | Runs as a separate service reachable by the gateway and supervisors |
`openshell/regex` is an example built-in middleware. It replaces only simple, self-contained token patterns in UTF-8 HTTP bodies and client WebSocket text messages; the initial pattern recognizes `sk-` tokens. It does not infer values from keyword assignments such as JSON `password` fields. This best-effort text transformation is not parser-aware and does not guarantee that it will detect or fully remove sensitive values. Its `config` accepts one field, `mode: redact`, which is also the default when the field is omitted. Unknown config fields and non-string values are rejected at policy validation. Custom expressions are not configurable yet.
Operator-run services expose bindings for supported operation and phase pairs. A binding is identified by its operation and phase. V1 supports `HttpRequest/pre_credentials` and `WebSocketMessage/pre_credentials`; a service may expose either or both. Policies attach the complete middleware by its operator-owned gateway registration name.
## Register a Middleware Service
Start an operator-run service before starting the gateway, then add a registration to the local gateway TOML:
```toml
[[openshell.supervisor.middleware]]
name = "local-content-guard"
grpc_endpoint = "https://content-guard.example:50051"
tls_ca_cert_path = "/etc/openshell/content-guard-ca.pem"
audience = "urn:example:content-guard"
max_payload_bytes = 262144
timeout = "500ms"
```
| Field | Description |
| --- | --- |
| `name` | Operator-owned registration name used by policy attachments and diagnostics. Names must be unique, and `openshell/` is reserved for built-ins. |
| `grpc_endpoint` | Service address reachable from both the gateway and sandbox supervisors. Authenticated extensions use TLS `https://`. |
| `tls_ca_cert_path` | Optional PEM trust roots for a private HTTPS service. Custom roots replace platform roots and retain hostname verification. |
| `audience` | Exact audience expected by the service. Defaults to `urn:openshell:extension:middleware:<name>`. |
| `allow_insecure_transport` | Opt this registration out of extension authentication, permitting a plaintext `http://` endpoint with no bearer credential. Defaults to `false`. Development and trusted-network deployments only. |
| `max_payload_bytes` | Shared operator limit applied to inspectable logical payloads across every binding exposed by the service, up to the 4 MiB platform maximum. It caps HTTP bodies and complete WebSocket text messages. |
| `timeout` | Optional service-wide RPC timeout using an integer with an `ms` or `s` suffix. Defaults to `500ms`; valid values range from `10ms` through `30s`. |
Each binding returned by `Describe` may advertise a shorter `timeout` using the same syntax and bounds. The operator-configured service timeout is a ceiling: OpenShell uses the smaller of the binding and service values. An omitted binding timeout inherits the service setting, and an omitted service setting uses the 500 ms platform default. OpenShell rejects an invalid timeout before accepting the manifest. The operator-configured service timeout applies to `Describe` and `ValidateConfig`. The effective binding timeout applies only to `EvaluateHttpRequest`, WebSocket preflight, and each WebSocket message. WebSocket streams have no connection-wide deadline.
The gateway connects to every registered service and verifies its capabilities before accepting traffic. Gateway startup fails when a service is unavailable, reports an invalid capability, or exposes more than one binding for the same operation and phase. The manifest `name` is diagnostic metadata and does not need to match the operator registration name. Operator-run registration names cannot claim the reserved `openshell/` namespace.
Registration is static. Restart the gateway after adding, removing, or changing a service. See [Gateway Configuration](/reference/gateway-config#supervisor-middleware-services) for the complete gateway TOML context.
### Authenticate OpenShell Callers
When gateway JWT signing is configured, OpenShell attaches a short-lived EdDSA bearer token to every remote middleware RPC. Gateway calls use `caller_kind: gateway`; sandbox supervisor calls use `caller_kind: supervisor` and include the sandbox ID. Supervisors request credentials by registration name through `RefreshSandboxToken`. The gateway derives the audience from operator-owned configuration and authorizes each name against the sandbox's effective policy.
Return your expected audience in the `expected_audience` field of your `Describe` manifest. After authenticated `Describe` succeeds, OpenShell compares the advertised value with its operator-configured audience and refuses to start when they differ. This is a post-authentication consistency assertion, not audience discovery: a strict verifier may reject an incorrect audience before returning the manifest, in which case startup reports an authentication failure. Leave the field empty to skip the consistency check.
Provision the trusted gateway URL, expected gateway ID, and public key or JWKS through the deployment. This operator-provisioned key material is the authoritative cold-start trust anchor. The expected issuer is exactly `openshell-gateway:<gateway_id>`; fetching JWKS does not establish that identity by itself. After initial trust is established, `GET /.well-known/openid-configuration` and its `jwks_uri` provide steady-state key refresh and operational convenience. The document is OIDC-shaped rather than OIDC-compliant: `issuer` is the gateway identity, not the URL serving the document, so compare `iss` against the configured value and fetch updates only over authenticated TLS at the trusted gateway URL.
Cache keys by `kid`. Validate, at minimum:
- `typ` is exactly `openshell-ext+jwt`. Extension tokens and sandbox-to-gateway bootstrap tokens share a signing key and differ only in audience; this header is a second, independent discriminator.
- `alg` is pinned to `EdDSA`. Never select the algorithm from the token.
- Signature, expected issuer, exact audience, and positive expiry.
- `caller_kind`, and the sandbox identity when your service scopes behavior per sandbox.
A sandbox-to-gateway JWT is not an extension credential even though both token types use the same signing key.
Each token carries a unique `jti` that identifies that token instance for correlation and future explicit revocation. OpenShell reuses a token across calls until rotation and does not track `jti`, so rejecting a repeated `jti` would reject legitimate requests. Per-request replay resistance requires a request nonce or signature, channel binding, or another proof-of-possession mechanism.
### Run Without Extension Authentication
Set `allow_insecure_transport = true` on a registration to keep a plaintext `http://` endpoint working. OpenShell then attaches no credential to that service, supervisors do not request one, and the gateway refuses to mint one if asked. The gateway logs a warning naming the registration at every startup.
The service cannot distinguish OpenShell from any other client that can reach it. Use this only where the network already provides that guarantee, and prefer `https://` everywhere else.
## Apply Middleware with Policy
Add middleware configs to the top-level `network_middlewares` map. Each key is the policy-local config name:
```yaml
network_middlewares:
regex-redactor:
name: Redact API tokens
middleware: openshell/regex
order: 10
config:
mode: redact
on_error: fail_closed
endpoints:
include: ["*.example.com"]
exclude: ["trusted.example.com"]
```
Each config has a stable policy-local identity from its map key, an optional human-readable `name` that defaults to that key, a built-in or operator-owned registration name in `middleware`, an integer `order`, implementation-owned `config`, failure behavior, and host selectors. The optional name does not replace the map key for attachment or future keyed updates. A policy accepts at most 10 middleware configs.
`include` selects destination hosts. `exclude` takes precedence and removes hosts from that selection. Each config accepts at most 32 combined include and exclude patterns. Matching is case-insensitive and uses the same exact-host and DNS glob behavior as network policy endpoints: `*` matches exactly one DNS label, `**` matches one or more labels, and intra-label patterns like `*-api.example.com` work. Brace alternates such as `{prod,staging}` are rejected at validation; list each host pattern separately.
Matching configs run once each by ascending `order`; lower values run first. Order values must be unique across the complete policy, even when endpoint selectors do not overlap. The default order is `0`, so policies with multiple configs normally set explicit values. Different map keys may attach the same middleware and run as separate stages. Map keys are structurally unique. Runtime selection defensively rejects chains with more than 10 stages.
See [Policy Schema](/reference/policy-schema#network-middleware) for the complete field reference.
## Configure Failure Behavior
`on_error` controls what happens after an operation binding is selected and middleware is unavailable, rejects its configuration, returns an invalid result, or exceeds the selected binding's payload limit. It does not turn an unadvertised operation or an unsupported WebSocket message class into a middleware failure.
| Value | Behavior |
| --- | --- |
| `fail_closed` | Denies the HTTP request or closes the WebSocket when the stage fails. This is the default. |
| `fail_open` | Skips the failed HTTP stage. For a broken WebSocket stage stream, disables that stage for the rest of the connection and continues the remaining chain. |
Use `fail_open` only when bypassing the middleware preserves the intended security policy. OpenShell emits a detection finding when a failed stage is bypassed and a separate state-change finding when a WebSocket stage is disabled for the session.
Capability coverage is separate from failure handling. A host-matched HTTP-only attachment does not join the WebSocket chain, regardless of `on_error`. Binary messages are outside the V1 text-message binding and pass through even when a selected stage is `fail_closed`. OpenShell records both states as informational coverage events so operators do not mistake pass-through traffic for inspected traffic. If a deployment requires all WebSocket message classes to be inspected, V1 cannot express that requirement.
An explicit deny decision always stops the chain and denies the request or WebSocket upgrade, regardless of `on_error`. A WebSocket preflight `DENY` is a successful policy decision, not a middleware failure; OpenShell rejects the upgrade before upstream contact and ends each still-writable stream opened by a successful preflight decision with `MIDDLEWARE_DENIAL`. The HTTP response uses `error: middleware_denied`, identifies the policy-local middleware config, and omits policy-advisor remediation because the network and L7 allow rules already matched. OpenShell never copies the free-form middleware `reason` into the response or security logs. HTTP results, WebSocket preflight decisions, and WebSocket message results can instead return an optional stable `reason_code`: 1–64 bytes, starting with a lowercase ASCII letter and containing only lowercase ASCII letters, digits, and underscores. Invalid codes make the result a middleware failure governed by `on_error`. Preflight findings and metadata use the same bounds and audit-safe handling as message results.
```json
{
"error": "middleware_denied",
"detail": "Request rejected by configured middleware",
"policy": "api-policy",
"middleware": "prototype-content-guard",
"reason_code": "content_match"
}
```
A failed `fail_closed` stage uses `error: middleware_failed` and a platform-owned `detail`. It also omits `rule_missing`, `next_steps`, and `agent_guidance`: the failure did not result from a missing network or L7 policy rule, and changing policy cannot repair it. Runtime diagnostic text is available only through sanitized operator telemetry.
Middleware decisions are enforced regardless of the endpoint's `enforcement` mode. `enforcement: audit` applies to an endpoint's network and L7 policy rules and does not bypass middleware: a middleware deny, or a failed `fail_closed` stage, blocks the request even on an audit endpoint. A middleware service that needs to observe traffic without blocking should return an allow decision with findings, which OpenShell emits as detection findings.
## Set Payload Limits
Every middleware binding declares the largest logical payload or replacement it supports through `max_payload_bytes`. For `HTTP_REQUEST`, that payload is one request body. For `WEBSOCKET_MESSAGE`, it is one complete message rather than the whole session.
- Built-in middleware uses its OpenShell-defined limit.
- Each operator-run registration sets one `max_payload_bytes` ceiling no higher than any binding's advertised `max_payload_bytes` capability.
- A selected chain buffers using its largest stage limit, so every stage that can process the body receives it.
- The same per-stage limit applies to request bodies and replacement bodies.
The gateway rejects a registration whose operator limit exceeds the service capability or the 4 MiB platform maximum instead of silently clamping it. OpenShell also bounds the non-payload protobuf components: 64 KiB for service config, 4 KiB for request context, 32 KiB for the target, and 128 request header lines totaling at most 64 KiB encoded. Results allow a 4 KiB discarded free-form reason, a 64-byte validated reason code, 64 header mutations totaling at most 64 KiB encoded, 32 findings of at most 4 KiB encoded each, and 64 metadata entries totaling at most 32 KiB. Middleware gRPC servers should configure request and response message limits to at least 4 MiB plus 293 KiB so every platform-valid envelope fits.
At request time, exceeding a selected stage's limit is a middleware failure for that stage alone and follows that config's `on_error` behavior; other stages in the chain still run against their own limits. OpenShell can apply `fail_open` to an oversized `Content-Length` before consuming body bytes. A chunked body can cross the limit only after bytes have been consumed, so OpenShell denies that request because it cannot safely resume the original stream.
For a WebSocket binding, `max_payload_bytes` covers complete client text messages and replacements. Exceeding a selected stage's effective text-message limit follows that stage's `on_error`. The 4 MiB parsed-text platform cap and other protocol-safety limits are independent of middleware failure policy. Binary messages are not delivered to middleware, so the operator ceiling does not become a binary relay limit; individual raw binary frames retain the 16 MiB relay-safety bound. Oversized parsed text closes the connection with code `1009`; invalid UTF-8 uses `1007`; protocol errors use `1002`; middleware or policy denials use `1008`; and policy reload uses `1012`.
## Mutate Request Headers
A middleware result can return ordered header mutations before OpenShell injects credentials. A `write` mutation adds a value when the case-insensitive header name is absent and selects one behavior when it is already present:
- `append` adds another field value.
- `overwrite` removes every existing value before adding the new value.
- `skip` leaves existing values unchanged.
A `remove` mutation removes every value for a case-insensitive header name. OpenShell applies each successful stage's mutations before invoking the next middleware, so later stages observe the accumulated header state.
Header writes must use the `x-openshell-middleware-` prefix. Removes may target other middleware-visible request headers. Protected credential, routing, framing, and hop-by-hop headers are always rejected. Header values must not contain control characters.
OpenShell validates and applies each stage's mutations atomically. An invalid operation discards every mutation from that stage and follows its `on_error` behavior. Built-in failures can name the offending header. Operator-run failures use a platform-owned error code so request-derived header text cannot reach logs or denied responses.
## Operate Middleware Services
Plan startup and updates around these boundaries:
- Start registered services before the gateway. The gateway validates every registration during startup.
- Keep service endpoints reachable from both the gateway and sandbox supervisors. The supervisors call operator-run services directly on the request path.
- Restart the gateway after changing registrations.
- Keep required services available before creating or updating policies. The gateway validates implementation-owned config before persisting a policy.
- Treat `fail_open` as an explicit availability-over-enforcement decision.
When the effective sandbox configuration changes, a running supervisor validates the new service registry before installing it. If the reload fails, the supervisor keeps its last-known-good registry and emits a configuration failure event.
## Observe Middleware
Middleware activity is emitted through OpenShell's OCSF logging:
- Each invocation records its policy-local config name, attached middleware name, decision, transformation state, and failure state.
- A denied invocation records a platform-owned reason derived from the policy-local config name and optional validated reason code. OpenShell does not record service-provided free-form reason text.
- A bypass under `fail_open` emits a detection finding.
- A required stage that fails closed emits a high-severity detection finding.
- A host-matched attachment without a WebSocket binding emits an informational `binding_not_selected` coverage event.
- A binary message encountered by an active WebSocket stage emits an informational `unsupported_message_type` coverage event with message type, sequence, and byte count. It is not reported as an invocation or failure.
- Built-in findings include their type, label, and aggregate count. Operator-run findings use the operator-owned registration name and a platform label plus the aggregate count; OpenShell does not log service-provided finding text or diagnostic metadata. A stage can return at most 32 findings. Exceeding the per-stage cap is an invalid response handled through `on_error`. A maximum 10-stage chain retains and emits up to 320 findings without silently dropping findings from later stages.
- Registry reload success and failure are emitted as configuration state changes.
See [Logging](/observability/logging) for log access and [OCSF JSON Export](/observability/ocsf-json-export) for structured export.
## Current Limitations
- Middleware applies only through operation bindings advertised by each implementation. For protocols that have no supported middleware operation at all, such as HTTP/2 prior knowledge or non-HTTP TCP, the existing uninspectable-traffic gate denies a host match containing `fail_closed` and relays an all-`fail_open` match with a detection finding.
- The typed operation and phase pairs are `HTTP_REQUEST/PRE_CREDENTIALS` and `WEBSOCKET_MESSAGE/PRE_CREDENTIALS`.
- A host match does not imply every advertised operation: an HTTP-only attachment can inspect the upgrade GET, then post-upgrade traffic passes with `binding_not_selected` coverage.
- The V1 WebSocket binding inspects complete client text messages only. Binary messages pass with `unsupported_message_type` coverage for active stages; control frames and upstream-to-client messages remain outside the middleware operation.
- Selection uses destination host include and exclude patterns.
- A fail-closed middleware cannot cover `tls: skip` endpoints because OpenShell cannot inspect that traffic. An all-`fail_open` match may cover the endpoint; OpenShell bypasses the middleware and emits a detection finding.
- Operator-run services use TLS `https://` when gateway JWT signing is enabled, unless the registration sets `allow_insecure_transport`. Certificates must chain to the configured custom CA or platform roots, and the endpoint hostname must match.
- Extension tokens and sandbox-to-gateway tokens are signed by the same key. They are separated by audience and by `typ`, but the extension credential path cannot yet be rotated or revoked independently of sandbox admission.
- OpenShell does not track or revoke `jti`; bearer tokens can be replayed until expiry. Per-request replay resistance requires proof of possession or request binding.
- mTLS client authentication, health checks, runtime registration, and overlapping signing-key rotation are not available.
@@ -0,0 +1,101 @@
---
# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
title: "Quickstart"
description: "Install the OpenShell CLI, connect to a gateway, and create your first sandboxed AI agent."
keywords: "Generative AI, Cybersecurity, AI Agents, Sandboxing, Installation, Quickstart, Gateway, Docker, Kubernetes, Podman"
position: 1
---
This page gets you from a reachable OpenShell gateway to a running, policy-enforced sandbox.
## Prerequisites
Before you begin, make sure you have:
- A reachable OpenShell gateway.
- At least one compute driver configured for the gateway: Kubernetes, Docker, Podman, or MicroVM.
- The OpenShell CLI installed on your workstation.
For a complete list of requirements, refer to [Support Matrix](/reference/support-matrix).
If you have not chosen a compute driver yet, refer to [Installation](/about/installation).
## Install the OpenShell CLI
Run the install script:
```shell
curl -LsSf https://raw.githubusercontent.com/NVIDIA/OpenShell/main/install.sh | sh
```
The install script uses Homebrew, RPM, or a Debian package based on your machine. It starts the local gateway server after installation.
After installing the CLI, run `openshell --help` in your terminal to view the full CLI reference.
<Tip>
You can also clone the [NVIDIA OpenShell GitHub repository](https://github.com/NVIDIA/OpenShell) and use the `/openshell-cli` skill to load the CLI reference into your agent.
</Tip>
## Create Your First OpenShell Sandbox
Create a sandbox and launch an agent inside it.
Choose the tab that matches your agent:
<Tabs>
<Tab title="Claude Code">
Run the following command to create a sandbox with Claude Code:
```shell
openshell sandbox create -- claude
```
The CLI prompts you to create a provider from local credentials.
Type `yes` to continue.
If `ANTHROPIC_API_KEY` is set in your environment, the CLI picks it up automatically.
If not, you can configure it from inside the sandbox after it launches.
<Note>
`ANTHROPIC_API_KEY` is an API key from [console.anthropic.com](https://console.anthropic.com), not a subscription token. Subscription users must generate a separate API key.
</Note>
</Tab>
<Tab title="OpenCode">
Run the following command to create a sandbox with OpenCode:
```shell
openshell sandbox create -- opencode
```
The CLI prompts you to create a provider from local credentials.
Type `yes` to continue.
If `OPENAI_API_KEY` or `OPENROUTER_API_KEY` is set in your environment, the CLI picks it up automatically.
If not, you can configure it from inside the sandbox after it launches.
</Tab>
<Tab title="Codex">
Run the following command to create a sandbox with Codex:
```shell
openshell sandbox create -- codex
```
The CLI prompts you to create a provider from local credentials.
Type `yes` to continue.
If `OPENAI_API_KEY` is set in your environment, the CLI picks it up automatically.
If not, you can configure it from inside the sandbox after it launches.
</Tab>
<Tab title="Base Sandbox">
Use the `--from` flag to create a sandbox from the base container:
```shell
openshell sandbox create --from base
```
</Tab>
</Tabs>
@@ -0,0 +1,195 @@
---
# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
title: "Run the Gateway with Docker Compose"
sidebar-title: "Docker Compose Setup"
slug: "get-started/tutorials/docker-compose"
description: "Run the OpenShell gateway as a Docker Compose service and create agent sandboxes."
keywords: "Generative AI, Docker Compose, Gateway, Sandbox, OpenClaw, Docker, Installation"
---
This tutorial shows how to run the OpenShell gateway as a Docker Compose service on a Linux host or on a machine running Docker Desktop (Windows or macOS).
After completing this tutorial you have:
- An OpenShell gateway running as a Compose service.
- The `openshell` CLI registered against that gateway.
- An AI provider configured with your API key.
- A running agent sandbox.
## Prerequisites
- Docker Desktop (Windows or macOS) or Docker Engine with the Compose plugin (Linux).
- The `openshell` CLI installed on your workstation. See [Install the CLI](#install-the-cli) below.
- Port 8080 available on the host.
## Compose files
The Compose configuration lives at [`deploy/docker/`](https://github.com/NVIDIA/OpenShell/tree/main/deploy/docker) in the repository.
| File | Purpose |
|---|---|
| `docker-compose.yml` | Gateway service, volumes, and environment variables |
| `gateway.toml` | TOML reference for release builds with config-file support |
## Port note
The Docker compute driver injects `host.openshell.internal:<gateway-port>` into every sandbox container as its callback address. The gateway listens on port 8080 inside the container, so **port 8080 must be published at the same number on the Docker host**. Publishing it as a different host port (for example `18080:8080`) causes sandbox containers to call back to the wrong port and remain stuck in the `Provisioning` phase.
If port 8080 is taken, change `OPENSHELL_SERVER_PORT` and update the port mapping to `<your-port>:8080`, then set `OPENSHELL_PORT=<your-port>` in an `.env` file.
## Data directory
The gateway extracts the `openshell-sandbox` supervisor binary from `ghcr.io/nvidia/openshell/supervisor:latest` on first start and caches it at:
```text
/var/lib/openshell/openshell/docker-supervisor/<digest>/openshell-sandbox
```
This path is used as a bind-mount source when Docker creates sandbox containers.
Docker resolves bind-mount sources against the **host filesystem**, not the container filesystem, so the data directory must be bind-mounted at the **same absolute path** in both the host and the container.
The Compose file uses `/var/lib/openshell` for this purpose and sets `create_host_path: true` so Docker creates it on first run.
## Start the gateway
```shell
cd deploy/docker
docker compose up -d
```
Verify the gateway is healthy:
```shell
curl -sf http://localhost:8080/healthz
```
## Install the CLI
**Binary (recommended — macOS / Linux / WSL):**
```shell
curl -LsSf https://raw.githubusercontent.com/NVIDIA/OpenShell/main/install.sh | sh
```
<Note>
The `openshell` package on PyPI provides the Python SDK only and does not install the CLI.
On Windows without WSL, install the CLI inside a WSL 2 distribution (for example AlmaLinux or Ubuntu) and run all `openshell` commands from that distribution.
</Note>
## Register the gateway
Run this once after the gateway starts:
```shell
openshell gateway add http://localhost:8080 --name openshell-docker
```
Verify the connection:
```shell
openshell status
```
The output should show `Status: Connected`.
## Configure an AI provider
Set your API key as an environment variable and create a provider:
<Tabs>
<Tab title="Anthropic (Claude)">
```shell
ANTHROPIC_API_KEY=sk-ant-... \
openshell provider create --name anthropic --type anthropic --from-existing
```
</Tab>
<Tab title="OpenAI">
```shell
OPENAI_API_KEY=sk-... \
openshell provider create --name openai --type openai --from-existing
```
</Tab>
</Tabs>
Confirm the provider was stored:
```shell
openshell provider list
```
## Pre-pull sandbox images (optional)
Sandbox images are pulled automatically on first use, but the initial pull can take several minutes for large images. Pre-pull to avoid long waits at sandbox creation time:
```shell
# Base image — includes Claude Code, OpenCode, Codex, and Copilot
docker pull ghcr.io/nvidia/openshell-community/sandboxes/base:latest
```
## Create a sandbox
<Tabs>
<Tab title="OpenClaw">
OpenClaw runs inside OpenShell through [NemoClaw](https://github.com/NVIDIA/NemoClaw), which manages the sandbox image, inference routing, and security policies.
Follow the [NemoClaw Quickstart](https://docs.nvidia.com/nemoclaw/latest/get-started/quickstart/) to set up an OpenClaw sandbox with managed inference.
</Tab>
<Tab title="Claude Code">
```shell
openshell sandbox create -- claude
```
</Tab>
<Tab title="OpenCode">
```shell
openshell sandbox create -- opencode
```
</Tab>
</Tabs>
Wait for the phase to change from `Provisioning` to `Ready`:
```shell
openshell sandbox list
```
Then connect:
```shell
openshell sandbox connect <sandbox-name>
```
## Manage the gateway
| Command | Purpose |
|---|---|
| `docker compose up -d` | Start or restart the gateway |
| `docker compose down` | Stop the gateway and remove the container |
| `docker compose logs -f` | Tail gateway logs |
| `docker compose pull` | Pull a new gateway image version |
## Linux notes
On Linux, `host.docker.internal` and `host.openshell.internal` are not automatically resolvable from containers. Add the following under the `gateway` service in `docker-compose.yml`:
```yaml
extra_hosts:
- "host.docker.internal:host-gateway"
- "host.openshell.internal:host-gateway"
```
## Next steps
- [First Network Policy](/get-started/tutorials/first-network-policy) — apply L7 policies to your sandbox.
- [GitHub Push Access](/get-started/tutorials/github-sandbox) — grant a sandbox scoped GitHub access.
@@ -0,0 +1,218 @@
---
# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
title: "Write Your First Sandbox Network Policy"
sidebar-title: "First Network Policy"
slug: "get-started/tutorials/first-network-policy"
description: "Learn how OpenShell network policies work by creating a sandbox, observing default-deny in action, and applying a fine-grained L7 read-only rule."
keywords: "Generative AI, Cybersecurity, Tutorial, Policy, Network Policy, Sandbox, Security"
---
This tutorial shows how OpenShell's network policy system works in under five minutes. You create a sandbox, watch a request get blocked by the default-deny policy, apply a fine-grained L7 rule, and verify that reads are allowed while writes are blocked, all without restarting anything.
After completing this tutorial, you understand:
- How default-deny networking blocks all outbound traffic from a sandbox.
- How to apply a network policy that grants read-only access to a specific API.
- How L7 enforcement distinguishes between HTTP methods such as GET and POST on the same endpoint.
- How to inspect deny logs for a complete audit trail.
## Prerequisites
- A working OpenShell installation. Complete the [Quickstart](/get-started/quickstart) before proceeding.
- Docker Desktop running on your machine.
<Tip>
To run every step of this tutorial, you can also use the automated demo script at the [examples/sandbox-policy-quickstart](https://github.com/NVIDIA/OpenShell/blob/main/examples/sandbox-policy-quickstart) directory in the NVIDIA OpenShell repository. It runs the full walkthrough in under a minute but without any user interaction.
```shell
bash examples/sandbox-policy-quickstart/demo.sh
```
</Tip>
<Steps toc={true}>
## Create a Sandbox
Start by creating a sandbox with no network policies. This gives you a clean environment to observe default-deny behavior.
```shell
openshell sandbox create --name demo --no-auto-providers
```
`--no-auto-providers` skips the provider setup prompt since this tutorial uses `curl` instead of an AI agent.
You land in an interactive shell inside the sandbox:
```text
sandbox@demo:~$
```
## Try to Reach the GitHub API
With no network policy in place, every outbound connection is blocked. Test this by making a simple API call from inside the sandbox:
```shell
curl -s https://api.github.com/zen
```
`https://api.github.com/zen` is a lightweight, unauthenticated GitHub REST endpoint that returns a random aphorism on each call. It requires no tokens or parameters, which makes it a convenient smoke-test target for verifying outbound HTTPS connectivity.
The request fails. By default, all outbound network traffic is denied. The sandbox proxy intercepted the HTTPS CONNECT request to `api.github.com:443` and rejected it because no network policy authorizes `curl` to reach that host.
```text
curl: (56) Received HTTP code 403 from proxy after CONNECT
```
Exit the sandbox. Sandboxes are kept running by default, so you can reconnect later. Use `--no-keep` at creation time if you want the sandbox deleted after exit:
```shell
exit
```
## Check the Deny Log
Every denied connection produces a structured log entry. Query the sandbox logs from your host to confirm the denial and inspect the reason.
```shell
openshell logs demo --since 5m
```
You see a line like:
```text
action=deny dst_host=api.github.com dst_port=443 binary=/usr/bin/curl deny_reason="no matching network policy"
```
Every denied connection is logged with the destination, the binary that attempted it, and the reason. Nothing gets out silently.
## Apply a Read-Only GitHub API Policy
To allow the sandbox to reach the GitHub API, define a network policy that grants read-only access. The policy specifies which host, port, binary, and HTTP methods are permitted. Create a file called `github_readonly.yaml` with the following content:
```yaml
version: 1
filesystem_policy:
include_workdir: true
read_only: [/usr, /lib, /proc, /dev/urandom, /app, /etc, /var/log]
read_write: [/tmp, /dev/null]
landlock:
compatibility: best_effort
network_policies:
github_api:
name: github-api-readonly
endpoints:
- host: api.github.com
port: 443
protocol: rest
enforcement: enforce
access: read-only
binaries:
- { path: /usr/bin/curl }
```
The `filesystem_policy` and `landlock` sections preserve the default sandbox settings, while process identity is omitted so the active compute driver can select it. These sections are required because `policy set` replaces the entire policy. The `network_policies` section is the key part: `curl` can make GET, HEAD, and OPTIONS requests to `api.github.com` over HTTPS. Everything else is denied. The proxy auto-detects TLS on HTTPS endpoints and terminates it to inspect each HTTP request and enforce the `read-only` access preset at the method level.
Apply it:
```shell
openshell policy set demo --policy github_readonly.yaml --wait
```
`--wait` blocks until the sandbox confirms the new policy is loaded. No restart required. Policies are hot-reloaded.
<Tip>
This tutorial uses `curl` and `read-only` access to keep things simple. When building policies for real workloads:
- To scope the policy to an agent, replace the `binaries` section with your agent's binary, such as `/usr/local/bin/claude`, instead of `curl`.
- To grant write access, change `access: read-only` to `read-write` or add explicit `rules` for specific paths. Refer to the [Policy Schema](/reference/policy-schema).
- To allow additional endpoints, stack multiple policies in the same file for PyPI, npm, or your internal APIs. Refer to [Policies](/sandboxes/policies) for examples.
</Tip>
## Verify If GET Requests Are Allowed
The policy is now active. Reconnect to the sandbox and retry the same request to confirm that read access works.
```shell
openshell sandbox connect demo
```
Retry the same request:
```shell
curl -s https://api.github.com/zen
```
```text
Anything added dilutes everything else.
```
It works. The `read-only` preset allows GET requests through.
## Try a Write
The read-only preset allows GET but blocks mutating methods like POST, PUT, and DELETE. Test this by sending a POST request to the GitHub API while still inside the sandbox:
```shell
curl -s -X POST https://api.github.com/repos/octocat/hello-world/issues \
-H "Content-Type: application/json" \
-d '{"title":"oops"}'
```
```json
{"error":"policy_denied","policy":"github-api-readonly","detail":"POST /repos/octocat/hello-world/issues not permitted by policy"}
```
The CONNECT request succeeded because `api.github.com` is allowed, but the L7 proxy inspected the HTTP method and returned `403`. `POST` is not in the `read-only` preset. An agent with this policy can read code from GitHub but cannot create issues, push commits, or modify anything.
Exit the sandbox:
```shell
exit
```
## Check the L7 Deny Log
L7 denials are logged separately from connection-level denials. The log entry includes the exact HTTP method and path that the proxy rejected.
```shell
openshell logs demo --level warn --since 5m
```
```text
l7_decision=deny dst_host=api.github.com l7_action=POST l7_target=/repos/octocat/hello-world/issues l7_deny_reason="POST /repos/octocat/hello-world/issues not permitted by policy"
```
The log captures the exact HTTP method, path, and deny reason. In production, pipe these logs to your SIEM for a complete audit trail of every request your agent makes.
<Tip>
To log violations without blocking requests, set `enforcement: audit` instead of `enforcement: enforce` in the policy. This is useful for building a policy iteratively: deploy in audit mode, review the logs, and switch to enforce when the rules are correct.
</Tip>
## Clean Up
Delete the sandbox to free resources. This stops all processes and purges any injected credentials.
```shell
openshell sandbox delete demo
```
<Tip>
To run this entire walkthrough non-interactively, use the automated demo script:
```shell
bash examples/sandbox-policy-quickstart/demo.sh
```
</Tip>
</Steps>
## Next Steps
- To walk through a full policy iteration with Claude Code, including diagnosing denials and applying fixes from outside the sandbox, refer to [GitHub Sandbox](/get-started/tutorials/github-sandbox).
@@ -0,0 +1,358 @@
---
# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
title: "Grant GitHub Push Access to a Sandboxed Agent"
sidebar-title: "GitHub Push Access"
slug: "get-started/tutorials/github-sandbox"
description: "Learn the iterative policy workflow by launching a sandbox, diagnosing a GitHub access denial, and applying a custom policy to fix it."
keywords: "Generative AI, Cybersecurity, Tutorial, GitHub, Sandbox, Policy, Claude Code"
---
This tutorial walks through an iterative sandbox policy workflow. You launch a sandbox, ask Claude Code to push code to GitHub, and observe the default network policy denying the request.
You then diagnose the denial from your machine and from inside the sandbox, apply a policy update, and verify that the policy update to the sandbox takes effect.
After completing this tutorial, you have:
- A running sandbox with Claude Code that can push to a GitHub repository.
- A custom network policy that grants GitHub access for a specific repository.
- Experience with the policy iteration workflow: fail, diagnose, update, verify.
<Note>
This tutorial shows example prompts and responses from Claude Code. The exact wording you see might vary between sessions. Use the examples as a guide for the type of interaction, not as expected output.
</Note>
## Prerequisites
This tutorial requires the following:
- A working OpenShell installation. Complete the [Quickstart](/get-started/quickstart) before proceeding.
- A GitHub personal access token (PAT) with `repo` scope. Generate one from the [GitHub personal access token settings page](https://github.com/settings/tokens) by selecting **Generate new token (classic)** and enabling the `repo` scope.
- An [Anthropic account](https://console.anthropic.com/) with access to Claude Code. OpenShell provides the sandbox runtime, not the agent. You must authenticate with your own account.
- A GitHub repository you own to use as the push target. A scratch repository is sufficient. You can [create one](https://github.com/new) with a README if needed.
This tutorial uses two terminals to demonstrate the iterative policy workflow:
- **Terminal 1**: The sandbox terminal. You create the sandbox in this terminal by running `openshell sandbox create` and interact with Claude Code inside it.
- **Terminal 2**: A terminal outside the sandbox on your machine. You use this terminal for viewing the sandbox logs with `openshell term` and applying an updated policy with `openshell policy set`.
Each section below indicates which terminal to use.
<Steps toc={true}>
## Set Up a Sandbox with Your GitHub Token
Depending on whether you start a new sandbox or use an existing sandbox, choose the appropriate tab and follow the instructions.
<Tabs>
<Tab title="Starting a new sandbox">
In terminal 2, create a new sandbox with Claude Code. The [default policy](/reference/default-policy) is applied automatically, which allows read-only access to GitHub.
Create a [credential provider](/sandboxes/manage-providers) that injects your GitHub token into the sandbox automatically. The provider reads `GITHUB_TOKEN` from your host environment and sets it as an environment variable inside the sandbox:
```shell
GITHUB_TOKEN=<your-token>
openshell provider create --name my-github --type github --from-existing
openshell sandbox create --provider my-github -- claude
```
`openshell sandbox create` keeps the sandbox running after Claude Code exits, so you can apply policy updates later without recreating the environment. Add `--no-keep` if you want the sandbox deleted automatically instead.
Claude Code starts inside the sandbox. It prints an authentication link. Open it in your browser, sign in to your Anthropic account, and return to the terminal. When prompted, trust the `/sandbox` workspace to allow Claude Code to read and write files.
</Tab>
<Tab title="Using an existing sandbox">
In terminal 1, connect to a sandbox that is already running and set your GitHub token as an environment variable:
```shell
openshell sandbox connect <sandbox-name>
export GITHUB_TOKEN=<your-token>
```
To find the name of running sandboxes, run `openshell sandbox list` in terminal 2.
</Tab>
</Tabs>
## Push Code to GitHub
In terminal 1, ask Claude Code to write a simple script and push it to your repository. Replace `<org>` with your GitHub organization or username and `<repo>` with your repository name.
```md title="Prompt" wordWrap showLineNumbers={false}
Write a `hello_world.py` script and push it to `https://github.com/<org>/<repo>`.
```
Claude recognizes that it needs GitHub credentials. It asks how you want to authenticate. Provide your GitHub personal access token by pasting it into the conversation. Claude configures authentication and attempts the push.
The push fails. Claude reports an error, but the failure is not an authentication problem. The default sandbox policy permits read-only access to GitHub and blocks write operations, so the proxy denies the push before the request reaches the GitHub server.
## Diagnose the Denial
In this section, you diagnose the denial from your machine and from inside the sandbox.
### View the Logs from Your Machine
In terminal 2, launch the OpenShell terminal:
```shell
openshell term
```
The dashboard shows sandbox status and a live stream of policy decisions. Look for entries with `l7_decision=deny`. Select a deny entry to see the full detail:
```text
l7_action: PUT
l7_target: /repos/<org>/<repo>/contents/hello_world.py
l7_decision: deny
dst_host: api.github.com
dst_port: 443
l7_protocol: rest
policy: github_rest_api
l7_deny_reason: PUT /repos/<org>/<repo>/contents/hello_world.py not permitted by policy
```
The log shows that the sandbox proxy intercepted an outbound `PUT` request to `api.github.com` and denied it. The `github_rest_api` policy allows read operations (GET) but blocks write operations (PUT, POST, DELETE) to the GitHub API. A similar denial appears for `github.com` if Claude attempted a git push over HTTPS.
### Ask Claude Code to Check the Sandbox Logs
In terminal 1, ask Claude Code to check the sandbox logs for denied requests:
```md title="Prompt" wordWrap showLineNumbers={false}
Check the sandbox logs for any denied network requests. What is blocking the push?
```
Claude reads the deny entries and identifies the root cause. It explains that the failure is a sandbox network policy restriction, not a token permissions issue. For example, the following is a possible response:
<Accordion title="Response" defaultOpen={true}>
The sandbox runs a proxy that enforces policies on outbound traffic.
The `github_rest_api` policy allows GET requests (used to read the file)
but blocks PUT/write requests to GitHub. This is a sandbox-level restriction,
not a token issue. No matter what token you provide, pushes through the API
are blocked until you update the policy.
</Accordion>
Both perspectives confirm the same thing: the proxy is doing its job. The default policy is designed to be restrictive. To allow GitHub pushes, you need to update the network policy.
Copy the deny reason from Claude's response. You paste it into an agent running on your machine in the next step.
## Update the Policy from Your Machine
In terminal 2, paste the deny reason from the previous step into your coding agent on your machine, such as Claude Code or Cursor, and ask it to recommend a policy update. The deny reason gives the agent the context it needs to generate the correct policy rules. After pasting the following prompt sample, properly provide the GitHub organization and repository names of the repository you are pushing to.
```md title="Prompt" wordWrap showLineNumbers={false}
Based on the following deny reasons, recommend a sandbox policy update that allows GitHub pushes to `https://github.com/<org>/<repo>`, and save to `/tmp/sandbox-policy-update.yaml`:
The `filesystem_policy` and `landlock` sections are static. They are read once at sandbox creation and cannot be changed by a hot reload. They are included here for completeness so the file is self-contained. Process identity is omitted so the active compute driver can select it, and only the `network_policies` section takes effect when you apply this to a running sandbox.
```
The following steps outline the expected process done by the agent:
1. Inspects the deny reasons.
2. Writes an updated policy that adds `github_git` and `github_api` blocks that grant write access to your repository.
3. Saves the policy to `/tmp/sandbox-policy-update.yaml`.
## Review the Generated Policy
Refer to the following policy example to compare with the generated policy before applying it. Confirm that the policy grants only the access you expect. In this case, `git push` operations and GitHub REST API access scoped to a single repository.
<Accordion title="Full reference policy">
The following YAML shows a complete policy that extends the [default policy](/reference/default-policy) with GitHub access for a single repository. Replace `<org>` with your GitHub organization or username and `<repo>` with your repository name.
The `filesystem_policy` and `landlock` sections are static. OpenShell reads them at sandbox creation, and a hot reload cannot change them. They are included here for completeness so the file is self-contained. Process identity is omitted so the active compute driver can select it, and only the `network_policies` section takes effect when you apply this to a running sandbox.
```yaml
version: 1
# ── Static (locked at sandbox creation) ──────────────────────────
filesystem_policy:
include_workdir: true
read_only:
- /usr
- /lib
- /proc
- /dev/urandom
- /app
- /etc
- /var/log
read_write:
- /tmp
- /dev/null
landlock:
compatibility: best_effort
# ── Dynamic (hot-reloadable) ─────────────────────────────────────
network_policies:
# Claude Code ↔ Anthropic API
claude_code:
name: claude-code
endpoints:
- { host: api.anthropic.com, port: 443, protocol: rest, enforcement: enforce, access: full }
- { host: statsig.anthropic.com, port: 443 }
- { host: sentry.io, port: 443 }
- { host: raw.githubusercontent.com, port: 443 }
- { host: platform.claude.com, port: 443 }
binaries:
- { path: /usr/local/bin/claude }
- { path: /usr/bin/node }
# NVIDIA inference endpoint
nvidia_inference:
name: nvidia-inference
endpoints:
- { host: integrate.api.nvidia.com, port: 443 }
binaries:
- { path: /usr/bin/curl }
- { path: /bin/bash }
- { path: /usr/local/bin/opencode }
# ── GitHub: git operations (clone, fetch, push) ──────────────
github_git:
name: github-git
endpoints:
- host: github.com
port: 443
protocol: rest
enforcement: enforce
rules:
- allow:
method: GET
path: "/<org>/<repo>.git/info/refs*"
- allow:
method: POST
path: "/<org>/<repo>.git/git-upload-pack"
- allow:
method: POST
path: "/<org>/<repo>.git/git-receive-pack"
binaries:
- { path: /usr/bin/git }
# ── GitHub: REST API ─────────────────────────────────────────
github_api:
name: github-api
endpoints:
- host: api.github.com
port: 443
path: "/repos/<org>/<repo>/**"
protocol: rest
enforcement: enforce
rules:
# Full read-write access to the repository
- allow:
method: "*"
path: "/repos/<org>/<repo>/**"
- host: api.github.com
port: 443
path: "/graphql"
protocol: graphql
enforcement: enforce
rules:
# GitHub GraphQL API (used by gh CLI)
- allow:
operation_type: query
- allow:
operation_type: mutation
fields: [createIssue, updateIssue, addComment]
deny_rules:
- operation_type: mutation
fields: [deleteRepository, deleteRef, updateBranchProtectionRule]
binaries:
- { path: /usr/local/bin/claude }
- { path: /usr/local/bin/opencode }
- { path: /usr/bin/gh }
- { path: /usr/bin/curl }
# ── Package managers ─────────────────────────────────────────
pypi:
name: pypi
endpoints:
- { host: pypi.org, port: 443 }
- { host: files.pythonhosted.org, port: 443 }
- { host: github.com, port: 443 }
- { host: objects.githubusercontent.com, port: 443 }
- { host: api.github.com, port: 443 }
- { host: downloads.python.org, port: 443 }
binaries:
- { path: /sandbox/.venv/bin/python }
- { path: /sandbox/.venv/bin/python3 }
- { path: /sandbox/.venv/bin/pip }
- { path: "/sandbox/.uv/python/**/python*" }
- { path: /usr/local/bin/uv }
- { path: "/sandbox/.uv/python/**" }
# ── VS Code Remote ──────────────────────────────────────────
vscode:
name: vscode
endpoints:
- { host: update.code.visualstudio.com, port: 443 }
- { host: "*.vo.msecnd.net", port: 443 }
- { host: vscode.download.prss.microsoft.com, port: 443 }
- { host: marketplace.visualstudio.com, port: 443 }
- { host: "*.gallerycdn.vsassets.io", port: 443 }
binaries:
- { path: /usr/bin/curl }
- { path: /usr/bin/wget }
- { path: "/sandbox/.vscode-server/**" }
- { path: "/sandbox/.vscode-remote-containers/**" }
```
The following table summarizes the two GitHub-specific blocks:
| Block | Endpoint | Behavior |
|---|---|---|
| `github_git` | `github.com:443` | Git Smart HTTP protocol. The proxy auto-detects and terminates TLS to inspect requests. Permits `info/refs` (clone/fetch), `git-upload-pack` (fetch data), and `git-receive-pack` (push) for the specified repository. Denies all operations on unlisted repositories. |
| `github_api` | `api.github.com:443` | REST API. The proxy auto-detects and terminates TLS to inspect requests. Permits all HTTP methods for the specified repository and GraphQL queries. Denies API access to unlisted repositories. |
The remaining blocks (`claude_code`, `nvidia_inference`, `pypi`, `vscode`) are identical to the [default policy](/reference/default-policy). The default policy's `github_ssh_over_https` and `github_rest_api` blocks are replaced by the `github_git` and `github_api` blocks above, which grant write access to the specified repository. Sandbox behavior outside of GitHub operations is unchanged.
For details on policy block structure, refer to [Policies](/sandboxes/policies).
</Accordion>
## Apply the Policy
After you have reviewed the generated policy, apply it to the running sandbox:
```shell
openshell policy set <sandbox-name> --policy /tmp/sandbox-policy-update.yaml --wait
```
Network policies are hot-reloadable. The `--wait` flag blocks until the policy engine confirms the new revision loaded, and the update takes effect immediately without restarting the sandbox or reconnecting Claude Code.
## Retry the Push
In terminal 1, ask Claude Code to retry the push:
```md title="Prompt" wordWrap showLineNumbers={false}
The sandbox policy has been updated. Try pushing to the repository again.
```
The push completes successfully. The `openshell term` dashboard now shows `l7_decision=allow` entries for `api.github.com` and `github.com` where it previously showed denials.
## Clean Up
When you are finished, delete the sandbox to free gateway compute resources:
```shell
openshell sandbox delete <sandbox-name>
```
</Steps>
## Next Steps
The following resources cover related topics in greater depth:
- To add per-repository access levels (read-write vs read-only) or restrict to specific API methods, refer to the [Policy Schema Reference](/reference/policy-schema).
- To learn the full policy iteration workflow (pull, edit, push, verify), refer to [Policies](/sandboxes/policies).
- To inject credentials automatically instead of pasting tokens, refer to [Manage Providers](/sandboxes/manage-providers)
@@ -0,0 +1,45 @@
---
# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
title: "Tutorials"
slug: "get-started/tutorials"
description: "Step-by-step walkthroughs for OpenShell, from first sandbox to production-ready policies."
keywords: "Generative AI, Cybersecurity, Tutorial, Sandbox, Policy"
position: 1
---
Hands-on walkthroughs that teach OpenShell concepts by building real configurations. Each tutorial builds on the previous one, starting with core sandbox mechanics and progressing to production workflows.
<Cards>
<Card title="First Network Policy" href="/get-started/tutorials/first-network-policy">
Create a sandbox, observe default-deny networking, apply a read-only L7 policy, and inspect audit logs. No AI agent required.
</Card>
<Card title="GitHub Push Access" href="/get-started/tutorials/github-sandbox">
Launch Claude Code in a sandbox, diagnose a policy denial, and iterate on a custom GitHub policy from outside the sandbox.
</Card>
<Card title="Microsoft Graph Provider Refresh" href="/get-started/tutorials/microsoft-graph-provider-refresh">
Configure a Providers v2 Microsoft Graph provider with gateway-managed OAuth2 refresh-token rotation.
</Card>
<Card title="Inference with Ollama" href="/get-started/tutorials/inference-ollama">
Route inference through Ollama using cloud-hosted or local models, and verify it from a sandbox.
</Card>
<Card title="Local Inference with LM Studio" href="/get-started/tutorials/local-inference-lmstudio">
Route inference to a local LM Studio server using the OpenAI-compatible or Anthropic-compatible APIs.
</Card>
<Card title="Docker Compose Setup" href="/get-started/tutorials/docker-compose">
Run the OpenShell gateway as a Docker Compose service and create agent sandboxes including OpenClaw.
</Card>
</Cards>
@@ -0,0 +1,211 @@
---
# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
title: "Run Local Inference with Ollama"
sidebar-title: "Inference with Ollama"
slug: "get-started/tutorials/inference-ollama"
description: "Run local and cloud models inside an OpenShell sandbox using the Ollama community sandbox, or route sandbox requests to a host-level Ollama server."
keywords: "Generative AI, Cybersecurity, Tutorial, Inference Routing, Ollama, Local Inference, Sandbox"
---
This tutorial covers two ways of running Ollama with OpenShell:
1. Ollama sandbox. This is the recommended way to run Ollama. A self-contained sandbox with Ollama, Claude Code, and Codex pre-installed. One command starts it.
2. Host-level Ollama. This is an alternative way to run Ollama. Run Ollama on the gateway host and route sandbox inference to it. Use this option when you want a single Ollama instance shared across multiple sandboxes.
After completing this tutorial, you know how to:
- Launch the Ollama community sandbox for a batteries-included experience.
- Use `ollama launch` to start coding agents inside a sandbox.
- Expose a host-level Ollama server to sandboxes through `inference.local`.
## Prerequisites
- A working OpenShell installation. Complete the [Quickstart](/get-started/quickstart) before proceeding.
## Option A: Ollama Community Sandbox (Recommended)
The Ollama community sandbox bundles Ollama, Claude Code, OpenCode, and Codex into a single image. Ollama starts automatically when the sandbox launches.
<Steps toc={true}>
### Create the Sandbox
```shell
openshell sandbox create --from ollama
```
This pulls the community sandbox image, applies the bundled policy, and drops you into a shell with Ollama running.
### Chat with a Model
Chat with a local model
```shell
ollama run qwen3.5
```
Or a cloud model
```shell
ollama run kimi-k2.5:cloud
```
Or use `ollama launch` to start a coding agent with Ollama as the model backend:
```shell
ollama launch claude
ollama launch codex
ollama launch opencode
```
For CI/CD and automated workflows, `ollama launch` supports a headless mode:
```shell
ollama launch claude --yes --model qwen3.5
```
</Steps>
### Model Recommendations
| Use case | Model | Notes |
|---|---|---|
| Smoke test | `qwen3.5:0.8b` | Fast, lightweight, good for verifying setup |
| Coding and reasoning | `qwen3.5` | Strong tool calling support for agentic workflows |
| Complex tasks | `nemotron-3-super` | 122B parameter model, needs 48GB+ VRAM |
| No local GPU | `qwen3.5:cloud` | Runs on Ollama's cloud infrastructure, no `ollama pull` required |
<Note>
Cloud models use the `:cloud` tag suffix and do not require local hardware.
```shell
openshell sandbox create --from ollama
```
</Note>
### Tool Calling
Agentic workflows (Claude Code, Codex, OpenCode) rely on tool calling. The following models have reliable tool calling support: Qwen 3.5, Nemotron-3-Super, GLM-5, and Kimi-K2.5. Check the [Ollama model library](https://ollama.com/library) for the latest models.
### Updating Ollama
To update Ollama inside a running sandbox:
```shell
update-ollama
```
Or auto-update on every sandbox start:
```shell
openshell sandbox create --from ollama -e OLLAMA_UPDATE=1
```
## Option B: Host-Level Ollama
Use this approach when you want a single Ollama instance on the gateway host, shared across multiple sandboxes through `inference.local`.
<Note>
This approach uses Ollama because it is easy to install and run locally, but you can substitute other inference engines such as vLLM, SGLang, TRT-LLM, and NVIDIA NIM by changing the startup command, base URL, and model name.
</Note>
<Steps toc={true}>
### Install and Start Ollama
Install [Ollama](https://ollama.com/) on the gateway host:
```shell
curl -fsSL https://ollama.com/install.sh | sh
```
Start Ollama on all interfaces so it is reachable from sandboxes:
```shell
OLLAMA_HOST=0.0.0.0:11434 ollama serve
```
<Tip>
If you see `Error: listen tcp 0.0.0.0:11434: bind: address already in use`, Ollama is already running as a system service. Stop it first:
```shell
systemctl stop ollama
OLLAMA_HOST=0.0.0.0:11434 ollama serve
```
</Tip>
### Pull a Model
In a second terminal, pull a model:
```shell
ollama run qwen3.5:0.8b
```
Type `/bye` to exit the interactive session. The model stays loaded.
### Create a Provider
Create an OpenAI-compatible provider pointing at the host Ollama:
```shell
openshell provider create \
--name ollama \
--type openai \
--credential OPENAI_API_KEY=empty \
--config OPENAI_BASE_URL=http://host.openshell.internal:11434/v1
```
OpenShell injects `host.openshell.internal` so sandboxes and the gateway can reach the host machine. You can also use the host's LAN IP.
### Set Inference Routing
```shell
openshell inference set --provider ollama --model qwen3.5:0.8b
```
Confirm:
```shell
openshell inference get
```
### Verify from a Sandbox
```shell
openshell sandbox create -- \
curl https://inference.local/v1/chat/completions \
--json '{"messages":[{"role":"user","content":"hello"}],"max_tokens":10}'
```
The response should be JSON from the model.
</Steps>
## Troubleshooting
Common issues and fixes:
- **Ollama not reachable from sandbox:** Ollama must be bound to `0.0.0.0`, not `127.0.0.1`. This applies to host-level Ollama only; the community sandbox handles this automatically.
- **`OPENAI_BASE_URL` wrong:** Use `http://host.openshell.internal:11434/v1`, not `localhost` or `127.0.0.1`.
- **Model not found:** Run `ollama ps` to confirm the model is loaded. Run `ollama pull <model>` if needed.
- **HTTPS instead of HTTP:** Code inside sandboxes must call `https://inference.local`, not `http://`.
- **AMD GPU driver issues:** Ollama v0.18+ requires ROCm 7 drivers for AMD GPUs. Update your drivers if you see GPU detection failures.
Useful commands:
```shell
openshell status
openshell inference get
openshell provider get ollama
```
## Next Steps
- To learn more about managed inference, refer to [Inference Routing](/sandboxes/inference-routing).
- To configure a different self-hosted backend, refer to [Inference Routing](/sandboxes/inference-routing#configure-inference-routing).
- To learn how sandbox containers are selected, refer to [Sandboxes](/sandboxes/manage-sandboxes#custom-containers).
@@ -0,0 +1,214 @@
---
# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
title: "Route Local Inference Requests to LM Studio"
sidebar-title: "Local Inference with LM Studio"
slug: "get-started/tutorials/local-inference-lmstudio"
description: "Configure inference.local to route sandbox requests to a local LM Studio server running on the gateway host."
keywords: "Generative AI, Cybersecurity, Tutorial, Inference Routing, LM Studio, Local Inference, Sandbox"
---
This tutorial describes how to configure OpenShell to route inference requests to a local LM Studio server.
<Note>
The LM Studio server provides easy setup with both OpenAI and Anthropic compatible endpoints.
</Note>
This tutorial covers:
- Expose a local inference server to OpenShell sandboxes.
- Verify end-to-end inference from inside a sandbox.
## Prerequisites
First, complete OpenShell installation and follow the [Quickstart](/get-started/quickstart).
[Install the LM Studio app](https://lmstudio.ai/download). Make sure that your LM Studio is running in the same environment as your gateway.
If you prefer to work without having to keep the LM Studio app open, download llmster (headless LM Studio) with the following command:
<Tabs>
<Tab title="Linux/Mac">
```shell
curl -fsSL https://lmstudio.ai/install.sh | bash
```
</Tab>
<Tab title="Windows">
```shell
irm https://lmstudio.ai/install.ps1 | iex
```
</Tab>
</Tabs>
And start llmster:
```shell
lms daemon up
```
<Steps toc={true}>
## Start LM Studio Local Server
Start the LM Studio local server from the Developer tab, and verify the OpenAI-compatible endpoint is enabled.
LM Studio listens to `127.0.0.1:1234` by default. For use with OpenShell, configure LM Studio to listen on all interfaces (`0.0.0.0`).
If you use the GUI, go to the Developer Tab, select Server Settings, then enable Serve on Local Network.
If you use llmster in headless mode, run `lms server start --bind 0.0.0.0`.
## Test with a small model
In the LM Studio app, head to the Model Search tab to download a small model like Qwen3.5 2B.
In the terminal, use the following command to download and load the model:
```shell
lms get qwen/qwen3.5-2b
lms load qwen/qwen3.5-2b
```
## Add LM Studio as a provider
Choose the provider type that matches the client protocol you want to route through `inference.local`.
<Tabs>
<Tab title="OpenAI-compatible">
Add LM Studio as an OpenAI-compatible provider through `host.openshell.internal`:
```shell
openshell provider create \
--name lmstudio \
--type openai \
--credential OPENAI_API_KEY=lmstudio \
--config OPENAI_BASE_URL=http://host.openshell.internal:1234/v1
```
Use this provider for clients that send OpenAI-compatible requests such as `POST /v1/chat/completions` or `POST /v1/responses`.
</Tab>
<Tab title="Anthropic-compatible">
Add a provider that points to LM Studio's Anthropic-compatible `POST /v1/messages` endpoint:
```shell
openshell provider create \
--name lmstudio-anthropic \
--type anthropic \
--credential ANTHROPIC_API_KEY=lmstudio \
--config ANTHROPIC_BASE_URL=http://host.openshell.internal:1234
```
Use this provider for Anthropic-compatible `POST /v1/messages` requests.
</Tab>
</Tabs>
## Configure LM Studio as the local inference provider
Set the managed inference route for the active gateway:
<div className="boxed-tabs">
<Tabs>
<Tab title="OpenAI-compatible">
```shell
openshell inference set --provider lmstudio --model qwen/qwen3.5-2b
```
If the command succeeds, OpenShell has verified that the upstream is reachable and accepts the expected OpenAI-compatible request shape.
</Tab>
<Tab title="Anthropic-compatible">
```shell
openshell inference set --provider lmstudio-anthropic --model qwen/qwen3.5-2b
```
If the command succeeds, OpenShell has verified that the upstream is reachable and accepts the expected Anthropic-compatible request shape.
</Tab>
</Tabs>
</div>
The active `inference.local` route is gateway-scoped, so only one provider and model pair is active at a time. Re-run `openshell inference set` whenever you want to switch between OpenAI-compatible and Anthropic-compatible clients.
Confirm the saved config:
```shell
openshell inference get
```
You should see either `Provider: lmstudio` or `Provider: lmstudio-anthropic`, along with `Model: qwen/qwen3.5-2b`.
## Verify from Inside a Sandbox
Run a simple request through `https://inference.local`:
<Tabs>
<Tab title="OpenAI-compatible">
```shell showLineNumbers={true}
openshell sandbox create -- \
curl https://inference.local/v1/chat/completions \
--json '{"messages":[{"role":"user","content":"hello"}],"max_tokens":10}'
openshell sandbox create -- \
curl https://inference.local/v1/responses \
--json '{
"instructions": "You are a helpful assistant.",
"input": "hello",
"max_output_tokens": 10
}'
```
</Tab>
<Tab title="Anthropic-compatible">
```shell
openshell sandbox create -- \
curl https://inference.local/v1/messages \
--json '{"messages":[{"role":"user","content":"hello"}],"max_tokens":10}'
```
</Tab>
</Tabs>
</Steps>
## Troubleshooting
If setup fails, check these first:
- LM Studio local server is running and reachable from the gateway host
- `OPENAI_BASE_URL` uses `http://host.openshell.internal:1234/v1` when you use an `openai` provider
- `ANTHROPIC_BASE_URL` uses `http://host.openshell.internal:1234` when you use an `anthropic` provider
- The gateway and LM Studio run on the same machine or a reachable network path
- The configured model name matches the model exposed by LM Studio
Useful commands:
```shell
openshell status
openshell inference get
openshell provider get lmstudio
openshell provider get lmstudio-anthropic
```
## Next Steps
- To learn more about using the LM Studio CLI, refer to [LM Studio docs](https://lmstudio.ai/docs/cli)
- To learn more about managed inference, refer to [Inference Routing](/sandboxes/inference-routing).
- To configure a different self-hosted backend, refer to [Inference Routing](/sandboxes/inference-routing#configure-inference-routing).
@@ -0,0 +1,188 @@
---
# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
title: "Refresh Microsoft Graph Credentials with Providers v2"
sidebar-title: "Microsoft Graph Provider Refresh"
slug: "get-started/tutorials/microsoft-graph-provider-refresh"
description: "Configure a Providers v2 Microsoft Graph profile with gateway-managed OAuth2 refresh-token rotation."
keywords: "Generative AI, Cybersecurity, Tutorial, Providers, Microsoft Graph, OAuth2, Credential Refresh, Sandbox"
---
Use Providers v2 to keep Microsoft Graph access tokens short lived while sandboxes receive a stable `MS_GRAPH_ACCESS_TOKEN` placeholder. OpenShell stores the non-injectable refresh material at the gateway, refreshes the Microsoft Graph access token before it expires, updates the provider record, and injects the current credential into newly launched sandbox processes.
After completing this tutorial, you have:
- A custom Microsoft Graph mail provider profile.
- A provider instance configured with `oauth2-refresh-token`.
- A sandbox that can use `curl` to read Microsoft Graph mail through provider-owned policy.
<Note>
This tutorial starts after your OAuth client has already completed the initial Microsoft sign-in flow. It does not publish a token bootstrap script. Use the Microsoft identity platform documentation for the [device authorization grant flow](https://learn.microsoft.com/en-ie/entra/identity-platform/v2-oauth2-device-code) or [authorization code flow](https://learn.microsoft.com/en-us/entra/identity-platform/v2-oauth2-auth-code-flow), and use any standards-compliant client that returns an access token, refresh token, and expiry.
</Note>
## Prerequisites
- A working OpenShell installation with an active gateway. Complete the [Quickstart](/get-started/quickstart) before proceeding.
- A Microsoft Entra app registration that can acquire delegated Microsoft Graph mail access.
- Delegated Microsoft Graph mail permission for the signed-in user. `Mail.Read` allows reading the signed-in user's mailbox; see the [Microsoft Graph permissions reference](https://learn.microsoft.com/en-us/graph/permissions-reference).
OAuth material from your initial Microsoft sign-in flow:
| Variable | Value |
|---|---|
| `MS_TENANT_ID` | Microsoft Entra tenant ID, domain, or `common`. |
| `MS_CLIENT_ID` | Microsoft Entra application client ID. |
| `MS_GRAPH_ACCESS_TOKEN` | Current delegated Microsoft Graph access token. |
| `MS_GRAPH_REFRESH_TOKEN` | Delegated OAuth refresh token. |
| `MS_GRAPH_ACCESS_TOKEN_EXPIRES_AT` | Absolute expiry for the current access token. |
`MS_GRAPH_ACCESS_TOKEN_EXPIRES_AT` can be an RFC3339 timestamp such as `2026-01-01T00:00:00Z` or a Unix epoch millisecond timestamp.
<Warning>
Do not commit access tokens, refresh tokens, or local `.env` files. The commands below pass token material to the gateway; they are not examples of values to store in source control.
</Warning>
<Steps toc={true}>
## Enable Providers v2
Enable provider profile policy composition on the active gateway:
```shell
openshell settings set --global --key providers_v2_enabled --value true --yes
```
## Create a Microsoft Graph Provider Profile
Create `microsoft-graph-mail.yaml` with this profile:
```yaml showLineNumbers={false}
# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
id: microsoft-graph-mail
display_name: Microsoft Graph Mail
description: Delegated Microsoft Graph mail read access
category: messaging
credentials:
- name: graph_access_token
description: Microsoft Graph delegated access token
env_vars: [MS_GRAPH_ACCESS_TOKEN]
required: true
auth_style: bearer
header_name: authorization
refresh:
strategy: oauth2_refresh_token
token_url: https://login.microsoftonline.com/common/oauth2/v2.0/token
scopes: [https://graph.microsoft.com/.default]
refresh_before_seconds: 600
max_lifetime_seconds: 3600
material:
- name: tenant_id
description: Microsoft Entra tenant ID
required: true
- name: client_id
description: Microsoft Entra application client ID
required: true
- name: refresh_token
description: Delegated OAuth refresh token
required: true
secret: true
endpoints:
- host: graph.microsoft.com
port: 443
protocol: rest
access: read-only
enforcement: enforce
binaries:
- /usr/bin/curl
- /usr/local/bin/curl
```
Lint and import the profile:
```shell
openshell provider profile lint -f microsoft-graph-mail.yaml
openshell provider profile import -f microsoft-graph-mail.yaml
```
The profile defines the refresh strategy and Graph network policy. The `tenant_id` refresh material selects the Microsoft token endpoint during gateway-managed refresh.
## Create the Provider
Create the provider with the current Microsoft Graph access token:
```shell
openshell provider create \
--name microsoft-mail \
--type microsoft-graph-mail \
--credential MS_GRAPH_ACCESS_TOKEN="$MS_GRAPH_ACCESS_TOKEN"
```
The current CLI requires an initial credential at provider creation time. Refresh material is configured separately and is not injected into the sandbox.
## Configure Refresh
Configure gateway-managed OAuth2 refresh-token rotation:
```shell
openshell provider refresh configure microsoft-mail \
--credential-key MS_GRAPH_ACCESS_TOKEN \
--strategy oauth2-refresh-token \
--material tenant_id="$MS_TENANT_ID" \
--material client_id="$MS_CLIENT_ID" \
--material refresh_token="$MS_GRAPH_REFRESH_TOKEN" \
--secret-material-key refresh_token \
--credential-expires-at "$MS_GRAPH_ACCESS_TOKEN_EXPIRES_AT"
```
`--secret-material-key refresh_token` names the material key to mark as sensitive. It is not the refresh-token value. If Microsoft returns a rotated refresh token during refresh, OpenShell stores the new `refresh_token` material and marks it secret automatically.
Force the first refresh immediately:
```shell
openshell provider refresh rotate microsoft-mail \
--credential-key MS_GRAPH_ACCESS_TOKEN
```
Check refresh status:
```shell
openshell provider refresh status microsoft-mail \
--credential-key MS_GRAPH_ACCESS_TOKEN
```
The status output shows refresh state, expiry, next refresh, and last refresh timing. It does not print access-token values or refresh material.
## Launch a Sandbox
Launch a sandbox with the Microsoft Graph provider attached:
```shell
openshell sandbox create \
--name microsoft-graph-mail \
--provider microsoft-mail \
--no-auto-providers \
-- /bin/sh
```
Provider policy allows `curl` to reach `graph.microsoft.com:443`. The sandbox process receives `MS_GRAPH_ACCESS_TOKEN` as an OpenShell placeholder, and the proxy resolves that placeholder to the current gateway-managed access token when `curl` sends it in the authorization header.
## Verify Microsoft Graph Access
Inside the sandbox, list a small page of mailbox messages:
```shell
curl -sS \
-H "Authorization: Bearer $MS_GRAPH_ACCESS_TOKEN" \
'https://graph.microsoft.com/v1.0/me/messages?$select=sender,subject&$top=5'
```
The request uses the [Microsoft Graph list messages API](https://learn.microsoft.com/en-us/graph/api/user-list-messages?view=graph-rest-1.0). If the token has delegated mail read permission, Microsoft Graph returns message metadata for the signed-in user's mailbox.
## Update Running Sandboxes
Provider refresh updates the provider record at the gateway. Running sandboxes poll for provider environment revisions, but already-running processes keep the environment they started with.
If you attach this provider to an existing sandbox or update provider credentials after a process has already started, launch a new process inside the sandbox before expecting `MS_GRAPH_ACCESS_TOKEN` to appear in that process environment.
</Steps>
+139
View File
@@ -0,0 +1,139 @@
---
# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
title: "NVIDIA OpenShell Developer Guide"
description: "OpenShell is the safe, private runtime for autonomous AI agents. Run agents in sandboxed environments that protect your data, credentials, and infrastructure."
keywords: "Generative AI, Cybersecurity, AI Agents, Sandboxing, Security, Privacy, Inference Routing"
position: 1
---
import { BadgeLinks } from "./_components/BadgeLinks";
import { CommandTerminal } from "./_components/CommandTerminal";
<BadgeLinks
badges={[
{
href: "https://github.com/NVIDIA/OpenShell",
src: "https://img.shields.io/badge/github-repo-green?logo=github",
alt: "GitHub",
},
{
href: "https://github.com/NVIDIA/OpenShell/blob/main/LICENSE",
src: "https://img.shields.io/badge/License-Apache_2.0-blue",
alt: "License",
},
{
href: "https://pypi.org/project/openshell/",
src: "https://img.shields.io/badge/PyPI-openshell-orange?logo=pypi",
alt: "PyPI",
},
]}
/>
NVIDIA OpenShell is the safe, private runtime for autonomous AI agents. It provides sandboxed execution environments
that protect your data, credentials, and infrastructure. Agents run with exactly the permissions they need and
nothing more, governed by declarative policies that prevent unauthorized file access, data exfiltration, and
uncontrolled network activity.
## Get Started
Install OpenShell and create your first sandbox in two commands.
<llms-ignore>
<CommandTerminal command="curl -LsSf https://raw.githubusercontent.com/NVIDIA/OpenShell/main/install.sh | sh" />
</llms-ignore>
<llms-only>
```shell
curl -LsSf https://raw.githubusercontent.com/NVIDIA/OpenShell/main/install.sh | sh
openshell sandbox create -- claude
```
</llms-only>
Refer to the [Quickstart](/get-started/quickstart) for more details.
---
## Explore
<div className="explore-cards">
<Cards>
<Card title="About OpenShell" href="/about/overview">
Learn about OpenShell and its capabilities.
<Badge intent="tip" minimal outlined>Concept</Badge>
</Card>
<Card title="Quickstart" href="/get-started/quickstart">
Install OpenShell and create your first sandbox in two commands.
<Badge intent="tip" minimal outlined>Tutorial</Badge>
</Card>
<Card title="Tutorials" href="/get-started/tutorials">
Hands-on walkthroughs from first sandbox to custom policies.
<Badge intent="tip" minimal outlined>Concept</Badge>
</Card>
<Card title="Gateways and Sandboxes" href="/sandboxes/manage-gateways">
Deploy gateways, create sandboxes, configure policies, providers, and community images for your AI agents.
<Badge intent="tip" minimal outlined>Concept</Badge>
</Card>
<Card title="Inference Routing" href="/sandboxes/inference-routing">
Keep inference traffic private by routing API calls to local or self-hosted backends.
<Badge intent="tip" minimal outlined>Concept</Badge>
</Card>
<Card title="Observability" href="/observability">
Understand sandbox logs, access them with the CLI and TUI, and export OCSF JSON records.
<Badge intent="tip" minimal outlined>How-To</Badge>
</Card>
<Card title="Reference" href="/reference/default-policy">
Policy schema, environment variables, and default policy details.
<Badge intent="tip" minimal outlined>Reference</Badge>
</Card>
<Card title="Security Best Practices" href="/security/best-practices">
Every configurable security control, its default, and the risk of changing it.
<Badge intent="tip" minimal outlined>Concept</Badge>
</Card>
</Cards>
</div>
---
<Warning title="Notice and Disclaimer">
This software automatically retrieves, accesses or interacts with external
materials. Those retrieved materials are not distributed with this software
and are governed solely by separate terms, conditions and licenses. You are
solely responsible for finding, reviewing and complying with all applicable
terms, conditions, and licenses, and for verifying the security, integrity and
suitability of any retrieved materials for your specific use case. This
software is provided "AS IS", without warranty of any kind. The author makes
no representations or warranties regarding any retrieved materials, and
assumes no liability for any losses, damages, liabilities or legal
consequences from your use or inability to use this software or any retrieved
materials. Use this software and the retrieved materials at your own risk.
</Warning>
+33
View File
@@ -0,0 +1,33 @@
# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
landing-page:
page: "Home"
path: index.mdx
navigation:
- folder: about
title: "About NVIDIA OpenShell"
- section: "Get Started"
slug: get-started
contents:
- page: "Quickstart"
path: get-started/quickstart.mdx
- folder: get-started/tutorials
skip-slug: true
- folder: sandboxes
title: "Manage OpenShell"
- folder: providers
title: "Providers"
- folder: extensibility
title: "Extensibility"
- folder: observability
title: "Observability"
- folder: kubernetes
title: "Kubernetes"
- folder: reference
title: "Reference"
- folder: security
title: "Security"
- folder: resources
title: "Resources"
@@ -0,0 +1,118 @@
---
# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
title: "Access Control"
sidebar-title: "Access Control"
description: "Configure OIDC user authentication or reverse-proxy auth termination for a Kubernetes-deployed OpenShell gateway."
keywords: "Generative AI, Cybersecurity, Kubernetes, Authentication, mTLS, OIDC, Keycloak, Entra ID, Okta, Gateway Auth"
position: 5
---
The OpenShell gateway supports two access-control models for human callers on Kubernetes:
| Model | When to use |
|---|---|
| OIDC (recommended) | Production deployments. Integrates with an existing identity provider, supports role-based access control, and gives each user their own identity without distributing certificates. |
| Reverse-proxy auth termination | An access proxy (Cloudflare Access, ngrok, corporate SSO) authenticates callers in front of the gateway. The gateway trusts the proxy and skips its own client-cert check. |
The Helm chart always generates mTLS certificates at install time. The gateway uses them for transport-layer security regardless of which access-control model you choose. The client bundle in the `openshell-client-tls` secret is used internally by sandbox supervisors, not for granting access to individual users.
For how the CLI resolves gateways and stores credentials, refer to [Gateway Authentication](/reference/gateway-auth).
## Sandbox Supervisor Identity
Kubernetes sandbox supervisors authenticate back to the gateway as sandbox workloads. By default, the gateway mints its own sandbox JWTs and Kubernetes sandboxes bootstrap them with a projected ServiceAccount token.
Dynamic provider token grants can use SPIFFE without changing supervisor-to-gateway authentication. Set `server.providerTokenGrants.spiffe.enabled=true` to mount the SPIFFE CSI Workload API socket into gateway and sandbox pods while keeping the projected ServiceAccount token bootstrap and gateway-minted sandbox JWT path.
Provider token grants require a SPIFFE implementation such as SPIRE and identities for the gateway and sandbox pods. The repository's local SPIRE overlay assigns sandbox IDs from the pod's `openshell.io/sandbox-id` annotation, but the gateway validation path only requires the supervisor SVID to be valid and in the same SPIFFE trust domain as the gateway SVID. Provider profiles with `token_grant` metadata cause the sandbox supervisor to request JWT-SVIDs and exchange them for upstream OAuth2 access tokens. Token-exchange profiles also require a gateway SPIFFE identity because the gateway brokers the intermediate token exchange with its own JWT-SVID.
The gateway verifies supervisor JWT-SVIDs with JWT bundles fetched from the SPIFFE Workload API, so intermediate token exchange does not require gateway access to the SPIRE OIDC discovery endpoint or its TLS CA.
## OIDC User Authentication
Set `server.oidc.issuer` to enable OIDC. The gateway validates the `Authorization: Bearer <token>` header on every request against the issuer's JWKS endpoint.
```shell
helm upgrade openshell \
oci://ghcr.io/nvidia/openshell/helm-chart \
--version <version> \
--namespace openshell \
--set server.oidc.issuer=https://your-idp.example.com/realms/openshell \
--set server.oidc.audience=openshell-cli
```
The `audience` value must match the client ID configured in your identity provider for the OpenShell resource server.
### OIDC values reference
| Value | Default | Purpose |
|---|---|---|
| `server.oidc.issuer` | `""` | OIDC issuer URL. Empty disables OIDC. |
| `server.oidc.audience` | `openshell-cli` | Expected `aud` claim in the JWT. |
| `server.oidc.jwksTtl` | `3600` | JWKS key cache TTL in seconds. |
| `server.oidc.rolesClaim` | `""` | Dot-separated path to the roles array in JWT claims. |
| `server.oidc.adminRole` | `""` | Role name that grants admin access. |
| `server.oidc.userRole` | `""` | Role name that grants standard user access. |
| `server.oidc.scopesClaim` | `""` | Dot-separated path to the scopes array in JWT claims. |
### Auth-only mode vs. RBAC mode
Leave both `adminRole` and `userRole` empty to use auth-only mode: any request with a valid JWT from the configured issuer is accepted, but no role distinction is enforced.
Set both values to enable RBAC mode, where the gateway checks the role claim and enforces access based on the assigned role:
```shell
helm upgrade openshell \
oci://ghcr.io/nvidia/openshell/helm-chart \
--version <version> \
--namespace openshell \
--set server.oidc.issuer=https://your-idp.example.com/realms/openshell \
--set server.oidc.audience=openshell-cli \
--set server.oidc.rolesClaim=realm_access.roles \
--set server.oidc.adminRole=openshell-admin \
--set server.oidc.userRole=openshell-user
```
Both `adminRole` and `userRole` must be set, or both must be empty. Setting only one is not supported.
OIDC RBAC is method-level authorization. It controls which API operations a caller can perform, but provider and sandbox records are not owned by individual OIDC subjects. In shared clusters, treat provider credentials as gateway-wide resources and use separate gateways or external tenancy controls when users must not see or attach each other's providers and sandboxes.
### Provider-specific rolesClaim paths
| Provider | rolesClaim value |
|---|---|
| Keycloak | `realm_access.roles` |
| Microsoft Entra ID | `roles` |
| Okta | `groups` |
## Reverse-Proxy Auth Termination
When an access proxy, such as Cloudflare Access, ngrok, or a corporate SSO gateway, handles authentication in front of the OpenShell gateway, you can explicitly allow unauthenticated user calls at the gateway:
```shell
helm upgrade openshell \
oci://ghcr.io/nvidia/openshell/helm-chart \
--version <version> \
--namespace openshell \
--set server.auth.allowUnauthenticatedUsers=true
```
The gateway still serves TLS and sandbox supervisors still authenticate with gateway-minted sandbox JWTs. User-facing CLI/API calls without OIDC or mTLS credentials are accepted as an unauthenticated local developer principal. The proxy is responsible for authenticating callers and forwarding only authorized traffic.
To also disable TLS entirely (when the proxy terminates TLS before the request reaches the gateway):
```shell
--set server.disableTls=true \
--set server.auth.allowUnauthenticatedUsers=true
```
<Warning>
Only enable unauthenticated users when the gateway is not reachable from outside a trusted local development environment or the proxy path is fully trusted. Never expose a plaintext, auth-disabled gateway to a public network.
</Warning>
Register the gateway with the CLI using the proxy's public URL. The browser-based login flow runs automatically on first use:
```shell
openshell gateway add https://gateway.example.com --name production
```
+144
View File
@@ -0,0 +1,144 @@
---
# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
title: "Ingress"
sidebar-title: "Ingress"
description: "Expose the OpenShell gateway externally using the Kubernetes Gateway API and a GRPCRoute."
keywords: "Generative AI, Cybersecurity, Kubernetes, Gateway API, Envoy Gateway, GRPCRoute, Ingress, External Access"
position: 4
---
By default, the OpenShell gateway is only reachable inside the cluster. To let CLI clients connect without a `kubectl port-forward`, expose the gateway through an ingress.
OpenShell uses the [Kubernetes Gateway API](https://gateway-api.sigs.k8s.io) for ingress. The chart creates a `GRPCRoute` that routes inbound gRPC traffic to the gateway pod. You need a Gateway API implementation installed on your cluster to fulfill the `GRPCRoute`. This page uses [Envoy Gateway](https://gateway.envoyproxy.io), which the chart is tested with.
## Install Envoy Gateway
Envoy Gateway installs the Gateway API CRDs and controller:
```shell
helm install eg \
oci://docker.io/envoyproxy/gateway-helm \
--version v1.8.1 \
--namespace envoy-gateway-system \
--create-namespace \
--wait
```
## Create the GatewayClass
Create the `eg` GatewayClass that the OpenShell chart references:
```shell
kubectl apply -f - <<'EOF'
apiVersion: gateway.networking.k8s.io/v1
kind: GatewayClass
metadata:
name: eg
spec:
controllerName: gateway.envoyproxy.io/gatewayclass-controller
EOF
```
Verify the GatewayClass is accepted:
```shell
kubectl get gatewayclass eg
```
The `ACCEPTED` column should show `True`.
## Install OpenShell with Gateway API enabled
Enable the GRPCRoute and let the chart create a Gateway resource in the `openshell` namespace:
```shell
helm upgrade --install openshell \
oci://ghcr.io/nvidia/openshell/helm-chart \
--version <version> \
--namespace openshell \
--set grpcRoute.enabled=true \
--set grpcRoute.gateway.create=true \
--set grpcRoute.gateway.className=eg
```
## Get the external address
After the Gateway is provisioned, Envoy Gateway creates a LoadBalancer service in the `openshell` namespace. Wait for it to get an external address:
```shell
kubectl -n openshell get svc -l gateway.envoyproxy.io/owning-gateway-name=openshell
```
After the `EXTERNAL-IP` is assigned, register the gateway with the CLI:
```shell
openshell gateway add http://<external-ip> --name production
openshell status
```
This setup is plaintext end-to-end and is intended for development. For external access, terminate TLS at the gateway as shown below.
## HTTPS (TLS termination)
Envoy Gateway can terminate TLS at the listener and forward plaintext to the OpenShell gateway pod:
```text
client → HTTPS → Envoy Gateway (terminate TLS) → plaintext → openshell gateway pod
```
Envoy Gateway only terminates TLS here — it does not perform OIDC. Do not enable an Envoy Gateway OIDC `SecurityPolicy` in front of the gateway: that flow relies on browser redirects and cookies and cannot work with the OpenShell CLI or headless agents. Instead, the OpenShell gateway validates an OIDC bearer token that the client sends in the gRPC `authorization` metadata, which Envoy forwards untouched.
Because Envoy terminates TLS, the OpenShell gateway never sees a client certificate, so client mTLS cannot provide identity on this path. Use OIDC bearer tokens for client identity instead.
For an interactive login on a headless machine, set `OPENSHELL_NO_BROWSER=1` before running `openshell gateway add`. The CLI uses the Device Authorization Grant with S256 PKCE and prompts the user to approve the login from another browser. For unattended agents and CI, set `OPENSHELL_OIDC_CLIENT_SECRET` to use the OAuth2 client-credentials grant instead. The client id comes from `--oidc-client-id` (default `openshell-cli`); pass it explicitly if your identity provider uses a different id. Interactive users with a local browser get the Authorization Code flow with PKCE by default.
### Provide a TLS certificate
Create a `kubernetes.io/tls` Secret in the `openshell` namespace with the certificate for your external hostname:
```shell
kubectl -n openshell create secret tls openshell-ingress-tls \
--cert=tls.crt --key=tls.key
```
The Secret may also be issued by cert-manager, or you can reference the chart's existing `openshell-server-tls` Secret if its SANs include the external hostname.
### Install with HTTPS termination
Enable an HTTPS listener, point it at the Secret, disable gateway-pod TLS so Envoy forwards plaintext, and configure an OIDC issuer for client identity:
```shell
helm upgrade --install openshell \
oci://ghcr.io/nvidia/openshell/helm-chart \
--version <version> \
--namespace openshell \
--set grpcRoute.enabled=true \
--set grpcRoute.gateway.create=true \
--set grpcRoute.gateway.className=eg \
--set grpcRoute.gateway.listener.protocol=HTTPS \
--set grpcRoute.gateway.listener.port=443 \
--set 'grpcRoute.gateway.listener.tls.certificateRefs[0].name=openshell-ingress-tls' \
--set server.disableTls=true \
--set server.oidc.issuer=https://<issuer> \
--set 'grpcRoute.hostnames[0]=<external-hostname>'
```
Keep the certificate Secret in the release namespace. Referencing a Secret in another namespace requires a `ReferenceGrant`.
### Register over HTTPS
```shell
openshell gateway add https://<external-hostname> --name production --oidc-issuer https://<issuer>
openshell status
```
See [Authentication](/kubernetes/setup) for OIDC issuer, audience, and roles configuration.
## SSH Relay
Sandbox SSH uses the gateway endpoint registered with the CLI. No separate Helm SSH host or port values are required.
## Next Steps
Return to [Setup](/kubernetes/setup) to complete the installation.
@@ -0,0 +1,122 @@
---
# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
title: "Managing Certificates"
sidebar-title: "Managing Certificates"
description: "Configure the OpenShell Helm chart to use cert-manager for mTLS certificate issuance and automatic renewal."
keywords: "Generative AI, Cybersecurity, Kubernetes, cert-manager, PKI, TLS, mTLS, Certificates"
position: 3
---
The OpenShell gateway uses mTLS certificates for transport between the gateway and sandbox supervisors. These certificates are not Kubernetes user authentication; configure OIDC or a trusted access proxy for user access. The Helm chart supports two ways to provision and manage the certificate bundle:
| Mode | When to use |
|---|---|
| Built-in `pkiInitJob` (default) | The default path. A pre-install Kubernetes Job generates a self-signed CA and certificates during installation. No additional dependencies. |
| cert-manager | Production deployments that need automatic certificate rotation managed by a running controller. |
The rest of this page covers switching to cert-manager. The built-in mode requires no configuration.
<Note>
When `certManager.enabled=true`, cert-manager owns TLS certificate generation.
The chart still runs a JWT-only initialization hook because cert-manager does
not create the sandbox JWT signing Secret required by the gateway. This
cert-manager precedence applies even if `pkiInitJob.enabled` remains true.
</Note>
## Install cert-manager
Install cert-manager from the OCI registry with CRD support enabled:
```shell
helm upgrade --install cert-manager oci://quay.io/jetstack/charts/cert-manager \
--version v1.20.3 \
--namespace cert-manager \
--create-namespace \
--set crds.enabled=true \
--wait
```
Verify the cert-manager pods are running:
```shell
kubectl -n cert-manager get pods
```
## Install OpenShell with cert-manager PKI
Pass the cert-manager values override when installing or upgrading the chart:
```shell
helm upgrade --install openshell \
oci://ghcr.io/nvidia/openshell/helm-chart \
--version <version> \
--namespace openshell \
--set certManager.enabled=true
```
The chart creates a self-signed CA, issues server and client certificates from it, and cert-manager handles renewal before expiry.
The chart also runs a pre-install hook in JWT-only mode to create the gateway's
sandbox JWT signing Secret. That Secret is separate from the cert-manager TLS
certificate Secrets and is mounted at `/etc/openshell-jwt`.
## Using a real Issuer for the server certificate
By default, cert-manager issues both the server and client certificates from
a self-signed CA the chart creates — this rotates automatically, but the
server certificate is still not publicly trusted. `certManager.serverIssuerRef`
overrides the `issuerRef` on the server `Certificate` resource to point at a
real `Issuer` or `ClusterIssuer` instead, for example an ACME issuer:
```shell
helm upgrade --install openshell \
oci://ghcr.io/nvidia/openshell/helm-chart \
--version <version> \
--namespace openshell \
--set certManager.enabled=true \
--set certManager.serverIssuerRef.name=letsencrypt-prod \
--set certManager.serverIssuerRef.kind=ClusterIssuer \
--set certManager.serverDnsNames[0]=openshell.example.com
```
### Dual certificate architecture
When `serverIssuerRef` is set, the chart creates **two** server certificates:
1. **Internal certificate** (`openshell-server-tls`): signed by the chart CA
with internal SANs (`*.svc.cluster.local`, `localhost`, etc.).
2. **External certificate** (`openshell-server-external-tls`): signed by the
configured issuer (e.g. ACME) with only the hostnames from
`certManager.serverDnsNames`.
The gateway uses **SNI** to select which certificate to present:
supervisors connect via internal service names and receive the internal
certificate (verified against the chart CA they already trust), while CLI
users connecting through a Route or ingress use the external hostname and
receive the ACME certificate. This keeps supervisor trust pinned to only
the operator's chart CA — no WebPKI root trust is needed.
<Note>
Public CAs such as Let's Encrypt reject certificate requests that include
internal-only names per CA/Browser Forum baseline requirements. The chart
validates this at install time and fails with an actionable error if
`certManager.serverDnsNames` contains internal-only entries while
`serverIssuerRef` is set.
You do **not** need to set `server.grpcEndpoint` to the external hostname.
Supervisors connect via the internal service name automatically. Setting
`server.grpcEndpoint` to an external hostname would cause supervisors to
receive the ACME certificate (via SNI) which they cannot verify against the
chart CA.
</Note>
The default `clientCaFromServerTlsSecret=true` is correct even when
`serverIssuerRef` is set: the internal server certificate is always signed
by the chart CA (the same CA that signs the client certificate), so its
`ca.crt` is the right trust anchor for mTLS verification.
## Next Steps
Return to [Setup](/kubernetes/setup) to complete the installation. For
exposing the gateway externally on OpenShift with a real certificate, see
[OpenShift](/kubernetes/openshift).
@@ -0,0 +1,142 @@
---
# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
title: "OpenShift"
sidebar-title: "OpenShift"
description: "Install the OpenShell Helm chart on OpenShift, including the SCC binding and chart overrides required by OpenShift's Security Context Constraints."
keywords: "Generative AI, Cybersecurity, Kubernetes, OpenShift, SCC, Security Context Constraints, Helm, Gateway, Installation"
position: 6
---
<Warning>
The OpenShift install path is experimental. It currently requires running sandbox pods under the `privileged` SCC and installing the gateway with TLS disabled. Use only for evaluation on a private network.
</Warning>
OpenShift's [Security Context Constraints](https://docs.openshift.com/container-platform/latest/authentication/managing-security-context-constraints.html) reject the chart's default pod security settings. Installing on OpenShift requires precreating the namespace, granting the `privileged` SCC to the sandbox service account, and overriding a few chart values so the cluster admission controller can assign UIDs and FS groups itself.
OpenShell installs sandbox nftables rules as individual commands. On OpenShift
nodes where optional conntrack or packet log expressions are unavailable, those
optional rules can fail without rolling back the required proxy bypass reject
rules.
## Prerequisites
- OpenShift 4.x cluster with `oc` configured
- Helm 3.x
- [Agent Sandbox](/kubernetes/setup#install-agent-sandbox) controller and CRDs installed
## Install
<Steps>
## Create the namespace
Pre-create the namespace so the SCC binding can be applied before the chart installs:
```shell
oc create ns openshell
```
## Grant the privileged SCC to sandbox pods
Sandbox pods run under the `openshell-sandbox` service account in the `openshell` namespace and require the `privileged` SCC:
```shell
oc adm policy add-scc-to-user privileged -z openshell-sandbox -n openshell
```
## Install the chart with OpenShift overrides
```shell
helm install openshell oci://ghcr.io/nvidia/openshell/helm-chart \
--version <version> \
--namespace openshell \
--set server.disableTls=true \
--set podSecurityContext.fsGroup=null \
--set securityContext.runAsUser=null
```
| Override | Reason |
|---|---|
| `server.disableTls=true` | Runs the gateway over plaintext HTTP for simpler evaluation. |
| `podSecurityContext.fsGroup=null` / `securityContext.runAsUser=null` | Clear the chart's hardcoded UID and fsGroup so OpenShift's SCC admission can assign them. |
## Wait for the gateway to be ready
```shell
oc -n openshell rollout status statefulset/openshell
```
If you set `workload.kind=deployment`, use
`oc -n openshell rollout status deployment/openshell` instead.
</Steps>
## Connect to the gateway
The gateway is now running over plaintext HTTP. Connect with `oc port-forward`:
```shell
oc -n openshell port-forward svc/openshell 8080:8080
```
Register the gateway with the CLI:
```shell
openshell gateway add http://127.0.0.1:8080 --local --name openshift
openshell status
```
## Production: expose externally with a real certificate
The steps above run the gateway over plaintext HTTP for quick evaluation. For
a real deployment, cert-manager can issue the gateway's server certificate
from a real Issuer or ClusterIssuer (for example, an ACME issuer), and an
OpenShift Route with TLS passthrough exposes it externally while the gateway
keeps terminating its own TLS and mTLS.
Install cert-manager and configure a working `ClusterIssuer` first — see
[Managing Certificates](/kubernetes/managing-certificates) for the
`certManager.serverIssuerRef` details. Configure an OIDC provider as described
in [Access Control](/kubernetes/access-control) — remote gateways authenticate
CLI users via OIDC, not mTLS, so the gateway must know the OIDC issuer URL.
Install the chart with:
```shell
helm install openshell oci://ghcr.io/nvidia/openshell/helm-chart \
--version <version> \
--namespace openshell \
--set podSecurityContext.fsGroup=null \
--set securityContext.runAsUser=null \
--set server.disableTls=false \
--set certManager.enabled=true \
--set certManager.serverIssuerRef.name=<cluster-issuer-name> \
--set certManager.serverIssuerRef.kind=ClusterIssuer \
--set certManager.serverDnsNames[0]=<external-hostname> \
--set openshiftRoute.enabled=true \
--set openshiftRoute.host=<external-hostname> \
--set server.oidc.issuer=<oidc-issuer-url> \
--set server.oidc.audience=<oidc-audience>
```
| Override | Reason |
|---|---|
| `certManager.serverIssuerRef` | Creates a second server certificate from your Issuer or ClusterIssuer for external clients. The gateway uses SNI to present this cert for the external hostname while continuing to present the internal (chart CA) cert to supervisors. The internal certificate's `ca.crt` is the chart CA that also signed the client cert, so the default `clientCaFromServerTlsSecret=true` is correct. |
| `openshiftRoute.enabled` / `openshiftRoute.host` | Creates an OpenShift Route with TLS passthrough — the router forwards the encrypted connection by SNI without decrypting, so the gateway uses the SNI hostname to select the external certificate. |
| `server.oidc.issuer` / `server.oidc.audience` | Configures server-side OIDC validation. Without these, the gateway expects mTLS client certificates and rejects OIDC-only CLI connections. See [Access Control](/kubernetes/access-control). |
Register the gateway with the CLI over OIDC. Remote gateways authenticate CLI
users via OIDC, not mTLS — see [Access Control](/kubernetes/access-control):
```shell
openshell gateway add https://<external-hostname> \
--name openshift \
--oidc-issuer <issuer-url>
openshell gateway login openshift
```
## Next Steps
- For more on certificate provisioning modes, refer to [Managing Certificates](/kubernetes/managing-certificates).
- To expose the gateway externally through the Kubernetes Gateway API instead of a Route, refer to [Ingress](/kubernetes/ingress).
- To configure OIDC authentication, refer to [Access Control](/kubernetes/access-control).
+297
View File
@@ -0,0 +1,297 @@
---
# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
title: "Set Up OpenShell on Kubernetes"
sidebar-title: "Setup"
description: "Deploy the OpenShell gateway to a Kubernetes cluster using the official Helm chart from GHCR."
keywords: "Generative AI, Cybersecurity, Kubernetes, Helm, Gateway, Deployment, OCI, GHCR, Installation"
position: 1
---
<Warning>
The OpenShell Helm chart is experimental and under active development. Templates, values, and defaults can change between releases. Do not use it in production.
</Warning>
Use the Kubernetes deployment when the gateway should run on a shared cluster, in a cloud environment, or as part of team infrastructure. The Helm chart handles PKI bootstrap, RBAC, sandbox namespace setup, and the gateway workload. It uses a StatefulSet by default for the SQLite database, and can render a Deployment when `server.externalDbSecret` points at an external database.
## Prerequisites
Make sure the following are in place before you install.
| Prerequisite | Required | Notes |
|---|---|---|
| Kubernetes 1.29+ with RBAC enabled | Yes | No additional notes. |
| Helm 3.x | Yes | No additional notes. |
| Agent Sandbox controller and CRDs | Yes | Install before the OpenShell chart. Refer to [Install Agent Sandbox](#install-agent-sandbox). |
| cert-manager | No | Refer to [Managing Certificates](/kubernetes/managing-certificates). Use cert-manager only if you prefer it over the built-in PKI job. |
| Kubernetes Gateway API | No | Refer to [Ingress](/kubernetes/ingress). Use it only for external access without port-forwarding. |
## Install Agent Sandbox
OpenShell uses the [Agent Sandbox](https://agent-sandbox.sigs.k8s.io) Kubernetes SIG project to provision sandbox pods. Install the Agent Sandbox controller and its CRDs on your cluster before installing the OpenShell Helm chart.
Apply the latest release manifest:
```shell
kubectl apply -f https://github.com/kubernetes-sigs/agent-sandbox/releases/latest/download/manifest.yaml
```
This creates the `agent-sandbox-system` namespace, installs the `sandboxes.agents.x-k8s.io` CRD, and starts the controller.
The Helm chart checks for a supported Agent Sandbox API before it creates
gateway resources. This preflight is enabled by default. Disable it only for
offline `helm template` rendering, where Helm cannot discover cluster APIs:
```shell
helm template openshell oci://ghcr.io/nvidia/openshell/helm-chart \
--version <version> \
--set agentSandbox.preflight.enabled=false
```
The chart does not install or upgrade the cluster-scoped Agent Sandbox CRDs or
controller.
<Note>
**Air-gapped clusters:** mirror the manifest above and the `registry.k8s.io/agent-sandbox/agent-sandbox-controller` image referenced inside it to your internal registry, then point the manifest's image reference at your mirror before applying. You will also need to mirror the OpenShell gateway and sandbox images — see the chart's `image.repository` value for the gateway and `server.sandboxImage` / `server.supervisorImage` for the sandbox runtime.
</Note>
Confirm the controller pod is running before proceeding:
```shell
kubectl -n agent-sandbox-system get pods
```
The controller pod should reach `Running` status within a few seconds. For cluster-specific setup instructions, including KinD and GKE walkthroughs, refer to the [Agent Sandbox getting started guide](https://agent-sandbox.sigs.k8s.io/docs/getting_started/).
### Upgrade Agent Sandbox
OpenShell detects the served Agent Sandbox `Sandbox` API when the Kubernetes gateway first needs it and caches that choice for the gateway process. If you upgrade Agent Sandbox in place, restart the OpenShell gateway after the Agent Sandbox controller and CRD rollout completes so the gateway can detect the served API versions again. Existing sandboxes keep running during the upgrade, and the restarted gateway can continue managing them.
## Install OpenShell
<Steps>
## Create the namespace
```shell
kubectl create namespace openshell
```
## Install the chart
Install from the OCI registry on GHCR. Replace `<version>` with the chart version you want to install.
```shell
helm upgrade --install openshell \
oci://ghcr.io/nvidia/openshell/helm-chart \
--version <version> \
--namespace openshell
```
To use the latest development build instead of a stable release:
```shell
helm upgrade --install openshell \
oci://ghcr.io/nvidia/openshell/helm-chart \
--version 0.0.0-dev \
--namespace openshell
```
The chart automatically generates PKI secrets on first install using pre-install Helm hooks. No manual secret creation is required.
## Wait for the gateway to be ready
```shell
kubectl -n openshell rollout status statefulset/openshell
```
If you set `workload.kind=deployment`, wait on the Deployment instead:
```shell
kubectl -n openshell rollout status deployment/openshell
```
## Connect to the gateway
For local evaluation, use a port-forward:
```shell
kubectl -n openshell port-forward svc/openshell 8080:8080
```
<Warning>
The port-forward is for local evaluation only. For shared environments, expose the gateway through your ingress controller or access proxy. Refer to [Ingress](/kubernetes/ingress) for an external access option.
</Warning>
## Install the TLS client bundle
The chart generates an mTLS bundle for transport security. Kubernetes deployments do not use that bundle as user authentication; configure OIDC or a trusted access proxy as described in [Access Control](/kubernetes/access-control). For local port-forwarded access, copy the generated bundle so the CLI can verify the gateway certificate:
```shell
mkdir -p ~/.config/openshell/gateways/k8s/mtls
kubectl -n openshell get secret openshell-client-tls \
-o jsonpath='{.data.ca\.crt}' | base64 -d > ~/.config/openshell/gateways/k8s/mtls/ca.crt
kubectl -n openshell get secret openshell-client-tls \
-o jsonpath='{.data.tls\.crt}' | base64 -d > ~/.config/openshell/gateways/k8s/mtls/tls.crt
kubectl -n openshell get secret openshell-client-tls \
-o jsonpath='{.data.tls\.key}' | base64 -d > ~/.config/openshell/gateways/k8s/mtls/tls.key
```
The server certificate SANs include `localhost` and `127.0.0.1`, so hostname verification passes over the port-forward without extra flags.
## Register with the CLI
In another terminal, register the gateway with the user authentication mode you configured and verify it is reachable. For example, with OIDC:
```shell
openshell gateway add https://127.0.0.1:8080 --local --name k8s \
--oidc-issuer https://your-idp.example.com/realms/openshell \
--oidc-client-id openshell-cli
openshell status
```
</Steps>
## Configure Chart Values
The most commonly changed values are:
| Value | Purpose |
|---|---|
| `image.repository` / `image.tag` | Gateway container image. Defaults to `ghcr.io/nvidia/openshell/gateway:latest`. |
| `replicaCount` | Number of gateway replicas. Leave at `1` unless you are explicitly testing multi-replica behavior. |
| `workload.kind` | Gateway workload controller. Use `statefulset` for SQLite or `deployment` with `server.externalDbSecret`. |
| `workload.allowMultiReplicaStatefulSet` | Allow `replicaCount > 1` with `workload.kind=statefulset`. Prefer Deployment for external database-backed multi-replica gateways. |
| `server.sandboxNamespace` | Namespace where sandbox pods are created. Defaults to the Helm release namespace when left empty. |
| `server.externalDbSecret` | Secret containing a PostgreSQL connection URI in the `uri` key. Use when the database is managed outside the chart. |
| `server.telemetryEnabled` | Enable anonymous OpenShell telemetry from the gateway and its sandbox supervisors. Set to `false` to opt out. |
| `server.sandboxImage` | Default sandbox image used when a sandbox does not specify one. |
| `server.sandboxImagePullSecrets` | Image pull secrets attached to sandbox pods. Referenced Secrets must exist in the sandbox namespace. |
| `server.grpcEndpoint` | Endpoint that sandbox supervisors use to call back to the gateway. Must be reachable from inside the cluster. |
| `server.appArmorProfile` | AppArmor profile requested for sandbox agent containers. Defaults to `Unconfined`. |
| `server.disableTls` | Run the gateway over plaintext HTTP. Use only behind a trusted transport. |
| `server.auth.allowUnauthenticatedUsers` | Accept user-facing calls without OIDC or mTLS credentials. Use only for trusted local development or a fully trusted access proxy. |
| `server.enableLoopbackServiceHttp` | Enable local plaintext HTTP for loopback sandbox service URLs. Defaults to `true`. |
| `pkiInitJob.serverDnsNames` / `certManager.serverDnsNames` | Additional gateway server DNS SANs. Wildcard SANs also enable sandbox service URLs under that domain. |
| `supervisor.sideloadMethod` | How the supervisor binary is delivered into sandbox pods. Leave empty to auto-detect based on cluster version: clusters running Kubernetes 1.35 or later use `image-volume` (ImageVolume GA in 1.36); older clusters use `init-container`. Set explicitly to `image-volume` on Kubernetes 1.33 or 1.34 with the ImageVolume feature gate enabled, or to `init-container` to force the legacy path on any version. |
| `supervisor.topology` | Sandbox pod topology. Refer to [Topology](/kubernetes/topology). |
| `supervisor.sidecar.proxyUid` | Non-root UID used when sidecar process/binary-aware network policy is disabled. The default binary-aware sidecar runs as UID 0 instead. The configured UID must not match the sandbox UID. |
| `upstreamProxy` | Operator-owned corporate HTTP forward proxy for policy-approved TLS egress. Refer to [Configure a Corporate Upstream Proxy](#configure-a-corporate-upstream-proxy). |
Use a values file for repeatable deployments:
```shell
helm upgrade --install openshell \
oci://ghcr.io/nvidia/openshell/helm-chart \
--version <version> \
--namespace openshell \
--values my-values.yaml
```
The chart defaults `server.appArmorProfile` to `Unconfined` because
runtime/default AppArmor profiles can block the supervisor's network namespace
mount setup on AppArmor-enabled nodes. Set `server.appArmorProfile` to an empty
string to omit the field, `RuntimeDefault` to force the runtime default, or
`Localhost/<profile-name>` when you load and manage a localhost profile on each
node.
To use private sandbox images, create a `kubernetes.io/dockerconfigjson` Secret
in the sandbox namespace and reference its name:
```shell
kubectl -n openshell create secret docker-registry regcred \
--docker-server=registry.example.com \
--docker-username="$REGISTRY_USER" \
--docker-password="$REGISTRY_TOKEN"
```
```yaml
server:
sandboxImage: registry.example.com/team/openshell-sandbox:latest
sandboxImagePullSecrets:
- name: regcred
```
## Configure a Corporate Upstream Proxy
Configure a corporate forward proxy when sandbox TLS egress cannot dial the Internet directly. OpenShell evaluates policy and SSRF checks before it opens an HTTP CONNECT tunnel through the proxy. The proxy URL is operator-owned configuration. Sandbox environment variables cannot select, replace, or bypass it.
Create the credential Secret in the sandbox namespace when the proxy requires Basic authentication. The Secret value uses the `user:pass` form.
```shell
kubectl -n openshell create secret generic corporate-proxy-auth \
--from-literal=credentials="$PROXY_USER:$PROXY_PASSWORD"
```
Add the proxy settings to your Helm values file. Replace the DNS suffixes and CIDRs in `noProxy` with values for your cluster. `noProxy` bypasses only the corporate proxy. OpenShell policy evaluation still applies.
```yaml
upstreamProxy:
url: http://proxy.corp.example:8080
noProxy: .svc,.svc.cluster.local,10.96.0.0/12,10.244.0.0/16
authSecret:
name: corporate-proxy-auth
key: credentials
authAllowInsecure: true
supervisor:
topology: sidecar
```
Use `authAllowInsecure: true` only when you accept that Basic authentication is cleartext on the connection to an `http://` proxy. The initial release supports `http://` proxy endpoints and TLS CONNECT egress. It does not support HTTPS-to-proxy, custom corporate CA bundles, or forwarding plain HTTP egress through the proxy.
Proxy credentials require `sidecar` topology. It mounts the credential only into the dedicated network supervisor container. OpenShell rejects credential Secrets with `combined` topology because Kubernetes `fsGroup` volume permission handling can make a shared credential mount readable by the sandbox group.
## RBAC
The chart creates the following RBAC resources in the release namespace:
| Resource | Scope | Name |
|---|---|---|
| ServiceAccount | Namespace | `openshell` |
| ServiceAccount | Namespace | `openshell-sandbox` (for sandbox pods) |
| Role + RoleBinding | Namespace | `openshell-sandbox` |
| ClusterRole + ClusterRoleBinding | Cluster | `openshell-node-reader` |
The namespaced Role covers sandbox lifecycle and identity:
| API Group | Resource | Verbs |
|---|---|---|
| `agents.x-k8s.io` | `sandboxes`, `sandboxes/status` | create, delete, get, list, patch, update, watch |
| `""` | `events` | get, list, watch |
| `""` | `pods` | get |
The ClusterRole grants node inspection and token validation:
| API Group | Resource | Verbs |
|---|---|---|
| `authentication.k8s.io` | `tokenreviews` | create |
| `""` | `nodes` | get, list, watch |
To use an existing ServiceAccount instead of creating one, set `serviceAccount.create=false` and supply its name:
```shell
helm upgrade --install openshell \
oci://ghcr.io/nvidia/openshell/helm-chart \
--version <version> \
--namespace openshell \
--set serviceAccount.create=false \
--set serviceAccount.name=my-existing-sa
```
The ServiceAccount must already have the Role and ClusterRole bindings described above.
## Probes
The gateway exposes `/healthz` for process liveness and `/readyz` for dependency-aware readiness on the health port. The Helm chart wires both into Kubernetes probes:
- `startupProbe` and `livenessProbe` use `/healthz`.
- `readinessProbe` uses `/readyz`, which reflects the latest result of an in-process background database check.
## Next Steps
- To choose between combined and sidecar sandbox pods, refer to [Topology](/kubernetes/topology).
- To enable automatic certificate rotation with cert-manager, refer to [Managing Certificates](/kubernetes/managing-certificates).
- To expose the gateway externally without port-forwarding, refer to [Ingress](/kubernetes/ingress).
- To configure OIDC or reverse-proxy authentication, refer to [Access Control](/kubernetes/access-control).
- To create your first sandbox, refer to [Manage Sandboxes](/sandboxes/manage-sandboxes).
+249
View File
@@ -0,0 +1,249 @@
---
# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
title: "Kubernetes Sandbox Topology"
sidebar-title: "Topology"
description: "Choose between combined and sidecar supervisor topology for Kubernetes sandbox pods."
keywords: "Generative AI, Cybersecurity, Kubernetes, Sandboxing, Sidecar, Network Policy, RuntimeClass"
position: 2
---
Kubernetes sandbox pods can run the OpenShell supervisor in `combined` or
`sidecar` topology. Choose the topology based on which controls you need inside
the pod and how much privilege your cluster allows on the agent container.
## Choose a Topology
The default `combined` topology preserves the full OpenShell enforcement model.
Use `sidecar` only when you accept network-focused enforcement in exchange for a
lower-privilege agent container.
| Topology | Use when | Main tradeoff |
|---|---|---|
| `combined` | You need OpenShell network, filesystem, and process controls in the sandbox workload. | The agent container carries the Linux capabilities the supervisor needs. |
| `sidecar` | You need the agent container to run as non-root without added Linux capabilities, and network policy is the primary control. | Privilege-dropping and supervisor mount isolation do not run in the agent container. |
## Privilege Model
The long-running container permissions differ by topology:
| Topology | Pod or container | UID/GID | Privilege escalation | Capabilities | Result |
|---|---|---|---|---|---|
| `combined` | Agent container, which also runs the supervisor | Not forced by topology | Not explicitly disabled by the driver | Adds `SYS_ADMIN`, `NET_ADMIN`, `SYS_PTRACE`, and `SYSLOG`; adds `SETUID`, `SETGID`, and `DAC_READ_SEARCH` when user namespaces are enabled | Full supervisor controls run in the agent container. |
| `sidecar` | Agent container, process-only supervisor (`network-only`) | `sandbox_uid:sandbox_gid` | `false` | Drops `ALL` | Agent and workload run without added Linux capabilities. |
| `sidecar` | Network supervisor sidecar, binary-aware mode (default) | `0:sandbox_gid` | `false` | Drops `ALL`; adds `SYS_PTRACE` and `DAC_READ_SEARCH` | Root sidecar inspects cross-UID workload `/proc` entries. The nftables fence exempts UID 0, so do not inject other root containers into these pods. |
| `sidecar` | Network supervisor sidecar, endpoint/L7-only mode | `proxyUid:sandbox_gid` | `false` | Drops `ALL` | Non-root sidecar enforces endpoint and L7 policy without matching `policy.binaries`. |
Short-lived setup containers still have the permissions needed to prepare the
pod:
| Topology | Setup container | UID/GID | Privilege escalation | Capabilities | Purpose |
|---|---|---|---|---|---|
| `combined` | Supervisor install init container | `0` | Not set | Not set | Copies the supervisor binary into the agent container volume. |
| `sidecar` | Network init container | `0` | `false` | Drops `ALL`; adds `NET_ADMIN`, `NET_RAW`, `CHOWN`, and `FOWNER` | Installs pod-local nftables rules and prepares shared sidecar state. |
## Combined Topology
Combined topology is the original Kubernetes mode and remains the default. The
agent container starts the OpenShell supervisor, and the supervisor launches the
workload after applying sandbox setup.
```mermaid
flowchart TB
Sandbox["agents.x-k8s.io Sandbox"]
subgraph Pod["Sandbox pod"]
subgraph Agent["agent container"]
Supervisor["OpenShell supervisor<br/>network + process + filesystem"]
Workload["Agent workload"]
end
end
Gateway["OpenShell Gateway"]
External["External services"]
Sandbox --> Pod
Supervisor --> Workload
Supervisor -->|"gateway callback / SSH relay"| Gateway
Supervisor -->|"policy-enforced egress"| External
```
Combined topology keeps these controls in one supervisor path:
- Network endpoint and L7 policy enforcement.
- Filesystem policy enforcement.
- Process and binary identity checks.
- Privilege drop into the sandbox user.
- Gateway relay, SSH sessions, exec, and file sync.
Because the supervisor performs network namespace setup and process/filesystem
controls from the agent container, Kubernetes grants that container elevated
Linux capabilities. Use this mode when you need the complete OpenShell sandbox
contract and your cluster policy permits those capabilities.
## Sidecar Topology
Sidecar topology splits the supervisor into a network sidecar and a
low-privilege process supervisor in the agent container.
```mermaid
flowchart TB
Sandbox["agents.x-k8s.io Sandbox"]
subgraph Pod["Sandbox pod"]
Init["network init container<br/>root setup capabilities"]
State["shared state + TLS volumes"]
NetNS["pod network namespace"]
subgraph Agent["agent container"]
ProcessSupervisor["process supervisor<br/>network-only"]
Workload["Agent workload"]
end
NetworkSidecar["network supervisor sidecar<br/>UID 0 by default"]
SshEndpoint["abstract SSH relay socket<br/>peer-PID authenticated"]
end
Gateway["OpenShell Gateway"]
External["External services"]
Sandbox --> Pod
Init -->|"installs nftables rules"| NetNS
ProcessSupervisor --> Workload
Workload -->|"egress redirected on loopback"| NetworkSidecar
NetworkSidecar -->|"gateway session + relays"| Gateway
NetworkSidecar -->|"policy-enforced egress"| External
NetworkSidecar -->|"control socket + proxy TLS"| State
ProcessSupervisor -->|"bootstrap + updates"| State
ProcessSupervisor --> SshEndpoint
NetworkSidecar -->|"SSH relay"| SshEndpoint
NetworkSidecar --- State
```
The pod contains these OpenShell-managed pieces:
| Component | Runs as | Purpose |
|---|---|---|
| Network init container | Root with setup capabilities | Installs pod-level nftables rules and prepares shared sidecar state. |
| Network sidecar | UID 0 by default; `supervisor.sidecar.proxyUid` when binary-aware policy is disabled | Runs the proxy, enforces network policy, owns gateway authentication and the gateway session, and serves local policy/provider state over the sidecar control socket. |
| Agent container | Resolved sandbox UID/GID | Runs the process supervisor and launches the user workload. |
In this topology, the agent container defaults to `runAsNonRoot: true`,
`allowPrivilegeEscalation: false`, and `capabilities.drop: ["ALL"]`. The
default binary-aware network sidecar runs as UID 0, drops default Linux
capabilities, and adds `SYS_PTRACE` plus `DAC_READ_SEARCH` for cross-UID workload
process identity resolution. Setting
`supervisor.sidecar.processBinaryAwareNetworkPolicy=false` runs the sidecar as
the configured non-root `proxyUid`, omits both capabilities, and downgrades
network policy to endpoint/L7 enforcement without binary matching. The root
init container keeps the setup capabilities needed to configure pod networking.
Sidecar mode preserves gateway session behavior, including SSH connectivity,
because the network sidecar owns the gateway session and bridges relay requests
to a Linux abstract SSH socket owned by the process supervisor. The relay
verifies the socket peer PID against the authenticated control connection, so
the workload cannot replace the relay endpoint. The agent container does not get a
gateway endpoint, gateway TLS material, or the sandbox bootstrap token in the
default sidecar path.
<Warning>
Sidecar mode runs the process supervisor in `network-only` mode. OpenShell still
enforces network endpoint and L7 policy through the sidecar, and the process
supervisor applies Landlock filesystem policy and child seccomp filters where
the kernel/runtime supports them. The process supervisor does not perform
root-to-sandbox privilege dropping because Kubernetes starts the container as
the sandbox UID/GID, and it does not perform supervisor identity mount
isolation because gateway credentials are not mounted into the agent container.
Sidecar pods use `shareProcessNamespace: true` so the network sidecar can
resolve workload process and binary identity through `/proc/<entrypoint-pid>`.
</Warning>
## Credential Exposure
Sidecar topology keeps gateway credentials in the network sidecar. The agent
container does not mount the projected ServiceAccount token used for sandbox
token bootstrap, does not mount the sandbox client TLS secret, and does not get
gateway callback environment variables.
The network sidecar serves the policy and workload-facing provider environment
over a Unix control socket in the shared sidecar state volume. Before launching
the workload, the process supervisor establishes the only accepted connection.
The sidecar validates its UID, GID, and PID with peer credentials, unlinks the
listener, derives the SSH target from trusted configuration, and rejects later
clients. The connection receives bootstrap state and provider-environment
updates after settings polls. If it closes, the network sidecar exits so
Kubernetes recreates the one-client bootstrap listener, and the process
supervisor exits so Kubernetes terminates the workload and restarts the agent
container. This symmetric failure behavior prevents a surviving workload from
claiming the new control listener after an isolated sidecar restart. Future
child processes can see refreshed provider env without giving the agent
container gateway authentication material. This does not mutate the environment
of the already-running workload entrypoint. Use `combined` topology when you
need the full single-supervisor enforcement path; use additional runtime
isolation when you need a stronger container boundary around sidecar workloads.
## RuntimeClass Isolation
Sidecar topology has been validated with Kata Containers. It does not currently
support gVisor because sidecar mode requires pod-local nftables setup, which
gVisor does not provide to the init container. A supported sandboxed runtime
strengthens the container boundary while OpenShell focuses on network policy
enforcement from the sidecar.
Runtime classes do not re-enable the OpenShell privilege-drop or supervisor
mount-isolation controls that sidecar mode relaxes. Use them as an additional
workload boundary, not as a replacement for the combined topology's full
supervisor controls.
You can set a default runtime class in the Kubernetes driver configuration or
override it per sandbox with driver config:
```shell
openshell sandbox create \
--driver-config-json '{"kubernetes":{"pod":{"runtime_class_name":"kata-containers"}}}' \
-- claude
```
## Enable Sidecar Mode
For direct gateway TOML configuration, set the Kubernetes driver fields:
```toml
[openshell.drivers.kubernetes]
topology = "sidecar"
[openshell.drivers.kubernetes.sidecar]
proxy_uid = 1337
```
`proxy_uid` configures only the relaxed endpoint/L7-only sidecar. It must be at
least `1000` and must not match the sandbox UID. The default binary-aware mode
runs the sidecar as UID 0 instead. The network init container exempts the
effective sidecar UID from proxy redirection so the sidecar can reach the
gateway.
When the Helm chart renders `gateway.toml`, set the equivalent chart values:
```yaml
supervisor:
topology: sidecar
sidecar:
proxyUid: 1337
processBinaryAwareNetworkPolicy: true
```
Leave `topology` unset, or set it to `combined`, to keep the original
single-container supervisor path. For Helm installs, leave
`supervisor.topology` unset or set it to `combined`.
Set `supervisor.sidecar.processBinaryAwareNetworkPolicy=false` only when you
accept downgrading sidecar network policy to endpoint/L7 enforcement without
matching `policy.binaries`. This changes the sidecar from UID 0 to `proxyUid`
and removes its `SYS_PTRACE` and `DAC_READ_SEARCH` capabilities, which are used
for cross-UID `/proc` inspection.
## Next Steps
- To install OpenShell on Kubernetes, refer to [Setup](/kubernetes/setup).
- To configure gateway authentication, refer to [Access Control](/kubernetes/access-control).
- To review the driver fields, refer to [Gateway Configuration File](/reference/gateway-config).
@@ -0,0 +1,92 @@
---
# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
title: "Accessing Logs"
description: "How to view sandbox logs through the CLI, TUI, and directly on the sandbox filesystem."
keywords: "Generative AI, Cybersecurity, Logging, CLI, TUI, Observability"
---
OpenShell provides three ways to access sandbox logs: the CLI, the TUI, and direct filesystem access inside the sandbox.
## CLI
Use `openshell logs` to stream logs from a running sandbox:
```shell
openshell logs smoke-l4 --source sandbox
```
The CLI receives logs from the gateway over gRPC. Each line includes a timestamp, source, level, and message:
```text
[1775014132.118] [sandbox] [OCSF ] [ocsf] NET:OPEN [INFO] ALLOWED /usr/bin/curl(58) -> api.github.com:443 [policy:github_api engine:opa]
[1775014132.190] [sandbox] [OCSF ] [ocsf] HTTP:GET [INFO] ALLOWED GET http://api.github.com/zen [policy:github_api]
[1775014132.690] [sandbox] [OCSF ] [ocsf] NET:OPEN [MED] DENIED /usr/bin/curl(64) -> httpbin.org:443 [policy:- engine:opa]
[1775014113.058] [sandbox] [INFO ] [openshell_sandbox] Starting sandbox
```
OCSF structured events show `OCSF` as the level. Standard tracing events show `INFO`, `WARN`, or `ERROR`.
Gateway-originated policy mutations also appear in this stream. When the gateway merges `openshell policy update` operations or approves or removes draft policy chunks, it emits `gateway` `OCSF` `CONFIG:*` lines for the affected sandbox so you can see the exact logical change that produced a new policy revision.
## TUI
The TUI dashboard displays sandbox logs in real time. Logs appear in the log panel with the same format as the CLI.
## Gateway Log Storage
The sandbox pushes logs to the gateway over gRPC in real time. The gateway stores a bounded buffer of recent log lines per sandbox. This buffer is not persisted to disk and is lost when the gateway restarts.
For durable log storage, use the log files inside the sandbox or enable [OCSF JSON export](/observability/ocsf-json-export) and ship the JSONL files to an external log aggregator.
## Direct Filesystem Access
Start an independent shell with `sandbox exec` to read log files directly:
```text
openshell sandbox exec --name my-sandbox --tty -- /bin/bash -l
sandbox@my-sandbox:~$ cat /var/log/openshell.2026-04-01.log
```
Or run a one-off command without an interactive shell:
```shell
openshell sandbox exec --name my-sandbox -- cat /var/log/openshell.2026-04-01.log
```
`sandbox connect` attaches to the sandbox's existing canonical main process; it
does not start a new shell.
The log files inside the sandbox contain the complete record, including events that the gRPC push channel can drop under load. The push channel is bounded and drops events rather than blocking.
## Filtering by Event Type
The shorthand format is designed for `grep`. Some useful patterns:
```shell
# All denied connections
grep "DENIED\|BLOCKED" /var/log/openshell.*.log
# All network events
grep "OCSF NET:" /var/log/openshell.*.log
# All L7 enforcement decisions
grep "OCSF HTTP:" /var/log/openshell.*.log
# Security findings only
grep "OCSF FINDING:" /var/log/openshell.*.log
# Policy changes
grep "OCSF CONFIG:" /var/log/openshell.*.log
# All OCSF events, excluding standard tracing
grep "^.* OCSF " /var/log/openshell.*.log
# Events at medium severity or above
grep "\[MED\]\|\[HIGH\]\|\[CRIT\]\|\[FATAL\]" /var/log/openshell.*.log
```
## Next Steps
- Learn how the [log formats](/observability/logging) work and how to read the shorthand.
- [Enable OCSF JSON export](/observability/ocsf-json-export) for machine-readable structured output.
@@ -0,0 +1,266 @@
---
# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
title: "Sandbox Logging"
sidebar-title: "Logging"
description: "How OpenShell logs sandbox activity using standard tracing and OCSF structured events."
keywords: "Generative AI, Cybersecurity, Logging, OCSF, Observability"
---
Every OpenShell sandbox produces a log that records network connections, process lifecycle events, filesystem policy decisions, and configuration changes. The log uses two formats depending on the type of event.
## Log Formats
### Standard tracing
Internal operational events use Rust's `tracing` framework with a conventional format:
```text
2026-04-01T03:28:39.160Z INFO openshell_sandbox: Fetching sandbox policy via gRPC
2026-04-01T03:28:39.175Z INFO openshell_sandbox: Creating OPA engine from proto policy data
```
These events cover startup plumbing, gRPC communication, and internal state transitions that are useful for debugging but do not represent security-relevant decisions.
### OCSF structured events
Network, process, filesystem, configuration, and API activity events use the [Open Cybersecurity Schema Framework (OCSF)](https://ocsf.io) format. OCSF is an open standard for normalizing security telemetry across tools and platforms. OpenShell maps sandbox events to OCSF v1.8.0 event classes, including API Activity [6003] with the `ai_operation` profile for inference observability.
In the log file, OCSF events appear in a shorthand format with an `OCSF` level label, designed for quick human and agent scanning:
```text
2026-04-01T04:04:13.058Z INFO openshell_sandbox: Starting sandbox
2026-04-01T04:04:13.065Z OCSF CONFIG:DISCOVERY [INFO] Server returned no policy; attempting local discovery
2026-04-01T04:04:13.074Z INFO openshell_sandbox: Creating OPA engine from proto policy data
2026-04-01T04:04:13.078Z OCSF CONFIG:VALIDATED [INFO] Validated 'sandbox' user exists in image
2026-04-01T04:04:32.118Z OCSF NET:OPEN [INFO] ALLOWED /usr/bin/curl(58) -> api.github.com:443 [policy:github_api engine:opa]
2026-04-01T04:04:32.190Z OCSF HTTP:GET [INFO] ALLOWED GET http://api.github.com/zen [policy:github_api engine:opa]
2026-04-01T04:04:32.690Z OCSF NET:OPEN [MED] DENIED /usr/bin/curl(64) -> httpbin.org:443 [policy:- engine:opa] [reason:no matching policy]
```
The `OCSF` label at column 25 distinguishes structured events from standard `INFO` tracing at the same position. Both formats appear in the same file.
When viewed through the CLI or TUI, which receive logs through gRPC, the same distinction applies:
```text
[1775014132.118] [sandbox] [OCSF ] [ocsf] NET:OPEN [INFO] ALLOWED /usr/bin/curl(58) -> api.github.com:443 [policy:github_api engine:opa]
[1775014132.690] [sandbox] [OCSF ] [ocsf] NET:OPEN [MED] DENIED /usr/bin/curl(64) -> httpbin.org:443 [policy:- engine:opa] [reason:no matching policy]
[1775014113.058] [sandbox] [INFO ] [openshell_sandbox] Starting sandbox
```
## OCSF Event Classes
OpenShell maps sandbox events to these OCSF classes:
| Shorthand prefix | OCSF class | Class UID | What it covers |
|---|---|---|---|
| `NET:` | Network Activity | 4001 | TCP proxy CONNECT tunnels, bypass detection, DNS failures |
| `HTTP:` | HTTP Activity | 4002 | HTTP FORWARD requests, L7 enforcement decisions |
| `SSH:` | SSH Activity | 4007 | SSH handshakes, authentication, channel operations |
| `PROC:` | Process Activity | 1007 | Process start, exit, timeout, signal failures |
| `FINDING:` | Detection Finding | 2004 | Security findings (nonce replay, proxy bypass, unsafe policy) |
| `CONFIG:` | Device Config State Change | 5019 | Policy load/reload, Landlock, TLS setup, inference routes |
| `LIFECYCLE:` | Application Lifecycle | 6002 | Sandbox supervisor start, SSH server ready |
| `API:INFERENCE` | API Activity | 6003 | AI model inference calls through `inference.local` (model, provider, latency, tokens) |
## Reading the Shorthand Format
The shorthand format follows this pattern:
```text
CLASS:ACTIVITY [SEVERITY] ACTION DETAILS [CONTEXT]
```
### Components
**Class and activity** (`NET:OPEN`, `HTTP:GET`, `PROC:LAUNCH`) identify the OCSF event class and what happened. The class name always starts at the same column position for vertical scanning.
**Severity** indicates the OCSF severity of the event:
| Tag | Meaning | When used |
|---|---|---|
| `[INFO]` | Informational | Allowed connections, successful operations |
| `[LOW]` | Low | DNS failures, operational warnings |
| `[MED]` | Medium | Denied connections, policy violations |
| `[HIGH]` | High | Security findings (nonce replay, bypass detection) |
| `[CRIT]` | Critical | Process timeout kills |
| `[FATAL]` | Fatal | Unrecoverable failures |
**Action** (`ALLOWED`, `DENIED`, `BLOCKED`) is the security control disposition. Not all events have an action; informational config events, for example, do not.
**Details** vary by event class:
- Network: `process(pid) -> host:port` with the process identity and destination
- HTTP: `METHOD url` with the HTTP method and target
- SSH: peer address and authentication type
- Process: `name(pid)` with exit code or command line
- Config: description of what changed
- Finding: quoted title with the stable finding type, optional confidence, and source-specific context attributes when available
**Context** in brackets provides structured fields such as policy provenance, source-specific attributes, and denial reasons.
### Examples
An allowed HTTPS connection:
```text
OCSF NET:OPEN [INFO] ALLOWED /usr/bin/curl(58) -> api.github.com:443 [policy:github_api engine:opa]
```
An L7 read-only policy denying a POST:
```text
OCSF HTTP:POST [MED] DENIED POST http://api.github.com/user/repos [policy:github_api engine:opa]
```
A connection denied because no policy matched:
```text
OCSF NET:OPEN [MED] DENIED /usr/bin/curl(64) -> httpbin.org:443 [policy:- engine:opa] [reason:no matching policy]
```
A connection denied because the destination resolves to an always-blocked address:
```text
OCSF NET:OPEN [MED] DENIED /usr/bin/curl(1618) -> 169.254.169.254:80 [policy:- engine:ssrf] [reason:resolves to always-blocked address]
```
An HTTP request to a non-default port. HTTP log URLs include the port whenever it differs from the scheme default (80 for `http`, 443 for `https`):
```text
OCSF HTTP:GET [INFO] ALLOWED GET http://api.internal.corp:8080/v1/status [policy:internal_api engine:opa]
```
A supervisor middleware HTTP event records whether it transformed the request. If the middleware also emits a finding, that remains a separate event:
```text
OCSF HTTP:POST [INFO] ALLOWED POST http://httpbin.org:443/anything [policy:httpbin engine:middleware] [failed:false transformed:true]
OCSF FINDING:CREATE [MED] "configured content matched" [type:content_guard.match count:1 middleware:prototype-content-guard]
```
WebSocket middleware emits one safe event per preflight, session-start, or client text-message decision. The event includes the policy-local config, registered implementation, sequence, byte counts, transformation flag, and validated reason code. It never includes the message payload or service-provided free-form reason:
```text
OCSF NET:OTHER [INFO] WEBSOCKET_MIDDLEWARE allow config=api-redactor implementation=openshell/regex sequence=3 input_bytes=128 replacement_bytes=96 transformed=true reason_code=-
```
Coverage events are separate from invocation decisions. `binding_not_selected` means a host-matched attachment did not advertise the WebSocket operation. `unsupported_message_type` means an active WebSocket stage encountered a binary message, which V1 passes through without inspection. Both are informational, do not apply `on_error`, and never claim the traffic was inspected:
```text
OCSF NET:OTHER [INFO] WEBSOCKET_MIDDLEWARE_COVERAGE state=binding_not_selected config=http-dlp implementation=example/http-dlp sequence=- message_type=- input_bytes=0
OCSF NET:OTHER [INFO] WEBSOCKET_MIDDLEWARE_COVERAGE state=unsupported_message_type config=api-redactor implementation=openshell/regex sequence=4 message_type=binary input_bytes=512
```
A fail-open stream error emits both a middleware failure and `openshell.middleware.websocket_stage_disabled`. The latter records that OpenShell will bypass that stage for later messages on the same connection. Waiting for saturated admission capacity also emits a detection finding without payload content. When both active capacity and the bounded wait queue are full, HTTP work is rejected before its payload is buffered, returns `503 Service Unavailable`, and emits `openshell.middleware.admission_exhausted`.
Proxy and SSH servers ready:
```text
OCSF NET:LISTEN [INFO] 10.200.0.1:3128
OCSF SSH:LISTEN [INFO]
```
An SSH connection accepted (one event per invocation, arriving over the supervisor's Unix socket, so there is no network peer address to log):
```text
OCSF SSH:OPEN [INFO] ALLOWED
```
A process launched inside the sandbox:
```text
OCSF PROC:LAUNCH [INFO] sleep(49)
```
A policy reload after a settings change:
```text
OCSF CONFIG:DETECTED [INFO] Settings poll: config change detected [old_revision:2915564174587774909 new_revision:11008534403127604466 policy_changed:true]
OCSF CONFIG:LOADED [INFO] Policy reloaded successfully [policy_hash:0cc0c2b525573c07]
```
## Denial Reasons
Denied `NET:` and `HTTP:` events carry a `[reason:...]` suffix that surfaces the decision detail from the event's `status_detail` field. The reason helps distinguish between policy misses, SSRF hardening, and L7 enforcement without inspecting the full OCSF JSONL record.
For supervisor middleware denials, `status_detail` contains a platform-owned reason derived from the policy-local middleware config name and optional validated reason code. Middleware failure details also use platform-owned error codes. OpenShell does not copy per-request service text or WebSocket message content into logs.
Common reason phrases emitted by the sandbox include:
| Reason | Meaning |
|---|---|
| `no matching policy` | OPA evaluated the request and no allow rule matched. |
| `resolves to always-blocked address` | The destination resolved to loopback, link-local, or unspecified. These ranges are always blocked, even when listed in `allowed_ips`. |
| `resolves to <ip> which is not in allowed_ips, connection rejected` | The destination resolved to an IP outside the policy's `allowed_ips` allowlist. |
| `DNS resolution failed for <host>:<port>` | The proxy could not resolve the destination. |
| `port <n> is a blocked control-plane port, connection rejected` | The destination port matches a control-plane port (etcd, Kubernetes API, kubelet) and is always blocked. |
| `request-target contains an encoded '/' (%2F)` | The L7 HTTP parser rejected an encoded slash. Configure `allow_encoded_slash: true` on a REST endpoint when the upstream requires encoded slashes. |
| `l7 deny` | An L7 policy rule denied the request. |
Invalid `allowed_ips` entries and entries that overlap always-blocked ranges are rejected at policy-load time, so they never reach the runtime denial path. The phrases above come from the proxy's per-CONNECT `allowed_ips` and SSRF checks, not from policy validation.
## Proxy Error Responses
When the HTTP CONNECT proxy denies a request or cannot reach the upstream, it returns an HTTP error response with a JSON body. Clients can parse the body to surface actionable failure details instead of treating the status code alone.
A denied CONNECT returns `403 Forbidden`:
```json
{
"error": "policy_denied",
"detail": "CONNECT api.example.com:443 not permitted by policy",
"reason": "binary '/usr/bin/node' not allowed in policy 'allow_api' (ancestors: [/usr/local/bin/claude])"
}
```
The `reason` field is included when the policy engine provides a specific denial reason (for example, which binary or rule caused the rejection). It is omitted when no additional detail is available.
An upstream that the proxy cannot reach returns `502 Bad Gateway`:
```json
{
"error": "upstream_unreachable",
"detail": "connection to api.example.com:443 failed"
}
```
The `error` field is a short machine-readable code (`policy_denied`, `middleware_denied`, `middleware_failed`, `ssrf_denied`, `upstream_unreachable`). The `detail` field is a human-readable explanation suitable for display in an agent transcript. The optional `reason` field, when present, provides the specific denial cause from the policy engine (for example, which binary was not allowed or which rule was missing).
For L7 REST policy denials, the body also includes structured policy fields such as `method`, `path`, `rule_missing`, and `next_steps`. When policy advisor is enabled, it also includes `agent_guidance`, a short plain-language instruction telling the agent to read `/etc/openshell/skills/policy_advisor.md`, propose the narrowest rule through `http://policy.local/v1/proposals`, wait for `policy_reloaded: true`, and retry. A middleware denial instead identifies the policy-local config in `middleware` and can include a validated `reason_code`. A fail-closed runtime failure uses `middleware_failed` with platform-owned text. Both middleware responses omit `rule_missing`, `next_steps`, and `agent_guidance` because no policy rule is missing.
## Filesystem Sandbox Logs
Landlock filesystem restrictions emit `CONFIG:` events at startup and whenever the sandbox has to skip a requested path.
On startup, the probe reports the kernel's supported Landlock ABI version alongside the requested path counts:
```text
OCSF CONFIG:ENABLED [INFO] Landlock filesystem sandbox available [abi:v2 compat:BestEffort ro:4 rw:2]
OCSF CONFIG:ENABLED [INFO] Applying Landlock filesystem sandbox [abi:V2 compat:BestEffort ro:4 rw:2]
OCSF CONFIG:ENABLED [INFO] Landlock ruleset built [rules_applied:5 skipped:1]
```
When `landlock.compatibility` is `best_effort` and a requested path fails to open for reasons other than `NotFound` (for example, permission denied or a symlink loop), the sandbox continues without that path and emits a `[MED]` event so the degradation is not silent:
```text
OCSF CONFIG:OTHER [MED] Skipping inaccessible Landlock path (best-effort) [path:/opt/data error:Permission denied (os error 13)]
```
Set `landlock.compatibility` to `hard_requirement` in the policy to make these failures fatal instead of degraded.
## Log File Location
Inside the sandbox, logs are written to `/var/log/`:
| File | Format | Rotation |
|---|---|---|
| `openshell.YYYY-MM-DD.log` | Shorthand + standard tracing | Daily, 3 files max |
| `openshell-ocsf.YYYY-MM-DD.log` | OCSF JSONL when enabled | Daily, 3 files max |
Both files rotate daily and retain the 3 most recent files to bound disk usage.
## Next Steps
- [Access logs](/observability/accessing-logs) through the CLI, TUI, or sandbox filesystem.
- [Enable OCSF JSON export](/observability/ocsf-json-export) for SIEM integration and compliance.
- Learn about [network policies](/sandboxes/policies) that generate these events.
@@ -0,0 +1,198 @@
---
# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
title: "OCSF JSON Export"
sidebar-title: "OCSF JSON Export"
description: "How to enable full OCSF JSON logging for SIEM integration, compliance, and structured analysis."
keywords: "Generative AI, Cybersecurity, OCSF, JSON, SIEM, Compliance, Observability"
---
The [shorthand log format](/observability/logging) is optimized for humans and agents reading logs in real time. For machine consumption, compliance archival, or SIEM integration, you can enable full OCSF JSON export. This writes every OCSF event as a complete JSON record in JSONL format, one JSON object per line.
## Enable JSON Export
Use the `ocsf_json_enabled` setting to toggle JSON export. The setting can be applied globally, for all sandboxes, or per-sandbox.
Global:
```shell
openshell settings set --global --key ocsf_json_enabled --value true
```
Per-sandbox:
```shell
openshell settings set my-sandbox --key ocsf_json_enabled --value true
```
The setting takes effect on the next poll cycle, by default every 10 seconds. No sandbox restart is required.
To disable:
```shell
openshell settings set --global --key ocsf_json_enabled --value false
```
## Output Location
When enabled, OCSF JSON records are written to `/var/log/openshell-ocsf.YYYY-MM-DD.log` inside the sandbox. The file rotates daily and retains the 3 most recent files, matching the main log file rotation.
## JSON Record Structure
Each line is a complete OCSF v1.8.0 JSON object. Here is an example of a network connection event:
```json
{
"class_uid": 4001,
"class_name": "Network Activity",
"category_uid": 4,
"category_name": "Network Activity",
"activity_id": 1,
"activity_name": "Open",
"severity_id": 1,
"severity": "Informational",
"status_id": 1,
"status": "Success",
"time": 1775014138811,
"message": "CONNECT allowed api.github.com:443",
"metadata": {
"product": {
"name": "OpenShell Sandbox Supervisor",
"vendor_name": "NVIDIA",
"version": "0.3.0"
},
"version": "1.8.0"
},
"action_id": 1,
"action": "Allowed",
"disposition_id": 1,
"disposition": "Allowed",
"dst_endpoint": {
"domain": "api.github.com",
"port": 443
},
"src_endpoint": {
"ip": "10.42.0.31",
"port": 37494
},
"actor": {
"process": {
"name": "/usr/bin/curl",
"pid": 57
}
},
"firewall_rule": {
"name": "github_api",
"type": "opa"
}
}
```
And a denied connection:
```json
{
"class_uid": 4001,
"class_name": "Network Activity",
"activity_id": 1,
"activity_name": "Open",
"severity_id": 3,
"severity": "Medium",
"status_id": 2,
"status": "Failure",
"action_id": 2,
"action": "Denied",
"disposition_id": 2,
"disposition": "Blocked",
"status_detail": "no matching policy",
"message": "CONNECT denied httpbin.org:443",
"dst_endpoint": {
"domain": "httpbin.org",
"port": 443
},
"actor": {
"process": {
"name": "/usr/bin/curl",
"pid": 63
}
},
"firewall_rule": {
"name": "-",
"type": "opa"
}
}
```
<Note>
The JSON examples above are formatted for readability. The actual JSONL file contains one JSON object per line with no whitespace formatting.
</Note>
## OCSF Event Classes in JSON
The `class_uid` field identifies the event type:
| `class_uid` | Class | Shorthand prefix |
|---|---|---|
| 4001 | Network Activity | `NET:` |
| 4002 | HTTP Activity | `HTTP:` |
| 4007 | SSH Activity | `SSH:` |
| 1007 | Process Activity | `PROC:` |
| 2004 | Detection Finding | `FINDING:` |
| 5019 | Device Config State Change | `CONFIG:` |
| 6002 | Application Lifecycle | `LIFECYCLE:` |
| 6003 | API Activity | `API:INFERENCE` |
### API Activity [6003] — AI Inference
When the inference proxy routes a model call through `inference.local`, an API Activity event is emitted with the `ai_operation` profile:
```json
{
"class_uid": 6003,
"class_name": "API Activity",
"activity_id": 99,
"activity_name": "Other",
"api": { "operation": "POST /v1/messages" },
"actor": { "process": { "name": "openshell-supervisor", "pid": 1 } },
"src_endpoint": { "ip": "127.0.0.1", "port": 3128 },
"ai_model": { "name": "claude-haiku-4-5-20251001", "ai_provider": "https://api.anthropic.com/v1" },
"metadata": { "version": "1.8.0", "profiles": ["container", "host", "ai_operation"] },
"status": "Success",
"unmapped": { "latency_ms": 701 }
}
```
The shorthand log renders this as:
```text
OCSF API:INFERENCE [INFO] Success claude-haiku-4-5-20251001 via https://api.anthropic.com/v1 701ms [POST /v1/messages]
```
## Integration with External Tools
The JSONL file can be shipped to any tool that accepts OCSF-formatted data:
| Tool | Integration Path |
|---|---|
| Splunk | Use the [Splunk OCSF Add-on](https://splunkbase.splunk.com/app/6943) to ingest OCSF JSONL files. |
| Amazon Security Lake | OCSF is the native schema for Security Lake. |
| Elastic | Use Filebeat to ship JSONL files with the OCSF field mappings. |
| Custom pipelines | Parse the JSONL file with `jq`, Python, or any JSON-capable tool. |
Example with `jq` to extract all denied connections:
```shell
cat /var/log/openshell-ocsf.2026-04-01.log | \
jq -c 'select(.action == "Denied")'
```
## Relationship to Shorthand Logs
The shorthand format in `openshell.YYYY-MM-DD.log` and the JSON format in `openshell-ocsf.YYYY-MM-DD.log` are derived from the same OCSF events. The shorthand is a human-readable projection; the JSON is the complete record. Both are generated at the same time from the same event data.
The shorthand log is always active. The JSON export is opt-in through `ocsf_json_enabled`.
## Next Steps
- Learn how to [read the shorthand format](/observability/logging) for real-time monitoring.
- Refer to the [OCSF specification](https://schema.ocsf.io/) for the full schema reference.
+173
View File
@@ -0,0 +1,173 @@
---
# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
title: "AWS SigV4 Credential Signing"
sidebar-title: "AWS SigV4"
description: "Configure proxy-side AWS SigV4 request signing so sandbox agents can reach AWS services through CONNECT tunnels without holding real credentials."
keywords: "Generative AI, Cybersecurity, AI Agents, AWS, SigV4, Bedrock, S3, Credential Signing, Sandbox"
---
AWS SigV4 credential signing lets sandbox agents call AWS services (Bedrock, S3, STS, and others) through the proxy's CONNECT tunnel. The proxy intercepts outbound requests, strips the sandbox client's placeholder `Authorization` header, and re-signs the request with real AWS credentials from the provider. The sandbox never sees the real credentials.
## Prerequisites
- A provider with `AWS_ACCESS_KEY_ID` and `AWS_SECRET_ACCESS_KEY` credentials configured. Optionally include `AWS_SESSION_TOKEN` for STS temporary credentials.
- An endpointless `aws` provider profile when the sandbox policy defines the service endpoints, or an endpoint-bearing service profile such as `aws-s3`.
- A sandbox policy with `credential_signing` enabled on the target endpoint. For the endpointless `aws` profile, the endpoint must also set `credential_binding.provider` to the attached provider name.
## Provider Setup
Create a provider with AWS credentials:
```shell
openshell provider create \
--name aws-prod \
--type aws \
--credential AWS_ACCESS_KEY_ID=AKIA... \
--credential AWS_SECRET_ACCESS_KEY=wJalr...
```
For STS temporary credentials, include the session token:
```shell
openshell provider create \
--name aws-sts \
--type aws \
--credential AWS_ACCESS_KEY_ID=ASIA... \
--credential AWS_SECRET_ACCESS_KEY=secret... \
--credential AWS_SESSION_TOKEN=FwoGZX...
```
To have the gateway mint and rotate STS credentials for you instead of supplying
them statically, use the `aws` or `aws-s3` profile with the `aws_sts_assume_role`
refresh strategy. A single `sts:AssumeRole` mints all three env vars
(`AWS_ACCESS_KEY_ID`, `AWS_SECRET_ACCESS_KEY`, `AWS_SESSION_TOKEN`) that the
signer reads. See [Manage Providers](/sandboxes/manage-providers#aws-sts).
## Policy Configuration
Enable SigV4 signing on a per-endpoint basis. The binding selects which
provider instance supplies credentials, while the signing fields control how
the proxy applies them:
| Field | Type | Required | Description |
|---|---|---|---|
| `credential_binding.provider` | string | For endpointless profiles | Exact name of the attached provider instance that supplies AWS credentials. |
| `credential_signing` | string | Yes | Signing mode: `sigv4`, `sigv4:body`, or `sigv4:no_body`. |
| `signing_service` | string | Yes | AWS service name for the SigV4 signature (e.g. `bedrock`, `s3`, `sts`). |
| `signing_region` | string | No | AWS region override. When omitted, extracted from the endpoint hostname. Required for non-standard endpoints. |
### Bedrock Example
```yaml
network_policies:
aws_bedrock:
endpoints:
- host: bedrock-runtime.us-east-1.amazonaws.com
port: 443
protocol: rest
credential_binding:
provider: aws-prod
credential_signing: sigv4
signing_service: bedrock
rules:
- allow:
method: POST
path: /model/*/invoke
```
<Note>
The Bedrock example uses `rules` for fine-grained access control. When `rules` are present, omit the `access` field — they are mutually exclusive.
</Note>
### S3 Example
```yaml
network_policies:
aws_s3:
endpoints:
- host: "*.s3.us-east-1.amazonaws.com"
port: 443
protocol: rest
access: full
credential_binding:
provider: aws-prod
credential_signing: sigv4
signing_service: s3
```
### STS Example
```yaml
network_policies:
aws_sts:
endpoints:
- host: sts.us-east-1.amazonaws.com
port: 443
protocol: rest
access: full
credential_binding:
provider: aws-prod
credential_signing: sigv4
signing_service: sts
```
## Signing Modes
The `credential_signing` field accepts three values:
| Value | Behavior | Use When |
|---|---|---|
| `sigv4` | Auto-detect payload mode from the client SDK's `x-amz-content-sha256` header. | Default. Works for most AWS services. |
| `sigv4:body` | Always buffer the request body and include its SHA-256 hash in the signature. Maximum body size: 10 MiB. | Services that require body signing (Bedrock). |
| `sigv4:no_body` | Sign headers only with `UNSIGNED-PAYLOAD`. Stream the body through without buffering. | Large uploads (S3 PutObject), chunked transfers, or any case where body buffering is impractical. |
In `sigv4` auto-detect mode, the proxy inspects the `x-amz-content-sha256` header sent by the client SDK:
- Hex hash → buffer body and sign it (same as `sigv4:body`).
- `UNSIGNED-PAYLOAD` → sign headers only (same as `sigv4:no_body`).
- `STREAMING-UNSIGNED-PAYLOAD-TRAILER` → sign headers only, stream body through.
- Absent → sign body if `Content-Length` is present, otherwise use unsigned payload.
<Note>
Chunk-signed streaming modes like `STREAMING-AWS4-HMAC-SHA256-PAYLOAD` are not supported. The proxy cannot reproduce per-chunk signatures. If your client SDK sends chunk-signed requests, use `sigv4:no_body` instead.
</Note>
## Region Detection
The proxy extracts the AWS region from the endpoint hostname automatically. It supports standard, dualstack, FIPS, virtual-hosted, GovCloud, and China partition hostnames.
For endpoints where the region cannot be inferred from the hostname, set `signing_region` explicitly:
```yaml
endpoints:
- host: custom-vpc-endpoint.example.com
port: 443
protocol: rest
access: full
credential_binding:
provider: aws-prod
credential_signing: sigv4
signing_service: s3
signing_region: us-west-2
```
## Restrictions
- `credential_signing` and `request_body_credential_rewrite` are mutually exclusive on the same endpoint. The policy validator rejects policies that set both.
- `credential_binding.provider` must name a provider attached to that sandbox. Use it only when the selected provider profile has no endpoints. Endpoint-bearing profiles already define their credential boundary.
- OpenShell rejects a signed sandbox policy before activation unless the endpoint has a resolvable AWS credential source. An endpoint-bearing profile must declare `AWS_ACCESS_KEY_ID` and `AWS_SECRET_ACCESS_KEY` and cover the signed host, port, and path. An endpointless profile must declare those keys and be selected with `credential_binding.provider` on the signed endpoint.
- The `sigv4:body` mode buffers at most 10 MiB. Requests with larger bodies are rejected. Use `sigv4:no_body` or `sigv4` (auto-detect) for large payloads.
- The active provider must contain current `AWS_ACCESS_KEY_ID` and `AWS_SECRET_ACCESS_KEY` values. If either becomes unavailable after policy activation, the request fails closed.
## Use from a Sandbox
Inside a sandbox, configure the AWS SDK with placeholder credentials. The proxy replaces them with real credentials during re-signing:
```shell
export AWS_ACCESS_KEY_ID=placeholder
export AWS_SECRET_ACCESS_KEY=placeholder
export AWS_DEFAULT_REGION=us-east-1
```
Then use any AWS SDK or CLI normally. The proxy transparently re-signs requests before forwarding to AWS.
@@ -0,0 +1,194 @@
---
# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
title: "Google Cloud"
sidebar-title: "Google Cloud"
description: "Authenticate with GCP APIs inside OpenShell sandboxes."
keywords: "Generative AI, Google Cloud, Vertex AI, GCP, OAuth2, Credentials, Sandbox"
---
The `google-cloud` provider gives sandboxes native GCP credentials so
any Google Cloud SDK works out of the box — Cloud Storage, BigQuery, Drive,
Maps, Discovery Engine, or any other GCP API. A GCE metadata server emulator
on loopback provides credential placeholders that the
sandbox proxy resolves to real tokens at request time. The sandbox process
never holds a real GCP credential.
## Quick Start
If you already have `gcloud` configured with Application Default Credentials,
create a provider with automatic credential refresh in one command:
```shell
openshell provider create \
--name my-gcp \
--type google-cloud \
--from-gcloud-adc \
--config project_id="$(gcloud config get-value project)" \
--config region=global
```
`--from-gcloud-adc` reads your ADC file, configures OAuth2 refresh on the
gateway, and mints the first access token before the command returns. The
gateway rotates the token automatically — no manual refresh needed.
## Authentication Flows
Two credential flows are supported. Choose based on your environment.
### Application Default Credentials (gcloud ADC)
Use credentials from `gcloud auth application-default login`. The gateway
exchanges the refresh token for short-lived access tokens automatically.
```shell
openshell provider create \
--name my-gcp \
--type google-cloud \
--config project_id=my-project \
--config region=us-central1 \
--credential GCP_ADC_ACCESS_TOKEN=placeholder
```
Configure credential refresh with the ADC JSON fields:
```shell
openshell provider refresh configure my-gcp \
--credential-key GCP_ADC_ACCESS_TOKEN \
--strategy oauth2-refresh-token \
--material client_id=YOUR_CLIENT_ID \
--material client_secret=YOUR_CLIENT_SECRET \
--material refresh_token=YOUR_REFRESH_TOKEN \
--secret-material-key client_secret \
--secret-material-key refresh_token
```
Find these values in your ADC file at
`~/.config/gcloud/application_default_credentials.json`.
Trigger the first token mint:
```shell
openshell provider refresh rotate my-gcp \
--credential-key GCP_ADC_ACCESS_TOKEN
```
### Service Account Key
Use a GCP service account JSON key file. The gateway signs JWTs and
exchanges them for access tokens using the `google-service-account-jwt`
strategy.
```shell
openshell provider create \
--name my-gcp \
--type google-cloud \
--config project_id=my-project \
--config region=us-central1 \
--credential GCP_SA_ACCESS_TOKEN=placeholder
```
```shell
openshell provider refresh configure my-gcp \
--credential-key GCP_SA_ACCESS_TOKEN \
--strategy google-service-account-jwt \
--material client_email=sa@my-project.iam.gserviceaccount.com \
--material private_key="$(jq -r .private_key /path/to/sa-key.json)" \
--secret-material-key private_key
```
```shell
openshell provider refresh rotate my-gcp \
--credential-key GCP_SA_ACCESS_TOKEN
```
## Configuration Keys
Set these with `--config key=value` during provider creation:
| Key | Description | Example |
|-----|-------------|---------|
| `project_id` | GCP project ID | `my-project-123` |
| `region` | GCP region | `us-central1` |
| `service_account_email` | SA email for metadata endpoint | `sa@proj.iam.gserviceaccount.com` |
## How It Works
When a sandbox starts with the `google-cloud` provider attached:
1. The gateway mints a fresh GCP access token and stores it in the
sandbox proxy's credential resolver.
2. A loopback HTTP server on `127.0.0.1:8174` emulates the GCE instance
metadata API, serving **credential placeholders** (not real tokens) to
GCP SDKs. The sandbox process never holds a real GCP credential.
3. When the SDK makes an API call, it sends the placeholder in the
`Authorization` header. The sandbox proxy TLS-terminates the
outbound connection, resolves the placeholder to the real token,
and forwards the request to GCP.
4. When the token approaches expiry, the gateway refreshes it. The
proxy's resolver is updated atomically — subsequent API calls
use the new token automatically.
Configuration values (`project_id`, `region`, `service_account_email`)
are **visible in plain text** inside the sandbox — they appear as
environment variables and are served by the metadata endpoint. These are
non-secret identifiers, not credentials. Access tokens are never exposed;
only placeholders reach the sandbox process.
### Injected Environment Variables
The provider automatically injects these into the sandbox. Non-secret
vars are resolved to real values at process spawn time; token vars stay
as placeholders for proxy-time resolution.
| Variable | Value | Purpose |
|----------|-------|---------|
| `GCE_METADATA_HOST` | `127.0.0.1:8174` | GCP SDK metadata discovery (loopback server) |
| `GCE_METADATA_IP` | `127.0.0.1:8174` | Python google-auth ping detection |
| `METADATA_SERVER_DETECTION` | `assume-present` | Node.js gcp-metadata skip detection |
| `GCP_PROJECT_ID` | from `project_id` config | GCP SDK project |
| `GOOGLE_CLOUD_PROJECT` | from `project_id` config | Alternative project var |
| `CLOUD_ML_REGION` | from `region` config | GCP region |
| `GCP_LOCATION` | from `region` config | Alternative region var |
## Using with GCP APIs
The metadata emulator serves tokens with the `cloud-platform` OAuth2 scope,
which grants access to any GCP API the underlying service account has IAM
permissions for. Add the target API hosts to your sandbox network policy:
```yaml
network_policies:
gcp_apis:
name: gcp-apis
endpoints:
- host: "*.googleapis.com"
port: 443
protocol: rest
access: read-write
enforcement: enforce
binaries:
- { path: /usr/bin/curl }
- { path: /usr/bin/node }
- { path: "/sandbox/.uv/python/**" }
- { path: "/sandbox/.venv/**" }
```
Or update a live sandbox directly:
```shell
openshell policy update my-sandbox \
--add-endpoint "*.googleapis.com:443:read-write:rest:enforce" \
--binary "/usr/bin/curl" \
--binary "/usr/bin/node" \
--binary "/sandbox/.uv/python/**" \
--binary "/sandbox/.venv/**" \
--wait
```
## Network Policy
The `google-cloud` provider type does not include any network policy
endpoints by default. You must add endpoint rules to your sandbox policy
for each GCP API the sandbox needs to reach. See "Using with GCP APIs"
above for an example.
@@ -0,0 +1,224 @@
---
# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
title: "Google Vertex AI"
sidebar-title: "Google Vertex AI"
description: "Configure OpenShell to route inference traffic through Google Vertex AI, including Anthropic Claude and Gemini models."
keywords: "Generative AI, Cybersecurity, AI Agents, Sandboxing, Google Vertex AI, Anthropic Claude, Inference Routing"
---
Google Vertex AI is a managed machine learning platform that hosts Anthropic Claude, Gemini, and third-party models through Google Cloud. OpenShell can route `inference.local` traffic to Vertex AI using gateway-managed credential refresh, so sandbox agents do not handle GCP credentials directly.
## Prerequisites
Before creating a Vertex AI provider, ensure you have:
- A GCP project with the [Vertex AI API](https://console.cloud.google.com/apis/library/aiplatform.googleapis.com) enabled.
- One of the following:
- A GCP service account with the **Vertex AI User** role and a downloaded JSON key file, for production use.
- The `gcloud` CLI with Application Default Credentials configured, for local development.
## Authentication
The `google-vertex-ai` provider supports two credential sources.
### Service Account Key
Supply the JSON key file content as the `GOOGLE_SERVICE_ACCOUNT_KEY` credential. OpenShell persists that value only as gateway-side refresh bootstrap material until you update or delete it. The raw service-account JSON and private key are not sandbox runtime credentials and are not exposed to sandboxes. Runtime inference requests use short-lived access tokens minted by the gateway and stored under a separate credential key.
```shell
openshell provider create \
--name vertex-prod \
--type google-vertex-ai \
--credential GOOGLE_SERVICE_ACCOUNT_KEY="$(cat /path/to/key.json)" \
--config VERTEX_AI_PROJECT_ID=my-gcp-project \
--config VERTEX_AI_REGION=us-central1
```
Then configure gateway-managed refresh so the gateway uses the private key as refresh bootstrap material and rotates access tokens:
```shell
openshell provider refresh configure vertex-prod \
--credential-key GOOGLE_VERTEX_AI_SERVICE_ACCOUNT_TOKEN \
--strategy google-service-account-jwt \
--material client_email="sa@my-gcp-project.iam.gserviceaccount.com" \
--material private_key="$(jq -r .private_key /path/to/key.json)" \
--secret-material-key private_key
```
### gcloud Application Default Credentials
For local development, configure ADC first, then pass `--from-gcloud-adc`:
```shell
gcloud auth application-default login
```
```shell
openshell provider create \
--name vertex-local \
--type google-vertex-ai \
--from-gcloud-adc \
--config VERTEX_AI_PROJECT_ID=my-gcp-project \
--config VERTEX_AI_REGION=us-central1
```
`--from-gcloud-adc` reads `GOOGLE_APPLICATION_CREDENTIALS` first, then falls back to `$CLOUDSDK_CONFIG/application_default_credentials.json` when that environment variable is set, then to `~/.config/gcloud/application_default_credentials.json`. It configures an OAuth2 refresh token flow on the gateway and immediately mints the first access token before the command returns. If the command succeeds, the provider is ready for inference right away. It only works with user credentials generated by `gcloud auth application-default login`. If your ADC file is a service account key, the CLI returns an error and directs you to use the service account key method above.
ADC-backed providers mint and rotate access tokens into `GOOGLE_VERTEX_AI_TOKEN`.
<Note>
`--from-gcloud-adc` is valid for `google-vertex-ai` and `google-cloud` providers.
</Note>
## Configuration Keys
Pass these as `--config KEY=VALUE` when creating the provider, or set them as environment variables and use `--from-existing`.
| Key | Required | Default | Description |
|---|---|---|---|
| `VERTEX_AI_PROJECT_ID` | Yes (unless `GOOGLE_VERTEX_AI_BASE_URL` or `VERTEX_AI_BASE_URL` is set) | — | GCP project ID. |
| `VERTEX_AI_REGION` | No | `us-central1` | Vertex location selector. Use a regional location such as `us-central1`, or `global`, `us`, or `eu` for the supported global and multi-region endpoints. |
| `GOOGLE_VERTEX_AI_BASE_URL` | No | — | Full base URL override for non-Anthropic routes. Must be an official Vertex AI HTTPS endpoint root. |
| `VERTEX_AI_BASE_URL` | No | — | Backward-compatible alias for `GOOGLE_VERTEX_AI_BASE_URL`. |
| `VERTEX_AI_PUBLISHER` | No | Inferred from model name | Set to `anthropic` to force Anthropic Messages API routing, or any other value for OpenAI-compatible routing. |
When `VERTEX_AI_PROJECT_ID` is set and no base URL override is present, the gateway maps `VERTEX_AI_REGION` to the Vertex host automatically:
- Regional locations such as `us-central1` use `https://<region>-aiplatform.googleapis.com`.
- `global` uses `https://aiplatform.googleapis.com`.
- `us` and `eu` use `https://aiplatform.<region>.rep.googleapis.com`.
For Anthropic models, OpenShell builds the publisher-model Vertex path automatically and injects `anthropic_version` into the request body. Vertex rawPredict does not receive `anthropic-version` as a header, and OpenShell strips `anthropic-beta` for Vertex Claude routes. For non-Anthropic models, OpenShell uses Vertex's OpenAI-compatible Chat Completions route under `.../endpoints/openapi/chat/completions`.
<Note>
Use `GOOGLE_VERTEX_AI_BASE_URL` or `VERTEX_AI_BASE_URL` only for non-Anthropic Vertex routes. OpenShell rejects Anthropic models when a base URL override is set because Anthropic routes require model-path shaping and `anthropic_version` body injection. Overrides must use `https://` and an official Vertex AI hostname such as `aiplatform.googleapis.com`, `aiplatform.us.rep.googleapis.com`, `aiplatform.eu.rep.googleapis.com`, or `<region>-aiplatform.googleapis.com`.
</Note>
## Supported Models
Vertex AI hosts Anthropic Claude models (claude-3-5-sonnet, claude-3-opus, and others) through a native Messages API integration, and Gemini and other third-party models through Vertex's OpenAI-compatible Chat Completions endpoint. OpenShell infers the routing path from the model name. For the full list of available models and regions, refer to the [Google Cloud model garden documentation](https://cloud.google.com/vertex-ai/generative-ai/docs/model-reference/overview).
Model names that match the `claude-*` pattern route through the Anthropic Messages API on Vertex. All other model names route through Vertex Chat Completions. Set `VERTEX_AI_PUBLISHER=anthropic` to force Anthropic routing when the model name does not follow the standard pattern.
OpenShell exposes Anthropic Vertex routes for inference only. It does not advertise OpenAI-style model discovery for those routes, so use the Google Cloud docs or Model Garden to discover supported Anthropic model IDs.
## Configure Inference Routing
Before configuring inference routing, enable provider endpoint injection so the Vertex AI network endpoints are automatically included in sandbox policies:
```shell
openshell settings set --global --key providers_v2_enabled --value true --yes
```
Then point `inference.local` at the provider:
```shell
openshell inference set \
--provider vertex-prod \
--model claude-sonnet-4-6
```
Use `--no-verify` if the endpoint verification fails. This is common with the `global` region, where the validation probe may not match the actual rawPredict path:
```shell
openshell inference set \
--provider vertex-prod \
--model claude-sonnet-4-6 \
--no-verify
```
Sandboxes on that gateway reach the model at `https://inference.local`. For full details on inference routing, refer to [Inference Routing](/sandboxes/inference-routing).
## Use from a Sandbox
Agents inside sandboxes should reach Vertex AI through `inference.local`, not by connecting to Vertex AI directly. The gateway manages GCP credential refresh and request translation; the agent only needs to point its SDK at the local endpoint.
The complete setup from scratch:
```shell
# 1. Enable provider endpoint injection
openshell settings set --global --key providers_v2_enabled --value true --yes
# 2. Create the provider
openshell provider create \
--name vertex-local \
--type google-vertex-ai \
--from-gcloud-adc \
--config VERTEX_AI_PROJECT_ID=my-gcp-project \
--config VERTEX_AI_REGION=us-central1
# 3. Configure inference routing
openshell inference set --provider vertex-local --model claude-sonnet-4-6 --no-verify
# 4. Create a sandbox with the provider attached
openshell sandbox create --name my-sandbox --provider vertex-local
```
Then inside the sandbox, launch the agent as shown below.
<Tabs>
<Tab title="Claude Code">
```shell
ANTHROPIC_BASE_URL="https://inference.local" ANTHROPIC_API_KEY=unused claude --bare
```
`--bare` skips the OAuth login flow and uses `ANTHROPIC_API_KEY` directly for authentication. The key value does not reach Vertex AI — `inference.local` strips it and injects the real GCP access token before forwarding.
<Warning>
Do not set `CLAUDE_CODE_USE_VERTEX=1` inside the sandbox. That flag makes Claude Code connect directly to Vertex AI and attempt GCP credential discovery (ADC file, metadata service), which fails because the sandbox does not expose GCP credentials. Use `inference.local` instead.
</Warning>
</Tab>
<Tab title="OpenCode">
```shell
ANTHROPIC_BASE_URL="https://inference.local/v1" ANTHROPIC_API_KEY=unused opencode
```
OpenCode requires `/v1` in the base URL. Without it, OpenCode sends `POST /messages` instead of `POST /v1/messages`, which does not match the inference pattern and is denied.
</Tab>
</Tabs>
### Policy Proposals
After running an agent, the TUI (`openshell term`) may show policy proposals for denied endpoints. Common ones for Vertex AI sandboxes:
| Endpoint | Action | Reason |
|---|---|---|
| `metadata.google.internal:80` | **Reject** | Resolves to `169.254.169.254` (GCE metadata service). Always blocked regardless of policy — the proxy blocks the resolved IP unconditionally to prevent credential exfiltration. |
| `downloads.claude.ai:443` | Approve if desired | Claude Code update checking and asset loading. Not required for inference. |
| `storage.googleapis.com:443` | Approve if desired | Google Cloud Storage. Used by some Claude Code features. Not required for inference. |
## From Existing Environment
If one of these token env vars is already set in your shell, create the provider with `--from-existing`:
- `GOOGLE_VERTEX_AI_TOKEN` or `VERTEX_AI_TOKEN`
- `GOOGLE_VERTEX_AI_SERVICE_ACCOUNT_TOKEN` or `VERTEX_AI_SERVICE_ACCOUNT_TOKEN`
OpenShell also reads these config env vars during `--from-existing`:
- `VERTEX_AI_PROJECT_ID`
- `VERTEX_AI_REGION`
- `GOOGLE_VERTEX_AI_BASE_URL` or `VERTEX_AI_BASE_URL`
- `VERTEX_AI_PUBLISHER`
Then create the provider:
```shell
openshell provider create \
--name vertex-env \
--type google-vertex-ai \
--from-existing
```
This reads credentials and config from the environment variables listed in the configuration keys table above.
## Next Steps
- To configure `inference.local` routing, refer to [Inference Routing](/sandboxes/inference-routing).
- To manage provider credentials and refresh, refer to [Providers](/sandboxes/manage-providers).
- To apply network policies to sandboxes using this provider, refer to [Policies](/sandboxes/policies).
@@ -0,0 +1,29 @@
---
# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
title: "Default Policy Reference"
sidebar-title: "Default Policy"
description: "Breakdown of the built-in default policy applied when you create an OpenShell sandbox without a custom policy."
keywords: "Generative AI, Cybersecurity, AI Agents, Sandboxing, Security, Policy"
position: 2
---
The default policy is the policy applied when you create an OpenShell sandbox without `--policy`. It is baked into the community base image ([`ghcr.io/nvidia/openshell-community/sandboxes/base`](https://github.com/NVIDIA/OpenShell-Community)) and defined in the community repo's `dev-sandbox-policy.yaml`.
## Agent Compatibility
The following table shows the coverage of the default policy for common agents.
| Agent | Coverage | Action Required |
|---|---|---|
| Claude Code | Full | None. Works out of the box. |
| OpenCode | Partial | Add `opencode.ai` endpoint and OpenCode binary paths. |
| Codex | None | Provide a complete custom policy with OpenAI endpoints and Codex binary paths. |
<Info>
If you run a non-Claude agent without a custom policy, the agent's API calls are denied by the proxy. You must provide a policy that declares the agent's endpoints and binaries.
</Info>
## Default Policy Blocks
The default policy blocks are defined in the community base image. Refer to the [OpenShell Community repository](https://github.com/NVIDIA/OpenShell-Community) for the full `dev-sandbox-policy.yaml` source.
@@ -0,0 +1,285 @@
---
# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
title: "Gateway Authentication"
description: "Gateway resolution, authentication modes, connection flow, OIDC support, and credential file layout."
keywords: "Generative AI, Cybersecurity, Gateway, Authentication, mTLS, OIDC, OpenID Connect, Edge Authentication, Reference"
position: 1
---
This page describes how the CLI resolves a gateway, authenticates with it, and where credentials are stored. For how to deploy or register gateways, refer to [Gateways](/sandboxes/manage-gateways).
## Gateway Resolution
When any CLI command needs to talk to the gateway, it resolves the target through a priority chain:
1. `--gateway-endpoint <URL>` flag (direct URL).
2. `-g <NAME>` flag.
3. `OPENSHELL_GATEWAY` environment variable.
4. Active gateway from `~/.config/openshell/active_gateway`, falling back to `/etc/openshell/active_gateway` when no user selection exists.
The CLI loads gateway metadata from `~/.config/openshell/gateways/<name>/metadata.json` first and falls back to `/etc/openshell/gateways/<name>/metadata.json` when no user entry exists.
## Authentication Status
`openshell status` reports gateway reachability and authentication separately. The public health RPC determines whether the gateway is connected and supplies its version. The CLI then calls the existing protected gateway-info capability query to verify the configured credentials. An authorization denial still confirms that the gateway authenticated the caller.
For example, an expired bearer token can produce `Status: Connected` and `Authentication: Failed`. This means the gateway is healthy but protected commands cannot use the current credentials. Run `openshell gateway login <name>` to re-authenticate. When a gateway predates the gateway-info capability query, the CLI reports `Authentication: Unverified` rather than treating a successful public health check as proof of authentication.
## Authentication Modes
The CLI uses one of these authentication modes depending on the gateway's configuration.
### mTLS
The default mode for local Docker, Podman, and VM gateways without OIDC. The CLI presents a client certificate during the TLS handshake, and the gateway can map the verified certificate subject to a local user principal when mTLS user authentication is enabled.
mTLS user authentication is for local single-user gateways. Kubernetes deployments must use OIDC or a trusted access proxy for user authentication; the Helm chart does not render `mtls_auth`.
Set these environment variables before starting the gateway:
| Environment variable | Purpose |
|---|---|
| `OPENSHELL_TLS_CERT` | Path to the gateway server certificate. |
| `OPENSHELL_TLS_KEY` | Path to the gateway server private key. |
| `OPENSHELL_TLS_CLIENT_CA` | Path to the CA certificate that verifies CLI client certificates. |
| `OPENSHELL_ENABLE_MTLS_AUTH` | Set to `true` to authenticate CLI callers from verified client certificates. Defaults on for local Docker, Podman, and VM gateways with no OIDC issuer. |
For local access, the server certificate must be valid for the endpoint the CLI uses. Include `localhost`, `127.0.0.1`, and `::1` in the certificate SANs when users connect to a local gateway through loopback.
Package-managed local gateways generate this bundle automatically for the `openshell` gateway name. Homebrew registers `https://localhost:17670`; Debian and RPM use `https://127.0.0.1:17670`.
When you register a package-managed local gateway with `openshell gateway add <endpoint> --local --name openshell`, the CLI refreshes its mTLS bundle from the package-managed TLS directory.
On Homebrew, the gateway service also mirrors the Docker sandbox client bundle into `$HOME/.local/state/openshell/homebrew/tls` before startup so Docker Desktop can bind-mount the files into sandbox containers.
The CLI loads its mTLS bundle from `~/.config/openshell/gateways/<name>/mtls/`:
| File | Purpose |
|---|---|
| `ca.crt` | CA certificate that verifies the gateway server certificate. |
| `tls.crt` | Client certificate. It must chain to `OPENSHELL_TLS_CLIENT_CA`. |
| `tls.key` | Client private key for `tls.crt`. |
The connection flow:
1. The CLI loads the certificate files.
2. Opens a TCP connection to the gateway endpoint.
3. Performs a TLS handshake, presenting the client certificate.
4. The gateway verifies the client certificate against its CA.
5. When mTLS user authentication is enabled, the gateway maps the verified certificate subject to a user principal.
6. The gateway authorizes the gRPC method.
### OIDC
Gateways can validate OpenID Connect access tokens on gRPC requests. Configure OIDC when you want users, operators, or automation to authenticate with an identity provider such as Keycloak, Entra ID, or Okta.
OIDC is application-layer authentication. TLS still controls the transport. If TLS client certificates remain required, the CLI must also have an mTLS bundle for the gateway.
Configure the gateway with an issuer and audience:
```shell
openshell-gateway \
--oidc-issuer https://idp.example.com/realms/openshell \
--oidc-audience openshell-cli \
--oidc-roles-claim realm_access.roles \
--oidc-admin-role openshell-admin \
--oidc-user-role openshell-user
```
The same settings are available through environment variables:
| Environment variable | Purpose | Default |
|---|---|---|
| `OPENSHELL_OIDC_ISSUER` | OIDC issuer URL. The gateway discovers `/.well-known/openid-configuration` from this URL. | None |
| `OPENSHELL_OIDC_AUDIENCE` | Expected JWT `aud` claim. | `openshell-cli` |
| `OPENSHELL_OIDC_JWKS_TTL` | JWKS cache TTL in seconds. The gateway also refreshes on an unknown key ID. | `3600` |
| `OPENSHELL_OIDC_ROLES_CLAIM` | Dot-separated claim path containing roles. | `realm_access.roles` |
| `OPENSHELL_OIDC_ADMIN_ROLE` | Role required for admin operations. | `openshell-admin` |
| `OPENSHELL_OIDC_USER_ROLE` | Role required for standard user operations. | `openshell-user` |
| `OPENSHELL_OIDC_SCOPES_CLAIM` | Dot-separated claim path containing scopes. Empty disables scope enforcement. | Empty |
For Helm deployments, set the same values under `server.oidc`:
```yaml
server:
oidc:
issuer: https://idp.example.com/realms/openshell
audience: openshell-cli
rolesClaim: realm_access.roles
adminRole: openshell-admin
userRole: openshell-user
scopesClaim: ""
```
Register an OIDC gateway with the CLI:
```shell
openshell gateway add https://gateway.example.com \
--name production \
--oidc-issuer https://idp.example.com/realms/openshell \
--oidc-client-id openshell-cli \
--oidc-audience openshell-cli
```
When you register or log in to an OIDC gateway, the CLI uses the Authorization Code flow with PKCE. It opens a browser, receives the authorization code on a localhost callback, exchanges the code for tokens, and stores the token bundle under the gateway credential directory. After `openshell gateway logout`, the next browser login asks the identity provider for a fresh login prompt so you can choose a different browser user instead of silently reusing the previous session. If `OPENSHELL_OIDC_CLIENT_SECRET` is set, the CLI uses the client credentials flow instead. Use that mode for CI and other non-interactive automation.
Official Python, TypeScript, and Go SDKs can perform renewable client-credentials
authentication directly. Configure the service account at the identity provider
with the audience, roles, scopes, and workspace membership required by the
gateway. The SDKs discover the token endpoint, attach the bearer token to each
RPC, and repeat the grant before expiry. They keep the client secret and access
token in memory and do not update the CLI's `oidc_token.json`. SDK clients
require TLS when sending these credentials to a non-loopback gateway.
<Tabs>
<Tab title="Python">
```python
from openshell import ClientCredentialsAuth, SandboxClient
auth = ClientCredentialsAuth(
client_secret=lambda: load_secret(),
# Omit issuer/client_id/scopes/audience to use active gateway metadata.
)
client = SandboxClient.from_active_cluster(client_credentials=auth)
```
</Tab>
<Tab title="TypeScript">
```ts
import { clientCredentials, OpenShellClient } from '@nvidia/openshell-sdk'
const client = await OpenShellClient.connect({
gateway: 'https://gateway.example.com',
oidcTokenProvider: clientCredentials({
issuer: 'https://idp.example.com/realms/openshell',
clientId: 'openshell-service',
clientSecret: () => loadSecret(),
audience: 'openshell-gateway',
scopes: ['sandbox:read', 'sandbox:write'],
}),
})
```
</Tab>
<Tab title="Go">
```go
auth, err := oidc.NewClientCredentialsAuth(
oidc.WithGateway("production"),
oidc.WithClientSecretProvider(loadSecret),
)
client, err := v1.NewClient(v1.Config{Address: address, Auth: auth, TLS: tlsConfig})
```
</Tab>
</Tabs>
The connection flow:
For a headless environment, set `OPENSHELL_NO_BROWSER=1` before registering or logging in to the gateway. When this variable is set and `OPENSHELL_OIDC_CLIENT_SECRET` is not configured, the CLI uses the Device Authorization Grant (RFC 8628) with S256 PKCE. This flow prompts the user to visit a verification URL on any device with a browser and enter a displayed code. The CLI polls the token endpoint until the user completes authorization. This requires the OIDC client to have the device authorization grant enabled on the identity provider.
1. The CLI loads the stored OIDC token bundle.
2. If the access token is expired and a refresh token is available, the CLI refreshes it with the OIDC scopes saved in the gateway metadata.
3. The CLI connects to the gateway and attaches `authorization: Bearer <token>` metadata to each gRPC request.
4. The gateway validates the JWT signature, issuer, audience, expiration, and key ID against the issuer's JWKS.
5. The gateway extracts roles and optional scopes from the configured claim paths.
6. The gateway authorizes the gRPC method. Platform-scoped methods require the configured admin role. Workspace-scoped methods require the configured user role and a sufficient membership in the target workspace. Admin role holders satisfy user-role checks and bypass workspace membership checks.
For the Platform Admin, Workspace Admin, and Workspace User permissions, refer
to [Manage Workspaces and Access](/sandboxes/manage-workspaces).
If `OPENSHELL_OIDC_SCOPES_CLAIM` is set, the gateway also enforces scopes. It accepts space-delimited scope strings such as `scope: "openid sandbox:read"` and JSON arrays such as `scp: ["sandbox:read"]`. Standard OIDC scopes such as `openid`, `profile`, `email`, and `offline_access` are ignored for authorization. `openshell:all` grants access to all scoped methods.
Supervisor-to-gateway RPCs do not use user OIDC tokens or mTLS user identity. Each sandbox supervisor presents a gateway-minted `Authorization: Bearer` token scoped to its sandbox ID. On Kubernetes, the gateway mints that token only after TokenReview validates the projected ServiceAccount token, the pod UID matches the live pod, and the pod's controlling `Sandbox` ownerReference matches the live Sandbox CR. Log upload, policy status, credential environment lookup, inference bundle lookup, and sandbox config sync run with sandbox-restricted scope, while CLI users authenticate with OIDC, edge auth, local mTLS user authentication, or an explicitly enabled unauthenticated local developer mode. `GetInferenceBundle` returns route material that includes provider credentials, so it requires a sandbox principal; user callers manage inference configuration through the user-facing inference APIs instead.
Re-authenticate an OIDC gateway with:
```shell
openshell gateway login production
```
Inspect the identity the gateway validated:
```shell
openshell whoami
openshell whoami --output json
```
The output includes the stable subject used for workspace membership, the
display name when available, identity provider, roles, and scopes. The gateway
returns its validated identity; the CLI does not infer these values from an
unverified local token payload. Use the `subject` value when adding the user to
a workspace. For membership commands, refer to
[Manage Workspaces and Access](/sandboxes/manage-workspaces).
### Edge JWT (cloud gateways)
For gateways behind a reverse proxy that handles authentication (e.g. Cloudflare Access), the CLI uses a browser-based login flow and routes traffic through a WebSocket tunnel.
**Registration flow** (`openshell gateway add https://gateway.example.com`):
1. The CLI stores gateway metadata with the edge authentication mode.
2. Opens your browser to the gateway's authentication endpoint.
3. The reverse proxy handles login (SSO, identity provider, etc.).
4. After authentication, the browser relays the authorization token back to the CLI through a localhost callback.
5. The CLI stores the token and sets the gateway as active.
**Connection flow** (subsequent commands):
1. The CLI starts a local proxy that listens on an ephemeral port.
2. The proxy opens a WebSocket connection (`wss://`) to the gateway, attaching the stored bearer token in the upgrade headers.
3. The reverse proxy authenticates the WebSocket upgrade request.
4. The gateway bridges the WebSocket into the same service that handles direct mTLS connections.
5. CLI commands send requests through the local proxy as plaintext HTTP/2 over the tunnel.
This is transparent to the user. All CLI commands work the same regardless of whether the gateway uses mTLS, OIDC, or edge authentication.
**Re-authentication**: If the token expires, run `openshell gateway login` to open the browser flow again and update the stored token.
### Plaintext
When a gateway is deployed with `server.disableTls=true`, TLS is disabled entirely. The CLI connects over plain HTTP/2. This mode is intended for local port-forwarding or gateways behind a trusted reverse proxy or tunnel that handles TLS termination externally.
Register a plaintext gateway with an explicit `http://` endpoint:
```shell
openshell gateway add http://127.0.0.1:17670 --local
```
For Kubernetes local development, the Helm Skaffold overlay enables `[openshell.gateway.auth] allow_unauthenticated_users = true` so a port-forwarded plaintext gateway works without OIDC or mTLS user credentials. Leave this disabled for shared and production clusters.
This stores the gateway with `auth_mode = plaintext`, skips mTLS client certificate lookup, and does not open the browser login flow.
## File Layout
User-managed gateway credentials and metadata are stored under `~/.config/openshell/`:
```text
openshell/
active_gateway # Plain text: active gateway name
gateways/
<name>/
metadata.json # Gateway metadata (endpoint, auth mode, type)
mtls/ # mTLS bundle (local and remote gateways)
ca.crt # CA certificate
tls.crt # Client certificate
tls.key # Client private key
edge_token # Edge auth JWT (cloud gateways)
oidc_token.json # OIDC access token, refresh token, and expiry metadata
last_sandbox # Last-used sandbox for this gateway
```
Installers can also seed read-only defaults under `/etc/openshell/` (or a non-empty absolute `OPENSHELL_SYSTEM_GATEWAY_DIR` override) using the same active-gateway and metadata layout:
```text
/etc/openshell/
active_gateway
gateways/
<name>/
metadata.json
```
Only `active_gateway` and `metadata.json` fall back to the system layout. mTLS bundles, OIDC tokens, edge tokens, and `last_sandbox` remain per-user state.
For OIDC gateways, `metadata.json` also stores the issuer, CLI client ID, optional audience, and requested scopes. Treat `oidc_token.json` as a credential. OpenShell writes it with owner-only file permissions.
@@ -0,0 +1,816 @@
---
# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
title: "Gateway Configuration File"
sidebar-title: "Gateway Config"
description: "Reference for the OpenShell gateway TOML configuration file (RFC 0003)."
keywords: "Generative AI, Cybersecurity, AI Agents, Sandboxing, Gateway, Configuration, TOML, Reference"
position: 5
---
The OpenShell gateway reads its configuration from a TOML file when `--config` or `OPENSHELL_GATEWAY_CONFIG` is set. When neither is set, the gateway reads `$XDG_CONFIG_HOME/openshell/gateway.toml` if that file exists. If no config file exists, the gateway starts from built-in defaults. Gateway process flags and gateway `OPENSHELL_*` environment variables override the file. Compute driver settings live in the driver TOML tables. See [RFC 0003](https://github.com/NVIDIA/OpenShell/blob/main/rfc/0003-gateway-configuration/README.md) for the full schema.
## Source Precedence
```text
Gateway CLI flag > gateway OPENSHELL_* env var > TOML file > built-in default
```
`database_url` is env-only. The loader rejects it when it appears in the file. When `OPENSHELL_DB_URL` is unset, the gateway stores its SQLite database under `$XDG_STATE_HOME/openshell/gateway/openshell.db`.
`name` assigns an operator-facing identity to the gateway installation. Set it with `[openshell.gateway].name`, `--name`, or `OPENSHELL_GATEWAY_NAME`. It defaults to `openshell`; the Helm chart defaults it to the chart fullname so all replicas in one installation share a name. Chart fullnames are only unique within their Kubernetes namespace, so set `server.name` explicitly when one collector receives telemetry from multiple namespaces or clusters. This identity is independent of client-side gateway aliases, TLS names, and `gateway_jwt.gateway_id`.
## Package-Managed Locations
Package-managed gateways do not require a TOML file. Create one at the package's optional config location when you need to override built-in defaults. Set `OPENSHELL_GATEWAY_CONFIG` in the launch environment to use a different file.
| Package | Optional Gateway TOML location |
|---|---|
| Homebrew | `$XDG_CONFIG_HOME/openshell/gateway.toml` when it exists, otherwise the Homebrew prefix config such as `/opt/homebrew/var/openshell/gateway.toml`. |
| Debian/Ubuntu | `$XDG_CONFIG_HOME/openshell/gateway.toml`, usually `~/.config/openshell/gateway.toml` for the systemd user service. |
| Fedora/RHEL RPM | `$XDG_CONFIG_HOME/openshell/gateway.toml`, usually `~/.config/openshell/gateway.toml` for the systemd user service. |
| Snap | `$SNAP_COMMON/gateway.toml`, usually `/var/snap/openshell/common/gateway.toml`. |
The Fedora/RHEL RPM template leaves `[openshell.gateway].bind_address` unset. The gateway therefore uses its built-in `127.0.0.1:17670` primary listener. The Podman driver negotiates separate, restricted listeners for sandbox callbacks, so the primary listener does not need a wildcard address. Set `bind_address` explicitly only when clients must reach the primary multiplexed API through another interface.
The Homebrew formula creates its prefix config without setting `bind_address`, so the gateway uses its built-in `127.0.0.1:17670` primary listener. Docker Desktop and Podman Machine reuse that listener for sandbox callbacks. A user config takes precedence. Upgrades preserve user-edited configs and migrate only an unchanged prefix config generated with the affected IPv6-loopback default.
## Layout
The file is rooted at `[openshell]`. Gateway-wide settings live under `[openshell.gateway]`. Each compute driver owns its own `[openshell.drivers.<name>]` table. Credential drivers own `[openshell.credential_drivers.<name>]` tables. Shared compute-driver keys set at gateway scope are inherited into compute driver tables when not overridden.
```toml
[openshell]
version = 1
[openshell.gateway]
# ... gateway-wide settings ...
[openshell.gateway.tls]
# ... gateway listener TLS ...
[openshell.gateway.oidc]
# ... JWT bearer auth ...
[openshell.drivers.kubernetes]
# ... driver-specific settings ...
[openshell.credential_drivers.kubernetes-secrets]
# ... credential-driver-specific settings ...
```
## Full Example
A complete gateway configuration covering every section. Trim to the fields you need.
```toml
# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
[openshell]
version = 1
[openshell.gateway]
name = "production-us-west"
bind_address = "0.0.0.0:8080"
health_bind_address = "0.0.0.0:8081"
metrics_bind_address = "0.0.0.0:9090"
log_level = "info"
# When empty, the gateway auto-detects Kubernetes, then Podman, then Docker.
# VM is never auto-detected and requires an explicit entry here.
compute_drivers = ["kubernetes"]
# Optional external provider credential storage backend. Omit this key to use
# the gateway's default encrypted database credential storage.
credential_drivers = ["kubernetes-secrets"]
sandbox_namespace = "openshell"
ssh_session_ttl_secs = 3600
# Reject invalid policy generations securely by default. Set
# "retain_last_valid" only when availability takes priority.
policy_validation_failure_mode = "fail_closed"
# Subject Alternative Names baked into the gateway server certificate.
# Wildcard DNS SANs (e.g. "*.dev.openshell.localhost") also enable sandbox
# service URLs under that domain.
server_sans = ["openshell", "*.dev.openshell.localhost"]
# Allow plaintext HTTP routing for loopback sandbox service URLs.
enable_loopback_service_http = true
# Set true only for local plaintext gateways or trusted TLS termination.
disable_tls = false
# Shared driver defaults. These inherit into [openshell.drivers.<name>] tables
# when the driver-specific table does not override them.
default_image = "ghcr.io/nvidia/openshell/sandbox:latest"
# Defaults to the gateway version; override to pin a specific build.
# supervisor_image = "ghcr.io/nvidia/openshell/supervisor:<version>"
client_tls_secret_name = "openshell-client-tls"
service_account_name = "openshell-sandbox"
host_gateway_ip = "10.0.0.1"
enable_user_namespaces = false
sa_token_ttl_secs = 3600
guest_tls_ca = "/etc/openshell/certs/ca.pem"
guest_tls_cert = "/etc/openshell/certs/client.pem"
guest_tls_key = "/etc/openshell/certs/client-key.pem"
# Optional gRPC rate limit. Both values must be positive to enable the limit.
# Set either value to 0, or omit both, to disable rate limiting.
grpc_rate_limit_requests = 120
grpc_rate_limit_window_seconds = 60
# Optional exact provider-profile source composition. When omitted, the
# gateway uses builtin + user.
provider_profile_sources = [
{ type = "builtin" },
{ type = "user" },
]
# Operator-run supervisor middleware. The gRPC endpoint must be reachable from
# both the gateway and sandbox supervisors.
[[openshell.supervisor.middleware]]
name = "local-content-guard"
grpc_endpoint = "https://host.openshell.internal:50051"
tls_ca_cert_path = "/etc/openshell/certs/content-guard-ca.pem"
audience = "urn:openshell:middleware:local-content-guard"
max_payload_bytes = 262144
timeout = "500ms"
# Gateway listener TLS (distinct from the per-driver guest_tls_*).
[openshell.gateway.tls]
cert_path = "/etc/openshell/certs/gateway.pem"
key_path = "/etc/openshell/certs/gateway-key.pem"
client_ca_path = "/etc/openshell/certs/client-ca.pem"
require_client_auth = false
# Optional: SNI-based dual certificate for external (e.g. ACME) TLS.
# external_cert_path = "/etc/openshell/certs/external.pem"
# external_key_path = "/etc/openshell/certs/external-key.pem"
# external_server_names = ["gateway.example.com"]
[openshell.gateway.gateway_jwt]
signing_key_path = "/etc/openshell/jwt/signing.pem"
public_key_path = "/etc/openshell/jwt/public.pem"
kid_path = "/etc/openshell/jwt/kid"
gateway_id = "openshell"
# Omit or set to 0 only for local single-player Docker, Podman, or VM gateways.
ttl_secs = 3600
[openshell.gateway.auth]
allow_unauthenticated_users = false
[openshell.gateway.mtls_auth]
enabled = false
# OTLP export. Omit this table entirely to disable it.
[openshell.gateway.otlp]
endpoint = "http://otel-collector.observability.svc:4317"
service_name = "openshell-gateway"
[openshell.gateway.oidc]
issuer = "https://idp.example.com/realms/openshell"
audience = "openshell-cli"
jwks_ttl_secs = 3600
roles_claim = "realm_access.roles"
admin_role = "openshell-admin"
user_role = "openshell-user"
scopes_claim = ""
[[openshell.gateway.interceptors]]
name = "quota"
grpc_endpoint = "unix:///run/openshell/interceptors/quota.sock"
audience = "urn:openshell:interceptor:quota"
order = 10
failure_policy = "fail_closed"
binding_policy = "allowlist"
timeout = "500ms"
max_response_bytes = 1048576
max_patches = 32
[[openshell.gateway.interceptors.bindings]]
rpc = "openshell.v1.OpenShell/CreateSandbox"
phases = ["modify_operation", "validate"]
failure_policy = "fail_closed"
[[openshell.gateway.interceptors.bindings]]
rpc = "openshell.v1.OpenShell/UpdateConfig"
phases = ["validate"]
[openshell.credential_drivers.kubernetes-secrets]
namespace = "openshell"
allow_reference_namespace = false
```
Local Docker, Podman, and VM gateways can also set `[openshell.gateway.mtls_auth] enabled = true` to authenticate CLI callers from verified client certificates. Kubernetes deployments must leave this unset and use OIDC or a trusted access proxy; the Helm chart does not render this table.
`[openshell.gateway.tls]` supports optional SNI-based dual-certificate mode for deployments that need separate internal and external server certificates. Set `external_cert_path` and `external_key_path` to point at the external (e.g. ACME/publicly-trusted) certificate and key. List the hostnames that should be served with the external certificate in `external_server_names`. Connections whose TLS SNI hostname matches one of those names receive the external certificate; all other connections (including those with no SNI) receive the primary internal certificate from `cert_path`/`key_path`. Both fields must be set together — providing only one is a configuration error. On Kubernetes with the Helm chart, the external certificate is managed automatically when `certManager.serverIssuerRef.name` is set; the chart populates these fields from the cert-manager-issued external server certificate.
`[openshell.gateway] policy_validation_failure_mode` controls what sandbox supervisors do when a complete candidate policy fails runtime validation. The default, `fail_closed`, deactivates the previous network policy, closes relays pinned to it, and denies new egress until a valid generation loads. `retain_last_valid` leaves the previous valid generation active. Both modes reject the candidate atomically; startup always fails closed when no previous valid generation exists. Gateway mutation paths that can preflight a known effective scope reject invalid candidates before persistence and leave the active policy unchanged regardless of this setting. Changing the value requires restarting the gateway so it can reload `gateway.toml` and distribute the new posture to sandbox supervisors.
`[openshell.gateway.gateway_jwt] ttl_secs` controls gateway-minted sandbox JWT lifetime. When omitted, it defaults to `0`: the token `exp` claim and `expires_at_ms` response field become `0`, and the sandbox JWT does not expire. Use that default only for local single-player Docker, Podman, or VM gateways. Kubernetes and other shared deployments should set a positive TTL; Helm renders `3600` seconds by default, and the gateway logs a warning when a Kubernetes gateway uses `0`.
`[openshell.gateway.auth] allow_unauthenticated_users = true` is an unsafe local-development and trusted-proxy escape hatch. It accepts user-facing CLI/API calls without OIDC or mTLS credentials while sandbox supervisors still authenticate with gateway-minted sandbox JWTs. Leave it false for shared and production gateways.
## OTLP Export
`[openshell.gateway.otlp]` enables OpenTelemetry export over OTLP/gRPC. Omit the table to disable export; there is no separate `enabled` flag.
The gateway already uses the Rust `tracing` framework for structured logs sent to stdout and the sandbox log stream. Enabling this section adds an OpenTelemetry layer to the same tracing subscriber. It exports span trees to an OTLP collector without exporting, replacing, or redirecting the existing log events.
```toml
[openshell.gateway.otlp]
endpoint = "http://otel-collector.observability.svc:4317"
service_name = "openshell-gateway"
```
`endpoint` is required and must be a valid URI. If it is malformed, the gateway logs the configuration error and continues with export disabled. It does not connect at startup: an unreachable collector produces export failures, never a failure to serve.
The transport is **OTLP over gRPC only**. HTTP/protobuf and HTTP/JSON are not supported, and `OTEL_EXPORTER_OTLP_PROTOCOL` has no effect. Point `endpoint` at a collector's gRPC receiver, conventionally port `4317`, not the HTTP receiver on `4318`. A URI alone cannot distinguish the two, so an HTTP endpoint is accepted at startup and then fails on export.
The OpenTelemetry SDK logs export failures after startup. Spans in a failed batch are dropped rather than retried.
`service_name` sets the gateway's `service.name` resource attribute and defaults to `openshell-gateway`. The gateway also reports `service.version`, `openshell.gateway.name` from the gateway's configured `name`, and `openshell.gateway.compute_driver`.
Only OpenTelemetry traces are exported. Inbound gRPC and HTTP requests produce server spans named for the RPC or HTTP method. Store and compute-driver operations appear as child spans. Internal reconciliation, credential-refresh, and driver-watch loops create operation roots for their store work because no inbound request supplies a parent. The gateway continues valid W3C `traceparent` context and starts a new trace when none is supplied. Request spans carry `method`, `path`, and the `request_id` that also appears in gateway logs. Health endpoint spans use DEBUG level and are not exported by the default INFO filter.
The gateway forwards the OTLP configuration, configured gateway name, and W3C trace context to managed external drivers. Built-in drivers also export their spans to the same collector through dedicated in-process providers. Driver spans retain the gateway trace context, use a distinct service name such as `openshell-driver-docker` or `openshell-driver-podman`, and carry the gateway name as the `openshell.gateway.name` resource attribute. Operator-run external drivers own their own telemetry configuration.
For Helm deployments, set `server.otlp.endpoint` to render this table. The
optional `server.otlp.serviceName` value overrides the gateway service name;
driver service names remain fixed.
The local `mise run helm:k3s:create` workflow installs a trace collector and UI,
enables Agent Sandbox controller tracing on supported releases, and configures
both the controller and the existing Skaffold deployment to export traces to it
automatically.
The local `gateway`, `gateway:docker`, `gateway:podman`, and `gateway:vm` tasks
also set the gateway installation name in their generated configuration. Their
defaults are driver-specific (`kubernetes-dev`, `docker-dev`, `podman-dev`, and
`vm-dev`), and their documented `OPENSHELL_*_GATEWAY_NAME` overrides update both
CLI registration and exported telemetry identity.
Run `mise run helm:k3s:forward` to forward OTLP/gRPC to local port `4317` and
the trace UI to local port `18888`. When Skaffold has deployed a Kubernetes
gateway, the task also forwards it to local port `8090`; otherwise it continues
with the collector ports only. The Skaffold run task registers the forwarded
gateway under the worktree-specific k3d cluster name; select it with
`openshell gateway select <name>`. The local Podman, Docker, and VM gateway
tasks export to the forwarded receiver automatically.
### Tuning
This table decides whether and where to export. How the SDK exports is controlled by the standard OpenTelemetry environment variables, which the gateway reads through the SDK rather than mirroring as TOML keys:
| Variable | Effect |
|---|---|
| `OTEL_TRACES_SAMPLER`, `OTEL_TRACES_SAMPLER_ARG` | Sampling strategy and ratio. Defaults to `parentbased_always_on`. |
| `OTEL_BSP_SCHEDULE_DELAY`, `OTEL_BSP_MAX_QUEUE_SIZE`, `OTEL_BSP_MAX_EXPORT_BATCH_SIZE`, `OTEL_BSP_EXPORT_TIMEOUT` | Batch span processor tuning. |
| `OTEL_RESOURCE_ATTRIBUTES` | Additional resource attributes, such as `deployment.environment=prod`. |
| `OTEL_SPAN_ATTRIBUTE_COUNT_LIMIT`, `OTEL_SPAN_EVENT_COUNT_LIMIT`, `OTEL_SPAN_LINK_COUNT_LIMIT` | Per-span limits. |
| `OTEL_EXPORTER_OTLP_HEADERS`, `OTEL_EXPORTER_OTLP_COMPRESSION`, `OTEL_EXPORTER_OTLP_TIMEOUT` | Exporter transport tuning. |
| `OTEL_EXPORTER_OTLP_PROTOCOL` | No effect. The gateway is built with the gRPC exporter only. |
To sample 10% of traces:
```shell
OTEL_TRACES_SAMPLER=parentbased_traceidratio OTEL_TRACES_SAMPLER_ARG=0.1
```
`OTEL_EXPORTER_OTLP_ENDPOINT` and `OTEL_EXPORTER_OTLP_TRACES_ENDPOINT` are deliberately ignored. Enablement has one source, so an environment variable cannot silently turn export on or redirect it. `OTEL_RESOURCE_ATTRIBUTES` does apply and adds attributes, but a `service_name` set here wins over `OTEL_SERVICE_NAME`.
The gateway flushes buffered spans during shutdown, so spans from in-flight requests survive a `SIGTERM`.
## Supervisor Middleware Services
Register operator-run supervisor middleware services with one or more `[[openshell.supervisor.middleware]]` entries. Registration is static and operator-owned; changing it requires restarting the gateway.
```toml
[[openshell.supervisor.middleware]]
name = "local-content-guard"
grpc_endpoint = "https://host.openshell.internal:50051"
tls_ca_cert_path = "/etc/openshell/certs/content-guard-ca.pem"
audience = "urn:openshell:middleware:local-content-guard"
max_payload_bytes = 262144
timeout = "500ms"
```
Each service implements the supervisor middleware gRPC contract and exposes bindings through `Describe`. Policies reference the operator-owned registration `name`, attaching the complete middleware and all of its bindings. Bindings are identified by operation and phase. A manifest may expose at most one binding for each operation and phase pair. V1 supports `HttpRequest/pre_credentials` and `WebSocketMessage/pre_credentials`, so a service can inspect HTTP, WebSocket, or both. Registration names must be unique, and operator-run registrations cannot claim the reserved `openshell/` namespace. The service-reported manifest name is diagnostic metadata and does not need to match the registration name.
The gateway connects to every registered service and validates `Describe` before it starts. The service must therefore be running before the gateway. Policy creation and full policy updates call `ValidateConfig`; an unavailable service or invalid middleware configuration rejects the policy before persistence.
`max_payload_bytes` is the shared operator limit for inspectable logical payloads across every binding exposed by the service. It caps HTTP request and replacement bodies as well as complete WebSocket text messages and replacements. The value must be greater than zero, no larger than each binding's advertised `max_payload_bytes` capability, and no larger than the 4 MiB platform maximum. OpenShell rejects oversized values instead of silently clamping them. Binary WebSocket messages are not exposed to V1 middleware, so this field does not limit binary pass-through. Middleware gRPC servers should allow messages of at least 4 MiB plus 293 KiB so a maximum-size payload and its protobuf envelope fit on the transport.
`timeout` is the operator-configured service-wide RPC timeout. It accepts the same compact duration syntax as gateway interceptors: an integer followed by `ms` or `s`, such as `500ms` or `2s`. Values must be between `10ms` and `30s`, inclusive. Omit the field to use the 500 ms platform default. A binding may advertise a shorter `timeout` in the `Describe` manifest, but it cannot extend the operator-configured deadline; OpenShell uses the smaller value. OpenShell validates both levels before accepting the service. The operator-configured service timeout applies to `Describe` and `ValidateConfig`. The effective binding timeout applies only to `EvaluateHttpRequest`, WebSocket preflight, and each WebSocket message. An accepted WebSocket stream has no connection-wide RPC deadline.
The service `grpc_endpoint` supports plaintext `http://` and TLS `https://`. HTTPS uses the platform trust store unless `tls_ca_cert_path` names a certificate-only PEM bundle. OpenShell rejects bundles containing private keys, loads the certificates at gateway startup, and distributes only public certificates to sandbox supervisors; normal TLS hostname verification still applies. `audience` sets the exact audience for gateway-minted service tokens and defaults to `urn:openshell:extension:middleware:<name>`. After authenticated `Describe` succeeds, OpenShell treats a non-empty manifest `expected_audience` as a consistency assertion and refuses to start when it differs from the configured audience. A strict verifier may reject an incorrect audience before returning the manifest.
When `gateway_jwt` is configured, OpenShell attaches short-lived bearer credentials to gateway and supervisor calls and requires `https://`. A middleware endpoint must be reachable from sandbox supervisors, so Unix sockets are not an option here. Set `allow_insecure_transport = true` on a registration to keep a plaintext `http://` endpoint: OpenShell then attaches no credential, supervisors do not request one, and the gateway logs a warning naming the registration at every startup. mTLS client authentication, health checks, and runtime registration are not currently supported. The endpoint must be reachable from both the gateway and sandbox supervisors; use `host.openshell.internal` or another shared address that can be resolved in both places.
See [Supervisor Middleware](/extensibility/supervisor-middleware) for selection, failure, payload-limit, and operational guidance.
## Gateway Interceptors
`[[openshell.gateway.interceptors]]` configures gateway-side interceptor services. The gateway calls each service's `Describe` RPC at startup, validates its declared OpenShell RPC bindings against the compiled service descriptor, and applies matching phases from a central gRPC middleware path. Interceptors can target only methods in the gateway's built-in allowlist of unary mutation RPCs. New RPCs are non-interceptable until they are deliberately added to that allowlist; adding one does not require handler-specific interceptor code. Request bodies are exposed as protobuf JSON objects. Fields marked secret in the protobuf schema are recursively omitted from requests and post-commit responses. Interceptors cannot patch an omitted field or a containing object.
HTTPS interceptor endpoints use the platform trust store by default. Set `tls_ca_cert_path` to a PEM certificate bundle for a private CA; normal TLS hostname verification still applies. `audience` sets the exact audience for gateway-minted service tokens and defaults to `urn:openshell:extension:interceptor:<name>`. After authenticated `Describe` succeeds, the gateway treats a non-empty manifest `expected_audience` as a consistency assertion and refuses to start when it differs from the configured audience. A strict verifier may reject an incorrect audience before returning the manifest. When `gateway_jwt` is configured, network interceptors must use HTTPS and receive short-lived gateway-caller bearer credentials; local Unix sockets are also supported. Set `allow_insecure_transport = true` to keep a plaintext `http://` interceptor endpoint with no credential attached and a startup warning.
### Extension Token Verification Endpoints
When `gateway_jwt` is configured the gateway publishes two unauthenticated documents that extension services use to verify OpenShell callers:
| Path | Contents |
| --- | --- |
| `/.well-known/jwks.json` | Single-key JWKS holding the Ed25519 public key and its `kid`. |
| `/.well-known/openid-configuration` | OIDC-shaped discovery metadata: the exact expected `iss`, an absolute `jwks_uri`, and `EdDSA` as the only supported signing algorithm. |
The discovery document is OIDC-shaped rather than OIDC-compliant: `issuer` is the gateway identity (`openshell-gateway:<gateway_id>`), not the URL the document is served from. Verifiers compare `iss` against that value and must not infer trust from the document's location; the serving TLS connection is what authenticates the gateway. Both endpoints return `404` when `gateway_jwt` is not configured.
`binding_policy` controls how the manifest and operator binding configuration combine:
- `dynamic` enables valid manifest bindings and treats configured entries as optional narrowing overrides. This is the compatibility default. The gateway logs a startup warning because the interceptor controls its non-secret RPC authority.
- `allowlist` enables only configured RPC selectors and phases. The gateway ignores and logs extra manifest declarations, but fails startup when the manifest omits a configured RPC or phase.
- `exact` requires the configured and manifest RPC selectors and phases to match exactly.
Bindings under `allowlist` and `exact` require `rpc` or `service` plus `method`, and a nonempty `phases` list. They match by RPC rather than manifest binding ID. Duplicate selectors, `id`, and `disabled` are invalid in these modes; omit a binding to disable it. Binding failure policy comes from the binding, then the configured service, then defaults to `fail_closed`. Manifest failure policies do not override strict-mode operator configuration.
Bindings that include `post_commit` must resolve to `failure_policy = "fail_open"`; the gateway rejects fail-closed post-commit configuration at startup. Post-commit evaluation is observational. If response observation or interceptor evaluation fails after a handler succeeds, the gateway logs and counts the failure but returns the committed response unchanged.
`provider_profile_sources` selects the exact ordered provider-profile source set. Omit it to use the built-in and user-managed sources. Use `{ type = "builtin" }`, `{ type = "user" }`, and `{ type = "interceptor", name = "<configured-interceptor-name>" }` entries to compose a catalog. An interceptor entry must reference a configured service whose manifest advertises `provider_profiles = true`. Selecting only an interceptor makes that catalog authoritative; include local sources explicitly to compose them. Empty or duplicate source lists, unknown interceptor names, and duplicate normalized profile IDs fail closed. Source order controls collection and diagnostics, not override precedence.
For an authoritative interceptor catalog, select only that interceptor:
```toml
[openshell.gateway]
provider_profile_sources = [
{ type = "interceptor", name = "provider-governance" },
]
```
The gateway validates snapshot structure and provider-profile semantics. It treats a configured interceptor as the source trust boundary and does not verify signature, hash, or key annotations in profile payloads.
`failure_policy` accepts `fail_closed` or `fail_open`. `timeout` accepts `ms` and `s` suffixes. In `dynamic` mode, binding overrides may select a manifest binding by `id`, `rpc`, or `service` plus `method`; they can disable a binding, narrow its phases, or override its failure policy.
`image_pull_policy` is intentionally not a shared gateway key. Kubernetes and Docker use `Always`, `IfNotPresent`, or `Never`. Podman uses `always`, `missing`, `never`, or `newer`. Set it inside the relevant driver table.
## Credential Drivers
Set `credential_drivers` only when the gateway should store provider secrets in an external credential backend. Provider secrets include injectable credentials and gateway-only refresh material such as OAuth refresh tokens, client secrets, and service-account private keys. OpenShell supports at most one enabled credential driver at a time. When `credential_drivers` is omitted, the gateway uses its default encrypted database credential storage. `credential_drivers = []` is invalid in the TOML file; omit the field for the default encrypted store, or select a backend such as `kubernetes-secrets` or `vault`.
Credential driver tables are backend-owned and live under `[openshell.credential_drivers.<name>]`. Built-in drivers default to in-tree transport, so they do not need a `transport` field. Use `transport = "uds"` with an absolute `socket_path` only for a remote gRPC driver over a Unix domain socket.
```toml
[openshell.gateway.credential_storage]
key_encryption_key_path = "/var/lib/openshell/credentials/key-encryption-key.bin"
```
For Kubernetes Secrets:
```toml
[openshell.gateway]
credential_drivers = ["kubernetes-secrets"]
[openshell.credential_drivers.kubernetes-secrets]
namespace = "openshell"
```
For Vault instead:
```toml
[openshell.gateway]
credential_drivers = ["vault"]
[openshell.credential_drivers.vault]
address = "http://vault.vault.svc.cluster.local:8200"
mount = "secret"
kv_version = "2"
auth_method = "kubernetes"
role = "openshell-gateway"
service_account_token_path = "/var/run/secrets/kubernetes.io/serviceaccount/token"
```
For the default encrypted database store, OpenShell stores provider credentials as JSON envelopes encrypted with AES-256-GCM in the gateway database. Each credential gets a random data-encryption key; the gateway wraps that key with a local key-encryption key. By default, the key-encryption key is created at `$XDG_STATE_HOME/openshell/gateway/credentials/key-encryption-key.bin` with owner-only permissions. Use `[openshell.gateway.credential_storage] key_encryption_key_env` instead of `key_encryption_key_path` to load a base64-encoded 32-byte key-encryption key from an environment variable. Back up the database and key-encryption key together; losing either makes stored credentials unrecoverable. In Kubernetes, the Helm chart creates a retained Secret containing the shared key-encryption key, injects it as `OPENSHELL_GATEWAY_CREDENTIAL_KEY_ENCRYPTION_KEY`, and renders `key_encryption_key_env` for the gateway when no external credential driver is enabled. Multi-replica deployments need every replica to use the same database and key-encryption key; the chart default handles the key-encryption key side.
For GitOps and `helm template` workflows where `lookup` returns empty, the chart-generated KEK Secret gets a random value on every render, making credentials unrecoverable. Set `server.credentialStorage.existingSecret` to the name of a pre-provisioned Secret containing the key-encryption key under the `key-encryption-key` data key. When set, the chart skips KEK Secret generation and references the provided Secret directly.
```yaml
server:
credentialStorage:
existingSecret: my-preprovisioned-kek-secret
```
For `kubernetes-secrets`, `namespace` sets where OpenShell-managed provider Secret objects are stored. When omitted, the driver uses the in-cluster ServiceAccount namespace when available, otherwise `default`. The Helm chart creates a Role granting the gateway access to all Secrets in the credential namespace because OpenShell-managed Secret names are dynamic SHA-256 hashes that cannot be restricted with `resourceNames`. Deploy credential Secrets in a dedicated namespace (`server.credentialDrivers.kubernetesSecrets.namespace`) to limit the RBAC blast radius.
For `vault`, `address` points at the Vault service, `mount` and `kv_version` describe the KV engine where OpenShell-managed provider secrets are stored, and `auth_method = "kubernetes"` logs in with the gateway Pod's ServiceAccount token. For local or development validation, use `auth_method = "token_file"` with `token_path = "/path/to/token"`. Do not put literal Vault tokens in TOML.
Provider records that already contain inline database credentials remain readable for upgrade compatibility. New provider create/update requests still submit credential values through the normal API, but the gateway stores those values through the active credential storage path and persists only handles. Before OpenShell 0.1.0, OpenShell does not automatically migrate inline refresh material or credential handles between drivers. Reconfigure refresh grants after an upgrade. Before changing credential drivers, remove affected credentials while the original driver is still available, then select the new driver and create them again. Do not run mixed gateway versions against the same refresh records.
For remote credential drivers, set `transport = "uds"` with `socket_path`. Omit `command`, `args`, and `startup_timeout_secs` when another service manager prestarts the driver socket. Keep backend tokens out of TOML; point the driver at mounted token files or native identity mechanisms instead.
The built-in `kubernetes-secrets` and `vault` drivers can also run out of
process over UDS. Set `command` to the standalone driver binary and pass
driver-specific settings through `args`; the gateway appends `--bind-socket
<socket_path>` when it launches the process.
```toml
[openshell.gateway]
credential_drivers = ["kubernetes-secrets"]
[openshell.credential_drivers.kubernetes-secrets]
transport = "uds"
socket_path = "/run/openshell/credential-drivers/kubernetes-secrets.sock"
command = "/usr/libexec/openshell/openshell-driver-kubernetes-secrets"
args = ["--namespace", "openshell"]
```
```toml
[openshell.gateway]
credential_drivers = ["vault"]
[openshell.credential_drivers.vault]
transport = "uds"
socket_path = "/run/openshell/credential-drivers/vault.sock"
command = "/usr/libexec/openshell/openshell-driver-vault"
args = [
"--address", "http://vault.vault.svc.cluster.local:8200",
"--auth-method", "kubernetes",
"--role", "openshell-gateway",
]
```
## Driver References
Each example is a complete TOML file for one compute driver. The examples repeat `[openshell]` and `[openshell.gateway]` so they stay copyable, and the driver tables list the accepted driver-specific keys. Driver-specific values override inherited gateway defaults. The gateway rejects unknown driver fields after inheritance is merged.
### Kubernetes
The gateway runs as a Pod and creates sandbox Pods in another namespace. mTLS material for sandboxes is delivered through a Kubernetes Secret rather than host-side file paths.
```toml
[openshell]
version = 1
[openshell.gateway]
bind_address = "0.0.0.0:8080"
health_bind_address = "0.0.0.0:8081"
metrics_bind_address = "0.0.0.0:9090"
log_level = "info"
compute_drivers = ["kubernetes"]
[openshell.gateway.tls]
cert_path = "/etc/openshell-tls/server/tls.crt"
key_path = "/etc/openshell-tls/server/tls.key"
client_ca_path = "/etc/openshell-tls/client-ca/ca.crt"
# When cert-manager serverIssuerRef is configured, these are populated by Helm:
# external_cert_path = "/etc/openshell-tls/server-external/tls.crt"
# external_key_path = "/etc/openshell-tls/server-external/tls.key"
# external_server_names = ["gateway.example.com"]
[openshell.drivers.kubernetes]
# Workspace isolation mode. "shared" renders all sandboxes into a single
# namespace. "managed" auto-creates a K8s namespace per workspace
# (openshell-{gateway_id}-{workspace}). "operator" maps each workspace to a
# pre-provisioned namespace discovered via label selector or drop-in file.
workspace_mode = "shared"
# Gateway identity used in managed-mode namespace naming. Defaults to the
# gateway JWT gateway_id. Must be a DNS-1123 label.
# gateway_id = "openshell"
namespace = "agents"
service_account_name = "openshell-sandbox"
default_image = "ghcr.io/nvidia/openshell/sandbox:latest"
image_pull_policy = "IfNotPresent"
image_pull_secrets = ["regcred"]
# Defaults to the gateway version; override to pin a specific build.
# supervisor_image = "ghcr.io/nvidia/openshell/supervisor:<version>"
supervisor_image_pull_policy = "IfNotPresent"
# Use the image volume on Kubernetes >= 1.35 (GA in 1.36); switch to "init-container"
# on older clusters or where the ImageVolume feature gate is off.
supervisor_sideload_method = "image-volume"
# "combined" runs the existing single supervisor container with full process,
# filesystem, and network enforcement in the agent container. "sidecar" moves
# pod-level network enforcement and gateway session handling into a network sidecar.
topology = "combined"
# Optional corporate HTTP forward proxy for policy-approved TLS egress. The
# sandbox workload cannot select or override these settings. Only http:// proxy
# endpoints and TLS CONNECT traffic are supported; plain HTTP egress remains
# direct. `no_proxy` bypasses only the corporate proxy, never OpenShell policy.
# https_proxy = "http://proxy.corp.example:8080"
# no_proxy = ".svc,.svc.cluster.local,10.96.0.0/12,10.244.0.0/16"
# Proxy credentials must be an existing Secret in the sandbox namespace. The
# key contains a `user:pass` value and is mounted only in the network
# supervisor container, never in workload environment or command arguments.
# proxy_auth_secret_name = "corporate-proxy-auth"
# proxy_auth_secret_key = "credentials"
# The gateway validates the Secret name/key syntax and their configuration
# relationship at startup; it does not read the Secret from the Kubernetes API.
# Kubernetes resolves the Secret when the Sandbox Pod starts. A missing key or
# Secret prevents that Pod from starting; unreadable or malformed `user:pass`
# content is validated fail-closed by the supervisor at startup and never
# falls back to direct egress.
# Proxy credential Secrets require `topology = "sidecar"`. Combined topology
# shares its credential mount with the workload and can make it readable by the
# sandbox group through Kubernetes `fsGroup` volume permission handling.
# Required with a credential Secret: Basic authentication to an http:// proxy
# is cleartext on the connection to that proxy.
# proxy_auth_allow_insecure = true
# Last resort for hostname-filtering proxy ACLs. The proxy resolves the target,
# so its ACL becomes part of the egress boundary for proxied connections.
# proxy_connect_by_hostname = true
grpc_endpoint = "https://openshell-gateway.agents.svc:8080"
ssh_socket_path = "/run/openshell/ssh.sock"
client_tls_secret_name = "openshell-client-tls"
host_gateway_ip = "10.0.0.1"
enable_user_namespaces = false
app_armor_profile = "Unconfined"
workspace_default_storage_size = "10Gi"
# Kubernetes StorageClass for the workspace PVC. Empty (default) omits the
# field, using the cluster's default StorageClass. Set this on clusters with no
# default StorageClass, otherwise the workspace PVC stays Pending.
# workspace_storage_class = "fast-ssd"
# Kubernetes RuntimeClass applied to sandbox pods when the API request does
# not specify one. Empty (default) = omit the field, using the cluster default.
# default_runtime_class_name = "kata-containers"
# Kubelet clamps projected tokens below 600 seconds. The driver caps values at 86400.
sa_token_ttl_secs = 3600
# Optional SPIFFE Workload API socket mounted into sandbox pods for dynamic
# provider token grants. Use an absolute path under a dedicated directory;
# shared roots such as /run, /var, /tmp, and /etc are rejected.
# Supervisor-to-gateway auth still uses gateway JWTs.
provider_spiffe_workload_api_socket_path = "/spiffe-workload-api/spire-agent.sock"
# Explicit sandbox UID/GID for the supervisor container securityContext and
# PVC init container. When unset, the driver auto-detects from OpenShift SCC
# namespace annotations (openshift.io/sa.scc.uid-range) if present, falling
# back to 1000 on non-OpenShift clusters. Any non-root Linux UID/GID is valid.
# sandbox_uid = 1500
# sandbox_gid = 1500
# Operator-mode namespace discovery. At least one must be set when
# workspace_mode = "operator". Both can be combined.
# operator_namespace_label discovers namespaces matching a K8s label selector.
# operator_namespace_label = "openshell.ai/workspace=true"
# operator_namespace_file reads allowed namespaces from a JSON/YAML file
# (hot-reloaded on change, e.g. via ConfigMap volume mount).
# operator_namespace_file = "/etc/openshell/workspace-namespaces.json"
[openshell.drivers.kubernetes.managed_ssh_ingress]
enabled = true
gateway_namespace = "openshell"
gateway_pod_selector = { "app.kubernetes.io/name" = "openshell", "app.kubernetes.io/instance" = "openshell" }
[openshell.drivers.kubernetes.sidecar]
# UID used by relaxed long-running network sidecars. Strict process/binary-aware
# sidecars run as UID 0 so Kubernetes grants the required /proc inspection
# capabilities into the effective set. In sidecar topology the network init
# container installs nftables rules that exempt the effective sidecar UID, so
# this dedicated infrastructure UID must remain at least 1000 and must not
# match the sandbox workload UID.
proxy_uid = 1337
# Keep process/binary-aware network policy enabled in sidecar topology. Set
# false to run the sidecar as proxy_uid, drop the sidecar's extra /proc
# inspection capabilities, and enforce endpoint/L7 policy without matching
# policy.binaries.
process_binary_aware_network_policy = true
```
In managed workspace mode, the Kubernetes driver copies each explicitly named
`image_pull_secrets` Secret from `namespace` into the managed workspace
namespace on sandbox creation. Shared and operator modes require the Secret to
already exist in the sandbox namespace.
For token-exchange provider profiles, the gateway also needs access to its own
SPIFFE Workload API socket. In Helm deployments, set
`server.providerTokenGrants.spiffe.enabled=true`; the chart mounts the socket
into the gateway pod and sets `OPENSHELL_GATEWAY_SPIFFE_WORKLOAD_API_SOCKET`.
The gateway verifies supervisor JWT-SVIDs with JWT bundles fetched from the
SPIFFE Workload API, so this validation path does not require gateway access to
the SPIRE OIDC discovery endpoint or its TLS CA.
### Docker
Sandboxes run as containers on a local bridge network. The supervisor binary is bind-mounted from the host (no in-cluster image pull required); guest mTLS material is supplied as host paths.
```toml
[openshell]
version = 1
[openshell.gateway]
bind_address = "127.0.0.1:17670"
log_level = "info"
compute_drivers = ["docker"]
[openshell.drivers.docker]
socket_path = "/var/run/docker.sock"
default_image = "ghcr.io/nvidia/openshell/sandbox:latest"
# Docker vocabulary: Always | IfNotPresent | Never. Empty behaves like IfNotPresent.
image_pull_policy = "IfNotPresent"
sandbox_namespace = "docker-dev"
# Empty auto-detects https://host.openshell.internal:<gateway-port> when guest TLS is set.
grpc_endpoint = "https://host.openshell.internal:17670"
# Skip the image-pull-and-extract step by pointing at a locally built binary.
supervisor_bin = "/usr/local/libexec/openshell/openshell-sandbox"
# When supervisor_bin is omitted, Docker extracts /openshell-sandbox from this image.
# Defaults to the gateway version; override to pin a specific build.
# supervisor_image = "ghcr.io/nvidia/openshell/supervisor:<version>"
guest_tls_ca = "/etc/openshell/certs/ca.pem"
guest_tls_cert = "/etc/openshell/certs/client.pem"
guest_tls_key = "/etc/openshell/certs/client-key.pem"
network_name = "openshell-docker"
host_gateway_ip = "172.17.0.1"
ssh_socket_path = "/run/openshell/ssh.sock"
# Unsafe operator override. Host bind mounts, including Docker local-driver
# bind-backed volumes, expose gateway-host paths inside sandboxes and can
# negate OpenShell isolation and filesystem controls.
enable_bind_mounts = false
# Set to 0 to leave Docker's runtime default unchanged.
sandbox_pids_limit = 2048
```
### Podman
Sandboxes run as Podman containers on a user-mode bridge network. The supervisor image is mounted read-only via Podman's `type=image` mount; guest mTLS material is supplied as host paths.
```toml
[openshell]
version = 1
[openshell.gateway]
bind_address = "127.0.0.1:17670"
log_level = "info"
compute_drivers = ["podman"]
[openshell.drivers.podman]
# Rootless socket path. For root Podman use /run/podman/podman.sock.
# Omit to auto-detect: the driver probes for a responsive Podman socket, then
# asks the podman CLI where its socket is, and fails to start if neither finds
# one. Set this to pin a specific Podman machine instead.
socket_path = "/run/user/1000/podman/podman.sock"
default_image = "ghcr.io/nvidia/openshell/sandbox:latest"
image_pull_policy = "missing" # always | missing | never | newer
grpc_endpoint = "https://host.containers.internal:17670"
# The gateway overwrites gateway_port from bind_address at runtime.
gateway_port = 17670
network_name = "openshell"
# Omit for the platform default: empty on Linux, 192.168.127.254 on macOS Podman machine.
# Set "" to force Podman's host-gateway resolver.
# host_gateway_ip = "192.168.127.254"
sandbox_ssh_socket_path = "/run/openshell/ssh.sock"
stop_timeout_secs = 45
# Defaults to the gateway version; override to pin a specific build.
# supervisor_image = "ghcr.io/nvidia/openshell/supervisor:<version>"
guest_tls_ca = "/etc/openshell/certs/ca.pem"
guest_tls_cert = "/etc/openshell/certs/client.pem"
guest_tls_key = "/etc/openshell/certs/client-key.pem"
# Unsafe operator override. Host bind mounts, including Podman local-driver
# bind-backed volumes, expose gateway-host paths inside sandboxes and can
# negate OpenShell isolation and filesystem controls.
enable_bind_mounts = false
# Set to 0 to leave Podman's runtime default unchanged.
sandbox_pids_limit = 2048
# Health check interval in seconds. Lower values detect readiness faster
# but increase process churn (each check spawns a conmon subprocess).
# Set to 0 to disable health checks entirely. Default: 10.
health_check_interval_secs = 10
# User namespace mode for sandbox containers. Omit to use the default.
# Supported modes: auto, host, keep-id, no-map, private.
# userns = "auto"
# Explicit UID/GID mappings for userns = "private". Each entry is
# "container_id:host_id:size". Required when mode is "private"; rejected
# for other modes. Rootless Podman uses intermediate IDs (0:0:1, 1:1:65535);
# rootful Podman uses absolute host IDs (0:1000:1, 1:100000:65536).
# uidmap = ["0:0:1", "1:1:65535"]
# gidmap = ["0:0:1", "1:1:65535"]
# Corporate forward proxy for sandbox egress. When set, the in-container
# supervisor chains policy-approved TLS tunnels through this proxy with HTTP
# CONNECT instead of dialing destinations directly. Plain-HTTP requests are
# not proxied and always dial the destination directly. http:// and https://
# proxy URLs in explicit scheme://host:port form are supported: the scheme and
# port are both required, and a URL carrying a path, query, or fragment is
# rejected rather than silently truncated. For an https:// proxy the supervisor
# wraps the connection to the proxy in TLS before the CONNECT handshake,
# verifying the proxy certificate against the built-in and system roots plus
# the optional proxy_ca_bundle below.
# NO_PROXY entries (hostnames, domain suffixes, IPs, CIDRs, each with an
# optional :port qualifier) are dialed directly. A port-qualified entry only
# bypasses that destination port. IP and CIDR entries also match hostnames
# through their validated DNS resolution; such a match dials directly only
# the resolved addresses inside the entry. This is an operator-owned egress
# boundary:
# sandbox and template environment cannot override it, and the conventional
# HTTPS_PROXY/HTTP_PROXY/NO_PROXY variables a sandbox sets do not affect it.
#
# The CONNECT request sent to the proxy targets a validated resolved IP,
# not the hostname, so the proxy performs no DNS resolution of its own and
# the tunnel stays bound to the address that passed the sandbox's SSRF and
# allowed_ips validation. The hostname still travels inside the tunnel (TLS
# SNI, application Host), so destination servers behave normally. In
# split-horizon networks, point the gateway host at the corporate resolver
# so internal names validate to their internal addresses. If the proxy's
# ACLs filter on hostnames and reject IP CONNECT targets, set
# proxy_connect_by_hostname = true as a last resort: the proxy then
# resolves the name itself, so a name that resolves differently at the
# proxy (split-horizon DNS, rebinding) can reach destinations the sandbox
# policy never approved, and the proxy's own ACLs become the effective
# egress control for proxied TLS.
#
# Configuration is fail-closed: an invalid proxy URL is rejected at gateway
# startup, setting no_proxy, proxy_auth_file, or proxy_connect_by_hostname
# without a proxy URL is rejected as well, and a set-but-invalid value
# reaching a sandbox (for example an unreadable or malformed auth file) is
# fatal to that sandbox's supervisor instead of silently falling back to
# direct or unauthenticated egress.
#
# Credentials must NOT be embedded in the URL (an inline user:pass@ is
# rejected at startup, since it would be stored here and exposed in container
# metadata). Instead point proxy_auth_file at a file containing "user:pass";
# the gateway delivers it to the supervisor through a root-only secret mount.
# The credential must use the user:pass form (non-empty user, no control
# characters); the same validation runs at sandbox-create time and in the
# supervisor, so a credential accepted here is never rejected in-container.
# Keep gateway.toml and the auth file owner-readable only (mode 0600).
#
# WARNING: with an http:// proxy the supervisor sends the credential as a
# Proxy-Authorization: Basic header over the plain-TCP connection. Basic
# auth is base64, not encryption: anyone on the network path between the
# sandbox host and the proxy can recover the credential. Because of that,
# proxy_auth_file with an http:// proxy requires the explicit
# acknowledgement proxy_auth_allow_insecure = true; without it the
# configuration is rejected at gateway startup. Only opt in when the path
# to the proxy is a trusted network segment. For an https:// proxy the
# credential is sent inside the verified TLS session, so the
# acknowledgement is not required (but tolerated if set).
#
# proxy_ca_bundle points at a PEM CA bundle (on the gateway host) trusted for
# the corporate proxy. A CA certificate is not secret, so the gateway
# bind-mounts it read-only into the sandbox. It is trusted in two places: the
# TLS handshake with an https:// proxy, and — because a TLS-intercepting proxy
# (mitmproxy, squid ssl-bump) re-signs tunneled server certificates with the
# same CA — the sandbox trust bundle and the supervisor's upstream
# re-encryption, so intercepted upstream TLS keeps working and sandbox
# workloads trust the re-signed certificates. It is valid with either an
# http:// or https:// proxy and requires a proxy URL; the file must exist and
# contain at least one certificate, or the sandbox fails closed at startup.
# https_proxy = "http://proxy.corp.com:8080"
# no_proxy = "*.svc.cluster.local,10.0.0.0/8"
# proxy_auth_file = "/etc/openshell/secrets/proxy-auth"
# proxy_auth_allow_insecure = true
# Last resort for hostname-filtering proxy ACLs; see the warning above.
# proxy_connect_by_hostname = true
# Corporate CA trusted for an https:// proxy and TLS-intercepting proxies.
# proxy_ca_bundle = "/etc/openshell/tls/proxy-ca.pem"
```
### MicroVM
Each sandbox runs inside its own libkrun microVM managed by the standalone `openshell-driver-vm` subprocess. Use this driver when you want stronger isolation than container namespaces alone.
```toml
[openshell]
version = 1
[openshell.gateway]
bind_address = "127.0.0.1:17670"
log_level = "info"
# VM is never auto-detected; an explicit entry here is required.
compute_drivers = ["vm"]
[openshell.drivers.vm]
state_dir = "/var/lib/openshell/vm"
# Where the gateway looks for the openshell-driver-vm subprocess binary.
driver_dir = "/usr/local/libexec/openshell"
default_image = "ghcr.io/nvidia/openshell/sandbox:latest"
grpc_endpoint = "https://host.containers.internal:17670"
# Empty falls back to default_image.
bootstrap_image = "ghcr.io/nvidia/openshell/sandbox:latest"
krun_log_level = 1
vcpus = 2
mem_mib = 2048
overlay_disk_mib = 4096
guest_tls_ca = "/var/lib/openshell/guest-tls/ca.pem"
guest_tls_cert = "/var/lib/openshell/guest-tls/client.pem"
guest_tls_key = "/var/lib/openshell/guest-tls/client-key.pem"
# Resolved sandbox UID/GID for the rootfs /etc/passwd entry.
# Defaults to 10001 when unset; matching GID is used if sandbox_gid is empty.
# Any non-root Linux UID/GID is valid.
# sandbox_uid = 20001
```
### Extension Driver
Extension drivers run outside the gateway and expose the
`compute_driver.proto` gRPC service on a Unix socket. Use a non-reserved driver
name; built-in names such as `vm`, `docker`, `podman`, and `kubernetes` cannot
be selected through unmanaged socket endpoints. The selected driver name is the
key used for driver-owned sandbox config such as `template.driver_config.<name>`.
```toml
[openshell]
version = 1
[openshell.gateway]
bind_address = "127.0.0.1:17670"
log_level = "info"
compute_drivers = ["kyma"]
[openshell.drivers.kyma]
socket_path = "/run/openshell/kyma-compute-driver.sock"
```
@@ -0,0 +1,591 @@
---
# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
title: "Policy Schema Reference"
sidebar-title: "Policy Schema"
description: "Complete field reference for the sandbox policy YAML including static and dynamic sections."
keywords: "Generative AI, Cybersecurity, Policy, Schema, YAML, Reference, Security"
position: 3
---
Complete field reference for the sandbox policy YAML. Each field is documented with its type, whether it is required, and whether it is static (locked at sandbox creation) or dynamic (hot-reloadable on a running sandbox).
## Top-Level Structure
A policy YAML file contains the following top-level fields:
```yaml showLineNumbers={false}
version: 1
filesystem_policy: { ... }
landlock: { ... }
process: { ... }
network_policies: { ... }
network_middlewares: { ... }
```
| Field | Type | Required | Category | Description |
|---|---|---|---|---|
| `version` | integer | Yes | -- | Policy schema version. Must be `1`. |
| `filesystem_policy` | object | No | Static | Controls which directories the agent can read and write. |
| `landlock` | object | No | Static | Configures Landlock LSM enforcement behavior. |
| `process` | object | No | Static | Sets the user and group the agent process runs as. |
| `network_policies` | map | No | Dynamic | Declares which binaries can reach which network endpoints. |
| `network_middlewares` | map | No | Dynamic | Attaches ordered middleware by destination host; each implementation's manifest selects its supported HTTP and WebSocket operations. |
Static fields are set at sandbox creation time. Changing them requires destroying and recreating the sandbox. Dynamic fields can be updated on a running sandbox with `openshell policy update` for incremental merges or `openshell policy set` for full replacement, and take effect without restarting.
## Version
The version field identifies which schema the policy uses:
| Field | Type | Required | Description |
|---|---|---|---|
| `version` | integer | Yes | Schema version number. Currently must be `1`. |
## Filesystem Policy
**Category:** Static
Controls filesystem access inside the sandbox. Paths not listed in either `read_only` or `read_write` are inaccessible.
| Field | Type | Required | Description |
|---|---|---|---|
| `include_workdir` | bool | No | When `true`, automatically adds the agent's working directory to `read_write`. |
| `read_only` | list of strings | No | Paths the agent can read but not modify. Typically system directories like `/usr`, `/lib`, `/etc`. |
| `read_write` | list of strings | No | Paths the agent can read and write. Typically `/tmp`; set `include_workdir: true` to add the driver-resolved working directory. |
**Validation constraints:**
- Every path must be absolute (start with `/`).
- Paths must not contain `..` traversal components. The server normalizes paths before storage, but rejects policies where traversal would escape the intended scope.
- Read-write paths must not be overly broad (for example, `/` alone is rejected).
- Each individual path must not exceed 4096 characters.
- The combined total of `read_only` and `read_write` paths must not exceed 256.
Policies that violate these constraints are rejected with `INVALID_ARGUMENT` at creation or update time. Disk-loaded YAML policies that fail validation fall back to a restrictive default.
Example:
```yaml showLineNumbers={false}
filesystem_policy:
include_workdir: true
read_only:
- /usr
- /lib
- /proc
- /dev/urandom
- /etc
read_write:
- /tmp
- /dev/null
```
## Landlock
**Category:** Static
Configures [Landlock LSM](https://docs.kernel.org/security/landlock.html) enforcement at the kernel level. Landlock provides mandatory filesystem access control below what UNIX permissions allow.
| Field | Type | Required | Values | Description |
|---|---|---|---|---|
| `compatibility` | string | No | `best_effort`, `hard_requirement` | How OpenShell handles Landlock failures. Refer to the behavior table below. |
**Compatibility modes:**
| Value | Kernel ABI unavailable | Individual path inaccessible | All paths inaccessible |
|---|---|---|---|
| `best_effort` | Warns and continues without Landlock. | Skips the path, applies remaining rules. | Warns and continues without Landlock (refuses to apply an empty ruleset). |
| `hard_requirement` | Aborts sandbox startup. | Aborts sandbox startup. | Aborts sandbox startup. |
`best_effort` (the default) is appropriate for most deployments. It handles missing paths gracefully. For example, `/app` might not exist in every container image but is included in the baseline path set for containers that do have it. Individual missing paths are skipped while the remaining filesystem rules are still enforced.
`hard_requirement` is for environments where any gap in filesystem isolation is unacceptable. If a listed path cannot be opened for any reason (missing, permission denied, symlink loop), sandbox startup fails immediately rather than running with reduced protection.
When a path is skipped under `best_effort`, the sandbox logs a warning that includes the path, the specific error, and a human-readable reason (for example, "path does not exist" or "permission denied").
Example:
```yaml showLineNumbers={false}
landlock:
compatibility: best_effort
```
## Process
**Category:** Static
Sets the OS-level identity for the agent process inside the sandbox.
| Field | Type | Required | Description |
|---|---|---|---|
| `run_as_user` | string | No | Overrides the user name or UID selected by the compute driver. Docker and Podman fall back to the image's OCI `USER`. |
| `run_as_group` | string | No | Overrides the group name or GID selected by the compute driver. Docker and Podman fall back to the image's OCI `USER`. |
**Validation constraint:** An explicit policy value must be `sandbox` or a
numeric UID/GID from `1` through `4294967294`. OpenShell rejects `0` as root
and `4294967295` as the invalid identity sentinel. Docker and Podman may select
other named identities through OCI `USER` fallback.
Omission is preserved independently for each field. For example, setting only
`run_as_user` keeps that explicit user while allowing the active driver to
select the group.
Example:
```yaml showLineNumbers={false}
process:
run_as_user: "1500"
run_as_group: "1500"
```
## Network Policies
**Category:** Dynamic
A map of named network policy entries. Each entry declares a set of endpoints and a set of binaries. Only the listed binaries are permitted to connect to the listed endpoints. The map key is a logical identifier. The `name` field inside the entry is the display name used in logs.
### Network Policy Entry
Each entry in the `network_policies` map has the following fields:
| Field | Type | Required | Description |
|---|---|---|---|
| `name` | string | No | Display name for the policy entry. Used in log output. Defaults to the map key. |
| `endpoints` | list of endpoint objects | Yes | Hosts and ports this entry permits. |
| `binaries` | list of binary objects | Yes | Executables allowed to connect to these endpoints. |
### Endpoint Object
Each endpoint defines a reachable destination and optional inspection rules.
| Field | Type | Required | Description |
|---|---|---|---|
| `host` | string | Conditional | Hostname or IP address. Required for `protocol: tcp`; transparent TCP requires a valid DNS hostname and rejects literal IPs. A non-TCP proxy endpoint may omit `host` only when `allowed_ips` supplies the destination constraint. Supports a `*` wildcard inside the first DNS label only: `*.example.com`, `**.example.com`, and intra-label patterns like `*-aiplatform.googleapis.com` are accepted; bare `*`/`**`, TLD wildcards (`*.com`), and wildcards outside the first label are rejected at load time. Prefer exact hosts for `protocol: tcp`: a wildcard authorizes DNS queries for all matching names and can provide a DNS-label exfiltration channel. |
| `port` | integer | Yes | TCP port number. |
| `path` | string | No | Optional HTTP path glob used to select between L7 endpoints that share the same host and port. Empty means all paths. Use this when REST and GraphQL live under the same host, such as `/repos/**` and `/graphql`. |
| `protocol` | string | No | Set to `tcp` with a valid DNS hostname to allow native TCP clients through policy DNS and transparent capture without payload inspection. Omit the field for L4 passthrough through an explicit proxy, including legacy hostless `allowed_ips` endpoints. Set to `rest` for HTTP method/path inspection, `websocket` for RFC 6455 upgrade and client text-message inspection, `graphql` for GraphQL-over-HTTP operation inspection, `mcp` for MCP Streamable HTTP request inspection, or `json-rpc` for generic JSON-RPC-over-HTTP method inspection. WebSocket endpoints can also use GraphQL operation rules for GraphQL-over-WebSocket traffic. Provider-credentialed endpoints require an inspected protocol unless `allow_uninspected_credentials` is explicitly set. |
| `tls` | string | No | TLS handling mode. The proxy auto-detects TLS by peeking the first bytes of each connection and terminates it for inspected HTTPS traffic, so this field is optional in most cases. Set to `skip` to disable auto-detection for edge cases such as client-certificate mTLS or non-standard protocols. Provider-credentialed endpoints reject `tls: skip` unless `allow_uninspected_credentials` is explicitly set. The values `terminate` and `passthrough` are deprecated and log a warning; they are still accepted for backward compatibility but have no effect on behavior. |
| `enforcement` | string | No | `enforce` actively blocks disallowed requests. `audit` logs violations but allows traffic through. |
| `access` | string | No | Access preset. One of `read-only`, `read-write`, or `full`. Mutually exclusive with `rules`. Not valid on `protocol: mcp` or `protocol: json-rpc`; MCP uses explicit rules unless `mcp.allow_all_known_mcp_methods: true` enables the endpoint method profile, and JSON-RPC always uses explicit rules. |
| `rules` | list of allow rule objects | No | Fine-grained protocol-specific allow rules. Mutually exclusive with `access`. |
| `deny_rules` | list of deny rule objects | No | L7 deny rules that block specific requests even when allowed by `access` or `rules`. Deny rules take precedence over allow rules. |
| `allowed_ips` | list of string | No | CIDR or IP allowlist for SSRF override. Exact user-declared hostname endpoints may resolve to RFC 1918 private addresses without this field, but wildcard, hostless, and policy-advisor-proposed endpoints still require `allowed_ips` for private resolved IPs. A hostless allowlist is valid only for the legacy proxy path and cannot be combined with `protocol: tcp`. Entries overlapping loopback (`127.0.0.0/8`), link-local (`169.254.0.0/16`), or unspecified (`0.0.0.0`) are rejected at load time. |
| `allow_encoded_slash` | bool | No | When `true`, L7 request parsing preserves `%2F` inside path segments instead of rejecting it. Use this for registries and APIs such as npm scoped packages (`/@scope%2Fname`). Defaults to `false`. |
| `websocket_credential_rewrite` | bool | No | When `true` on a `protocol: rest` or `protocol: websocket` endpoint, OpenShell rewrites credential placeholders in client-to-server WebSocket text messages after an allowed HTTP `101` upgrade. On provider-credentialed endpoints without `allow_uninspected_credentials`, OpenShell uses the parsed relay and rejects binary frames; text frames containing placeholders fail closed when rewrite is disabled. Defaults to `false`. |
| `request_body_credential_rewrite` | bool | No | When `true` on a `protocol: rest` endpoint, OpenShell rewrites credential placeholders in UTF-8 `application/json`, `application/x-www-form-urlencoded`, and `text/*` request bodies before forwarding upstream. The proxy buffers at most 256 KiB and updates `Content-Length` after rewriting. For chunked requests, the limit counts framing, extensions, and trailers. When rewrite is disabled and the sandbox has provider credentials, ordinary bodies continue to stream, but a reserved credential placeholder is rejected before its marker reaches upstream, including for providers without endpoint profiles. Defaults to `false`. Mutually exclusive with `credential_signing`. |
| `allow_uninspected_credentials` | bool | No | Explicit security-sensitive opt-in that permits a provider-credentialed endpoint to use traffic paths OpenShell cannot inspect or rewrite, including L4-only and `tls: skip` tunnels. Defaults to `false`. Policy proposals that set it require explicit security-flagged approval. |
| `credential_signing` | string | No | Proxy-side credential signing mode. When set, the proxy strips the sandbox client's `Authorization` header and re-signs with real provider credentials. Values: `sigv4` (auto-detect payload mode from client headers), `sigv4:body` (buffer and hash body, max 10 MiB), `sigv4:no_body` (unsigned payload, stream body). Mutually exclusive with `request_body_credential_rewrite`. See [AWS SigV4](/providers/aws-sigv4). |
| `signing_service` | string | No | AWS service name for SigV4 signing (e.g. `bedrock`, `s3`, `sts`). Required when `credential_signing` is set. |
| `signing_region` | string | No | AWS region override for SigV4 signing (e.g. `us-east-1`). When omitted, the region is extracted from the endpoint hostname. Required for non-standard AWS endpoints where the region cannot be inferred. |
| `credential_binding` | object | No | Binds static credentials from an attached provider to this endpoint when that provider's profile defines no endpoints. This field is valid only in a sandbox-scoped policy. |
| `credential_binding.provider` | string | Yes with `credential_binding` | Exact name of the provider instance attached to the sandbox. The referenced provider must have a profile, and that profile must define no endpoints. |
| `persisted_queries` | string | No | GraphQL hash-only behavior for `protocol: graphql` and GraphQL-over-WebSocket operation policy. Default is `deny`; use `allow_registered` only with `graphql_persisted_queries`. |
| `graphql_persisted_queries` | map | No | Trusted GraphQL persisted-query registry keyed by hash or saved-query ID. Values contain `operation_type`, optional `operation_name`, and optional root `fields`. |
| `graphql_max_body_bytes` | integer | No | Maximum GraphQL-over-HTTP request body bytes buffered for inspection. Defaults to `65536`. |
| `mcp` | object | No | MCP endpoint options for `protocol: mcp`. MCP endpoints must set a concrete `host` and `port` or `ports`; `protocol: mcp` alone is invalid and is not treated as a wildcard endpoint. |
| `mcp.max_body_bytes` | integer | No | Maximum MCP JSON-RPC-over-HTTP request body bytes buffered for inspection. Defaults to `65536`. |
| `mcp.strict_tool_names` | bool | No | Defaults to `true`. Requires `tools/call` `params.name` values to match `^[A-Za-z0-9_.-]{1,128}$` before policy evaluation. Set to `false` only for compatibility with MCP servers that intentionally use non-recommended tool names. Wildcard `tool` matchers require this to remain enabled. |
| `mcp.allow_all_known_mcp_methods` | bool | No | Defaults to `false`. When `true`, enables the endpoint MCP method profile: omitted `rules` allow all MCP-family methods and all tools before `deny_rules`, and omitted rule `method` uses that profile. When unset or `false`, explicit MCP method rules are required; rules with `tool` or `params.name` must set `method: tools/call`. |
| `json_rpc` | object | No | JSON-RPC endpoint options. For `protocol: json-rpc`, `json_rpc.max_body_bytes` sets the maximum JSON-RPC-over-HTTP request body bytes buffered for inspection. Defaults to `65536`. |
**Validation constraints:**
- `access` and `rules` are mutually exclusive; setting both is rejected.
- `protocol: tcp` requires a valid DNS hostname. Hostless `allowed_ips`, IP-literal hosts, trailing-dot names, and malformed DNS selectors are rejected with a policy-validation error.
- `protocol: tcp` requires at least one port and rejects L7-only fields, including `path`, `enforcement`, `access`, `rules`, `deny_rules`, request rewriting and credential signing fields, and GraphQL, JSON-RPC, or MCP options.
- A sandbox runtime must support policy DNS and transparent TCP capture before it can activate a policy containing `protocol: tcp`. Docker and Podman provide this runtime support.
- Adding the first `protocol: tcp` endpoint to a running sandbox that started without one is rejected atomically because its DNS and capture substrate is startup infrastructure. Recreate the sandbox with a TCP endpoint. A sandbox that started with the substrate can remove and re-add TCP endpoints dynamically.
- When `protocol` is set, at least one of `access` or `rules` is required for `rest`, `websocket`, `graphql`, and `sql`.
- `mcp` and `json-rpc` reject `access` presets; use explicit `rules`.
- `json-rpc` requires explicit `rules` with `allow.method`.
- `mcp` requires `rules` unless `mcp.allow_all_known_mcp_methods: true`.
- `deny_rules` require `protocol`. For non-MCP protocols, `deny_rules` also require `rules` or `access` to define the base allow set. MCP `deny_rules` may omit both when `mcp.allow_all_known_mcp_methods: true` supplies the base allow set.
- `rules: []` (empty list) is rejected; use `access: full` or remove `rules`.
- Non-empty `rules` must contain at least one effective allow clause; rules where every entry lacks an allow are rejected as deny-all.
- `deny_rules: []` (empty list) is rejected; remove it if no denials are needed.
- `credential_signing` requires a resolvable AWS credential source before a sandbox policy can activate. Use an attached endpoint-bearing profile that declares `AWS_ACCESS_KEY_ID` and `AWS_SECRET_ACCESS_KEY` and covers the signed endpoint, or bind an attached endpointless profile that declares those keys with `credential_binding.provider`.
Credential rewrite recognizes the canonical `openshell:resolve:env:KEY` placeholder form and whole-token provider-shaped aliases such as `provider-OPENSHELL-RESOLVE-ENV-API_TOKEN` when the referenced environment key exists in the configured provider credentials.
Static provider placeholders also require the request host, port, and path to
match their credential binding. Profile endpoints supply this boundary by
default. An endpointless profile can instead use a sandbox policy endpoint with
`credential_binding.provider` set to the exact attached provider name. OpenShell
rejects unattached providers, profileless providers, endpointful profiles, and
global policies that use this field. Network policy admission does not expand
the credential boundary unless the endpoint explicitly supplies this binding.
OpenShell rejects a request mismatch with HTTP 403 and
`credential_endpoint_mismatch`. Refer to [Static Credential Endpoint
Binding](/sandboxes/providers-v2#understand-static-credential-endpoint-binding).
This example allows the sandbox to reach Google Cloud Storage and binds the
static credentials from the attached `work-gcp` provider to that endpoint:
```yaml showLineNumbers={false}
network_policies:
gcp_storage:
endpoints:
- host: storage.googleapis.com
port: 443
protocol: rest
access: full
credential_binding:
provider: work-gcp
```
#### Access Levels
The `access` field accepts one of the following values on REST, WebSocket, and GraphQL endpoints. MCP and JSON-RPC endpoints reject `access` because HTTP method/path presets cannot authorize JSON-RPC safely. Use explicit MCP rules, set `mcp.allow_all_known_mcp_methods: true` for the MCP method profile, or use explicit JSON-RPC rules.
| Value | REST expansion | WebSocket expansion | GraphQL expansion | MCP / JSON-RPC expansion |
|---|---|---|---|---|
| `full` | All methods and paths. | WebSocket upgrade and all inspected client text-message paths. | All operation types. | Rejected. |
| `read-only` | `GET`, `HEAD`, `OPTIONS`. | WebSocket upgrade handshake only. | `query` operations. | Rejected. |
| `read-write` | `GET`, `HEAD`, `OPTIONS`, `POST`, `PUT`, `PATCH`. | WebSocket upgrade handshake and client text messages. | `query` and `mutation` operations. | Rejected. |
For MCP endpoints, configure explicit `rules` with `method`, optional `tool`, and supported `params`. For generic JSON-RPC endpoints, configure explicit `rules` with `method`; JSON-RPC policy `params` matchers are not presently supported.
#### Allow Rule Objects
Used when `access` is not set. Each entry in `rules` contains an `allow` object. The tables below list the fields inside that `allow` object.
##### REST Allow Rule (`protocol: rest`)
REST allow rules match HTTP requests by method, path, and optional query parameters.
| Field | Type | Required | Description |
|---|---|---|---|
| `method` | string | Yes | HTTP method to allow (for example, `GET`, `POST`). `*` matches any method. |
| `path` | string | Yes | URL path glob. `*` and `**` match zero or more characters and may cross `/`; `?` matches one character; bracket classes such as `[0-9]` and `[!0]` are supported. |
| `query` | map | No | Query parameter matchers keyed by decoded param name. Matcher value can be a glob string (`tag: "foo-*"`) or an object with `any` (`tag: { any: ["foo-*", "bar-*"] }`). |
Example REST allow rules:
```yaml showLineNumbers={false}
rules:
- allow:
method: GET
path: /**/info/refs*
query:
service: "git-*"
- allow:
method: POST
path: /**/git-upload-pack
query:
tag:
any: ["v1.*", "v2.*"]
```
##### WebSocket Allow Rule (`protocol: websocket`)
WebSocket allow rules match the RFC 6455 HTTP upgrade by path and match client-to-server text messages on the same upgraded connection with the synthetic `WEBSOCKET_TEXT` method. Binary frames are relayed but are not rewritten.
| Field | Type | Required | Description |
|---|---|---|---|
| `method` | string | Yes | `GET` allows the upgrade handshake, `WEBSOCKET_TEXT` allows client text messages after upgrade, and `*` matches both inspected actions. |
| `path` | string | Yes | URL path pattern from the original upgrade request. Supports `*` and `**` glob syntax. |
| `query` | map | No | Query parameter matchers from the original upgrade request. Matcher syntax is the same as REST allow rules. |
Example WebSocket allow rules:
```yaml showLineNumbers={false}
rules:
- allow:
method: GET
path: /v1/realtime/**
- allow:
method: WEBSOCKET_TEXT
path: /v1/realtime/**
```
##### GraphQL Allow Rule (`protocol: graphql` or GraphQL-over-WebSocket)
GraphQL allow rules match parsed GraphQL operations by operation type, optional operation name, and optional root fields. On `protocol: graphql`, they apply to GraphQL-over-HTTP `GET` and `POST` requests. On `protocol: websocket`, include a separate `GET` allow rule for the RFC 6455 upgrade, then use GraphQL allow rules for client operation messages using the `graphql-transport-ws` `subscribe` message type or the legacy `graphql-ws` `start` message type.
| Field | Type | Required | Description |
|---|---|---|---|
| `operation_type` | string | Yes | GraphQL operation type: `query`, `mutation`, `subscription`, or `*`. |
| `operation_name` | string | No | GraphQL operation-name glob. Omit to match any operation name. |
| `fields` | list of string | No | GraphQL root-field globs. Every selected root field must match one configured glob. Omit to match all root fields. |
Example GraphQL allow rules:
```yaml showLineNumbers={false}
rules:
- allow:
operation_type: query
fields: [viewer, repository]
- allow:
operation_type: mutation
operation_name: Issue*
fields: [createIssue]
```
Example GraphQL-over-WebSocket allow rules:
```yaml showLineNumbers={false}
rules:
- allow:
method: GET
path: /graphql
- allow:
operation_type: subscription
fields: [messageAdded]
- allow:
operation_type: query
fields: [viewer]
```
Do not combine `method`, `path`, or `query` with `operation_type`, `operation_name`, or `fields` inside the same WebSocket rule. When a WebSocket endpoint has GraphQL operation policy, use GraphQL rules for client messages instead of a raw `WEBSOCKET_TEXT` allow rule.
##### MCP Allow And Deny Rules (`protocol: mcp`)
MCP rules match sandbox-to-server MCP Streamable HTTP request bodies by MCP method and optional tool selectors. OpenShell parses the underlying JSON-RPC 2.0 envelope, validates known MCP request and notification params, and preserves unknown extension methods as policy-addressable literal method strings. Until OpenShell exposes explicit MCP version profiles, `mcp.allow_all_known_mcp_methods` defaults to `false`, so endpoints require explicit MCP method rules. Set `mcp.allow_all_known_mcp_methods: true` to enable the endpoint method profile; in that mode, rules can omit `method`, and tool selectors are normalized to `tools/call` internally. By default, `tools/call` `params.name` must match the MCP-recommended tool-name pattern `^[A-Za-z0-9_.-]{1,128}$`; configure `mcp.strict_tool_names: false` on the endpoint only to allow a server that intentionally uses names outside that pattern. Wildcard `tool` matchers require `mcp.strict_tool_names` to remain enabled. JSON-RPC responses and server-to-client MCP messages on response bodies or SSE streams are relayed but are not currently parsed for policy enforcement.
Use `rules` for MCP allow rules and `deny_rules` for MCP deny rules. Deny rules take precedence over allow rules. If an MCP endpoint sets `mcp.allow_all_known_mcp_methods: true` and omits `rules`, OpenShell allows all MCP-family methods and all tools, then applies any `deny_rules`. Otherwise, the endpoint must define explicit rules. A broad allow or deny rule whose method matcher includes `tools/call` cannot be combined with tool-specific allow rules because it would bypass or erase the tool filter; add `tool` or `params.name` to scope `tools/call`, or remove the tool-specific rules. In a batch request, one denied call denies the full batch.
| Field | Type | Required | Description |
|---|---|---|---|
| `method` | string | No | MCP method name, such as `initialize`, `tools/list`, `tools/call`, or an unknown extension method. Globs are accepted only for the `tools/` method family, such as `tools/*`. Required unless `mcp.allow_all_known_mcp_methods` is `true`; when that option is true, omitted method uses the endpoint method profile. Do not use `method: "*"` for MCP; omit `method` only when using the allow-all MCP method profile. |
| `tool` | string or matcher | No | Convenience matcher for `tools/call` `params.name`. Supports a glob string or `{ any: [...] }`. Requires `method: tools/call` unless `mcp.allow_all_known_mcp_methods` is `true`; validation fails otherwise. Omit to match every tool. |
| `params` | map | No | MCP currently accepts only `params.name` as a lower-level tool-name matcher. Requires `method: tools/call` unless `mcp.allow_all_known_mcp_methods` is `true`; validation fails otherwise. Tool argument matching is not supported yet; allowed tools accept all argument payloads by default. |
An MCP client first sends `initialize`. After the server returns a successful response, the client sends `notifications/initialized`. After initialization completes and the server advertises the `tools` capability, the client can call an advertised tool. The response does not need an allow rule because these rules inspect messages sent from the client to the server. This example adds both client initialization messages to the existing tool rules. It omits `tools/list` because it assumes the client already knows the tool names; add that method when the client performs discovery.
```yaml showLineNumbers={false}
endpoints:
- host: mcp.example.com
port: 443
path: /mcp
protocol: mcp
enforcement: enforce
mcp:
max_body_bytes: 131072
rules:
- allow:
method: initialize
- allow:
method: notifications/initialized
- allow:
method: tools/call
tool: search_web
- allow:
method: tools/call
tool:
any: [create_issue, list_issues]
deny_rules:
- method: tools/call
tool: send_email
- method: tools/call
tool: execute_code
```
##### JSON-RPC Allow Rule (`protocol: json-rpc`)
JSON-RPC allow rules match sandbox-to-server JSON-RPC-over-HTTP request objects by RPC method. They apply to single JSON-RPC requests and batch requests. For a batch, OpenShell evaluates each call independently. Client-to-server JSON-RPC response frames in POST bodies are denied. Server-to-client messages on HTTP response bodies or MCP SSE streams are relayed but are not currently parsed for policy enforcement.
| Field | Type | Required | Description |
|---|---|---|---|
| `method` | string | Yes | Exact JSON-RPC method name such as `initialize` or `reports.search`. Use `*` only as the allow-all sentinel. Other wildcard or glob patterns are rejected for generic JSON-RPC endpoints. |
Generic JSON-RPC policy `params` matchers are not supported. Allow rules match only the JSON-RPC method.
Example JSON-RPC allow rules:
```yaml showLineNumbers={false}
endpoints:
- host: jsonrpc.example.com
port: 443
path: /rpc
protocol: json-rpc
enforcement: enforce
json_rpc:
max_body_bytes: 131072
rules:
- allow:
method: initialize
- allow:
method: reports.list
- allow:
method: reports.search
```
#### Deny Rule Objects
Blocks specific operations on endpoints that otherwise have broad access. Deny rules are evaluated after allow rules and take precedence: if a request matches any deny rule, it is blocked regardless of what the allow rules or access preset permit.
##### REST Deny Rule (`protocol: rest`)
REST deny rules use the same field names as REST allow rules, but they appear directly under each `deny_rules` entry instead of under an `allow` wrapper.
| Field | Type | Required | Description |
|---|---|---|---|
| `method` | string | Yes | HTTP method to deny (for example, `POST`, `DELETE`). `*` matches any method. |
| `path` | string | Yes | URL path pattern. Same glob syntax as allow rules. Use `**` to match any path. |
| `query` | map | No | Query parameter matchers. Same syntax as allow rule `query`. |
Example REST deny rules:
```yaml showLineNumbers={false}
endpoints:
- host: api.github.com
port: 443
protocol: rest
enforcement: enforce
access: read-write
deny_rules:
- method: POST
path: "/repos/*/pulls/*/reviews"
- method: PUT
path: "/repos/*/branches/*/protection"
- method: "*"
path: "/repos/*/rulesets"
```
##### WebSocket Deny Rule (`protocol: websocket`)
WebSocket deny rules use the same field names as WebSocket allow rules, but they appear directly under each `deny_rules` entry instead of under an `allow` wrapper.
| Field | Type | Required | Description |
|---|---|---|---|
| `method` | string | Yes | `GET` denies matching upgrade handshakes, `WEBSOCKET_TEXT` denies matching client text messages after upgrade, and `*` matches both inspected actions. |
| `path` | string | Yes | URL path pattern from the original upgrade request. Same glob syntax as allow rules. |
| `query` | map | No | Query parameter matchers from the original upgrade request. Same syntax as allow rule `query`. |
Example WebSocket deny rules:
```yaml showLineNumbers={false}
endpoints:
- host: realtime.example.com
port: 443
protocol: websocket
enforcement: enforce
access: read-write
deny_rules:
- method: WEBSOCKET_TEXT
path: "/v1/admin/**"
```
##### GraphQL Deny Rule (`protocol: graphql` or GraphQL-over-WebSocket)
GraphQL deny rules use the same field names as GraphQL allow rules, but they appear directly under each `deny_rules` entry instead of under an `allow` wrapper. On WebSocket GraphQL endpoints, they apply only to classified GraphQL operation messages; protocol lifecycle messages such as `connection_init`, `ping`, `pong`, and `complete` are allowed as WebSocket control-plane messages and are not payload-logged.
| Field | Type | Required | Description |
|---|---|---|---|
| `operation_type` | string | Yes | GraphQL operation type to deny: `query`, `mutation`, `subscription`, or `*`. |
| `operation_name` | string | No | GraphQL operation-name glob. |
| `fields` | list of string | No | GraphQL root-field globs. Any matching root field blocks the request. Omit to deny every operation that matches `operation_type` and `operation_name`. |
Example GraphQL deny rules:
```yaml showLineNumbers={false}
endpoints:
- host: api.github.com
port: 443
protocol: graphql
enforcement: enforce
access: read-write
deny_rules:
- operation_type: mutation
fields: [deleteRepository]
- operation_type: mutation
operation_name: Admin*
```
##### JSON-RPC Deny Rule (`protocol: json-rpc`)
JSON-RPC deny rules use the same field names as JSON-RPC allow rules, but they appear directly under each `deny_rules` entry instead of under an `allow` wrapper. Deny rules take precedence over allow rules. In a batch request, one denied call denies the full batch.
| Field | Type | Required | Description |
|---|---|---|---|
| `method` | string | Yes | Exact JSON-RPC method name to deny, or `*` to deny all JSON-RPC methods. Other wildcard or glob patterns are rejected for generic JSON-RPC endpoints. |
JSON-RPC deny rules do not support policy `params` matchers yet. Use MCP `tool` rules for MCP tool calls, or deny generic JSON-RPC methods by `method`.
Example JSON-RPC deny rules:
```yaml showLineNumbers={false}
endpoints:
- host: jsonrpc.example.com
port: 443
path: /rpc
protocol: json-rpc
enforcement: enforce
rules:
- allow:
method: "*"
deny_rules:
- method: reports.delete
```
### Binary Object
Identifies an executable that is permitted to use the associated endpoints.
| Field | Type | Required | Description |
|---|---|---|---|
| `path` | string | Yes | Filesystem path to the executable. Supports glob patterns with `*` and `**`. For example, `/sandbox/.vscode-server/**` matches any executable under that directory tree. |
## Network Middleware
**Category:** Dynamic
A map of up to 10 middleware configs selected after network and L7 policy admit an HTTP request or WebSocket upgrade. Each map key is the stable policy-local config identity. Middleware selection is independent of the network policy entry that admitted the traffic. Every matching config runs once by ascending `order` before provider credential injection. WebSocket-capable bindings continue on client text messages after the upgrade. Order values must be unique across the policy, and runtime selection also enforces the 10-stage maximum.
```yaml showLineNumbers={false}
network_middlewares:
regex-redactor:
name: Redact API tokens
middleware: openshell/regex
order: 10
config:
mode: redact
on_error: fail_closed
endpoints:
include: ["*.example.com"]
exclude: ["trusted.example.com"]
```
| Field | Type | Required | Description |
|---|---|---|---|
| `name` | string | No | Human-readable name for the middleware config. Defaults to the map key, which remains its stable identity. |
| `middleware` | string | Yes | Built-in middleware name or operator-owned registration name. `openshell/` is reserved for built-ins. |
| `order` | integer | No | Execution priority. Lower values run first, and values must be unique across the policy. Defaults to `0`; therefore, policies with multiple configs normally specify it explicitly. |
| `config` | object | No | Implementation-owned configuration validated by the selected middleware. |
| `on_error` | string | No | Applies only after an advertised operation binding is selected. `fail_closed` denies the HTTP request or closes the WebSocket when that stage fails; `fail_open` skips a failed HTTP stage or disables a broken WebSocket stage for the rest of that connection. Defaults to `fail_closed`. |
| `endpoints` | object | Yes | Host selector with required non-empty `include` and optional `exclude` lists, limited to 32 combined patterns. Exclusions take precedence. |
Host selectors use the same case-insensitive exact and DNS glob semantics as network endpoints: `*` matches exactly one DNS label and `**` matches one or more labels, so `**.example.com` covers subdomains but not `example.com` itself. Brace alternates are rejected at validation. A matching attachment joins only the operation chains its implementation advertises. An HTTP-only attachment may inspect a WebSocket upgrade GET without joining the post-upgrade chain; OpenShell permits the messages and records `binding_not_selected` coverage regardless of `on_error`. WebSocket bindings inspect complete client text messages. Binary messages pass with `unsupported_message_type` coverage for active stages. A fail-closed selector that can cover a `tls: skip` endpoint is rejected because OpenShell cannot inspect that traffic through any operation. An all-`fail_open` match may cover the endpoint; the supervisor bypasses the middleware and emits a detection finding.
See [Supervisor Middleware](/extensibility/supervisor-middleware) for registration, failure behavior, body limits, and operational guidance.
## Full Example
The following policy grants read-only GitHub API access and npm registry access:
```yaml showLineNumbers={false}
network_policies:
github_rest_api:
name: github-rest-api
endpoints:
- host: api.github.com
port: 443
protocol: rest
enforcement: enforce
access: read-only
binaries:
- path: /usr/local/bin/claude
- path: /usr/bin/node
- path: /usr/bin/gh
npm_registry:
name: npm-registry
endpoints:
- host: registry.npmjs.org
port: 443
protocol: rest
access: read-only
allow_encoded_slash: true
binaries:
- path: /usr/bin/npm
- path: /usr/bin/node
```
@@ -0,0 +1,585 @@
---
# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
title: "Sandbox Compute Drivers"
sidebar-title: "Compute Drivers"
description: "Reference for Docker, Podman, MicroVM, and Kubernetes sandbox compute drivers."
keywords: "Generative AI, Cybersecurity, AI Agents, Sandboxing, Docker, Podman, MicroVM, Kubernetes, Reference"
position: 4
---
The gateway's configured compute driver determines how OpenShell creates each sandbox. The CLI workflow stays the same across drivers: you create, connect to, inspect, stop, start, and delete sandboxes through the gateway API.
Every compute driver runs the OpenShell supervisor inside the sandbox workload. The supervisor launches the agent process, applies policy, routes egress through the proxy, injects configured credentials, and maintains the gateway session.
Stop stops compute but retains the sandbox record and the driver's
persistent workspace boundary. Start reactivates the same driver resource.
Delete remains independent and removes compute plus driver-owned persistent
state. While a sandbox is stopped, gateway access paths and exposed services
remain unavailable.
Restarting the gateway preserves this intent. The gateway does not stop Docker
or Podman containers during shutdown. At startup it sends idempotent start
requests for Docker, Podman, and MicroVM sandboxes that were intended to run;
already-running resources are unchanged, retained stopped compute is restarted,
and explicitly stopped sandboxes remain stopped. Kubernetes workloads continue
running independently of the gateway process.
The gateway forwards one exact, persisted main-process specification to every
driver. Drivers serialize that specification in
`OPENSHELL_MAIN_PROCESS_SPEC`; they do not install an idle `sleep` workload or
reconstruct argv with shell parsing. Runtime restart policies are disabled so
an exited canonical process remains a terminal sandbox error.
## Configure a Compute Driver
Configure the compute driver on the gateway. Current releases accept one driver per gateway. Set `compute_drivers` in the gateway TOML file:
```toml
[openshell.gateway]
compute_drivers = ["docker"]
```
Reserved built-in values are `docker`, `podman`, `kubernetes`, and `vm`.
Non-reserved names select an extension driver and require a
`socket_path` in `[openshell.drivers.<name>]`.
When `compute_drivers` is unset, the gateway auto-detects Kubernetes, then Podman, then Docker. Local container runtimes must respond to an API probe before the gateway selects them. The VM driver is never auto-detected; configure it explicitly with `compute_drivers = ["vm"]` or set `OPENSHELL_DRIVERS=vm` in the launch environment.
Common gateway options:
| Gateway TOML option | Description |
|---|---|
| `compute_drivers = ["<driver>"]` | Select the compute driver. Built-in values are `docker`, `podman`, `kubernetes`, and `vm`; custom names require `[openshell.drivers.<name>].socket_path`. |
Set driver-specific values such as sandbox images, callback endpoints, network names, TLS material, and VM sizing in the gateway TOML file. See the [Gateway Configuration File](./gateway-config) reference for the full `[openshell.drivers.<name>]` schema.
Extension drivers use the same `compute_driver.proto` gRPC surface as the
managed VM driver. For an out-of-tree driver, choose a driver name and point
the gateway at the Unix socket the operator has already provisioned:
```toml
[openshell.gateway]
compute_drivers = ["kyma"]
[openshell.drivers.kyma]
socket_path = "/run/openshell/kyma.sock"
```
For a launch-time socket override, pass the selected driver name with the
socket path. The endpoint replaces normal driver construction for that name,
including canonical built-in names:
```shell
openshell-gateway --drivers kyma --compute-driver-socket /run/openshell/kyma.sock
openshell-gateway --drivers docker --compute-driver-socket /run/openshell/docker.sock
```
The gateway connects to the operator-provided endpoint; it does not provision
or supervise the remote driver. The operator must protect the socket so only
the gateway uid can access it.
Sandbox create supports `--cpu` and `--memory` for per-sandbox compute sizing.
Docker and Podman apply them as runtime limits. Kubernetes applies them as both
container requests and limits. The VM driver accepts the fields but currently
ignores them.
Sandbox create also accepts experimental driver-owned config through
`--driver-config-json`. The value is a JSON object keyed by driver name. The
gateway forwards only the block for the active driver, so a Kubernetes gateway
receives the `kubernetes` object from a value such as:
Nested keys inside each driver block use snake_case. The top-level envelope keys
are driver names, such as `kubernetes`, and are not part of the nested schema.
```shell
openshell sandbox create \
--driver-config-json '{"kubernetes":{"pod":{"runtime_class_name":"kata-containers","priority_class_name":"batch-low"}}}' \
-- claude
```
Driver config is for fields without a stable public flag. Prefer `--cpu`,
`--memory`, and `--gpu` for supported resource intent. When `--gpu` is present
without a count, OpenShell treats it as a request for one GPU. Pass
`--gpu COUNT` when requesting more than one GPU.
Kubernetes maps the GPU count to the `nvidia.com/gpu` pod resource limit.
Docker and Podman satisfy count-only GPU requests by selecting the requested
number of NVIDIA CDI devices from the local CDI inventory in round-robin order.
The drivers refresh the CDI inventory before validating or creating the
sandbox, so CDI devices added or removed after gateway start can affect later
creates. On WSL2 all-only runtimes, Docker or Podman can use
`nvidia.com/gpu=all` as a compatibility fallback, where it counts as one
selectable device.
Exact GPU device selection remains driver-owned and requires `--gpu`. Docker
and Podman accept `cdi_devices` as opaque CDI device names; replace the
top-level `docker` key with `podman` when using the Podman driver, for example
`{"docker":{"cdi_devices":["nvidia.com/gpu=0"]}}`. Explicit CDI device lists
must not contain duplicates, and their length must match the effective GPU
count. A single exact CDI device is compatible with the default `--gpu`
request. The VM driver accepts `gpu_device_ids`, for example
`{"vm":{"gpu_device_ids":["0000:2d:00.0"]}}`; the current VM implementation
accepts at most one entry and allows either `--gpu` or `--gpu 1` when
`gpu_device_ids` is set.
For Kubernetes, `pod.runtime_class_name` maps to PodSpec `runtimeClassName`.
It overrides the gateway's configured default runtime class for that sandbox,
while a typed `SandboxTemplate.runtime_class_name` value from the API still
takes precedence.
Docker and Podman report the address through which their sandboxes can reach
the gateway. If the primary listener covers that address, the gateway reuses
it and sandbox JWT authentication restricts the supervisor to its callback RPC
allowlist. If the primary listener is not reachable through that address, the
gateway creates an additional callback-only listener. Use the primary endpoint
for CLI, administrator, health, reflection, inference-route management, and
HTTP requests. A `PermissionDenied` response from an additional callback-only
listener is expected for those requests. Do not broaden the primary listener
to `0.0.0.0` solely to make sandbox callbacks reachable.
## Docker Driver
[Docker](https://www.docker.com/get-started/)-backed sandboxes run as containers on the gateway host. Use Docker for local development, single-machine gateways, and hosts that already use Docker Desktop or Docker Engine.
The gateway talks to the Docker daemon to create sandbox containers. Docker is also required for local image builds from directories or Dockerfiles.
Docker Desktop and compatible macOS runtimes route `host.openshell.internal`
through an IPv4 host-gateway alias. The gateway reuses an IPv4 primary listener
that already covers loopback. Otherwise, the Docker driver requests a separate
`127.0.0.1:<gateway-port>` callback-only listener.
For maintainer-level implementation details, refer to the [Docker driver README](https://github.com/NVIDIA/OpenShell/blob/main/crates/openshell-driver-docker/README.md).
Select Docker with `compute_drivers = ["docker"]` in `[openshell.gateway]`. Configure Docker driver values such as `socket_path`, `grpc_endpoint`, `network_name`, `supervisor_bin`, `supervisor_image`, `image_pull_policy`, `ssh_socket_path`, `sandbox_pids_limit`, and `guest_tls_*` in `[openshell.drivers.docker]`. When `socket_path` is unset, the driver uses the same responsive local socket selected by auto-detection. An explicitly selected Docker driver falls back to `/var/run/docker.sock` when no candidate responds.
When operating `openshell-driver-docker` as an external driver, set
`OPENSHELL_OTLP_ENDPOINT` to export its spans. The driver continues W3C trace
context from gateway RPCs and reports as `openshell-driver-docker`.
Stop stops the existing Docker container without removing its writable
layer or attached volumes. Start starts that same container. A durably
stopped container stays stopped across gateway restart, and delete remains
responsible for removing it. Graceful gateway shutdown stops running-intent
Docker containers through the driver RPC without recording an explicit
sandbox stop. On startup, the gateway reconciles that retained intent with an
idempotent start request. Explicitly stopped sandboxes remain stopped.
For GPU-backed Docker sandboxes, configure Docker CDI before starting the gateway so OpenShell can detect the daemon capability.
### Docker Driver Config Mounts
Docker driver config accepts user-supplied `volume` and `tmpfs` mounts. It also
accepts `bind` mounts when `[openshell.drivers.docker]` sets
`enable_bind_mounts = true` in `gateway.toml`. See Docker's [storage documentation](https://docs.docker.com/engine/storage/) for more information.
Docker local-driver named volumes created with bind options also expose
gateway-host paths, so OpenShell treats them like bind mounts and requires
`enable_bind_mounts = true`.
Use a `volume` mount for existing Docker named volumes:
```shell
docker volume create openshell-work
openshell sandbox create \
--driver-config-json '{"docker":{"mounts":[{"type":"volume","source":"openshell-work","target":"/sandbox/work","read_only":false}]}}' \
-- claude
```
<Warning>
Bind mounts share gateway-host filesystem resources with the sandbox. They may
be considered insecure because they can negate OpenShell controls such as
workspace isolation and filesystem policy. Use them only when you understand and
accept that loss of isolation.
</Warning>
Use a `bind` mount only after enabling it in the Docker driver table:
```toml
[openshell.drivers.docker]
enable_bind_mounts = true
```
```shell
openshell sandbox create \
--driver-config-json '{"docker":{"mounts":[{"type":"bind","source":"/srv/openshell/work","target":"/sandbox/work","read_only":false}]}}' \
-- claude
```
Docker mount schema:
| Type | Fields |
|---|---|
| `bind` | `source`, `target`, optional `read_only` (`true` by default), optional `selinux_label` (`shared` for `:z` or `private` for `:Z`). `source` must be an absolute host path. Requires `enable_bind_mounts = true`. |
| `volume` | `source`, `target`, optional `read_only` (`true` by default), optional `subpath`. The named volume must already exist. Docker local-driver bind-backed volumes require `enable_bind_mounts = true`. |
| `tmpfs` | `target`, optional `options`, optional `size_bytes`, optional `mode`. |
OpenShell rejects mount `source`, `target`, and Docker volume `subpath` values
with surrounding whitespace. OpenShell also rejects mount targets that replace
the workspace root or container root, or contain or are contained by the
configured SSH socket or reserved `/opt/openshell`, `/etc/openshell`,
`/etc/openshell-tls`, `/run/openshell`, `/run/openshell-sidecar`, and network
namespace roots. These checks do not make host bind mounts safe.
## Podman Driver
[Podman](https://podman.io/)-backed sandboxes run as rootless containers on the gateway host. Use Podman for Linux workstation workflows that avoid a rootful Docker daemon.
The gateway talks to the Podman API socket. The Podman driver requires Podman 5.x, cgroups v2, rootless networking, and an active Podman user socket. When `socket_path` is not set, the driver probes for a responsive Podman socket and fails to start if none respond.
For maintainer-level implementation details, refer to the [Podman driver README](https://github.com/NVIDIA/OpenShell/blob/main/crates/openshell-driver-podman/README.md) and [Podman networking notes](https://github.com/NVIDIA/OpenShell/blob/main/crates/openshell-driver-podman/NETWORKING.md).
Select Podman with `compute_drivers = ["podman"]` in `[openshell.gateway]`. Configure Podman driver values such as `socket_path`, `network_name`, `supervisor_image`, `stop_timeout_secs`, `image_pull_policy`, `grpc_endpoint`, `host_gateway_ip`, `sandbox_ssh_socket_path`, `sandbox_pids_limit`, and `guest_tls_*` in `[openshell.drivers.podman]`.
Podman sandboxes default to a 45-second graceful stop window before Podman escalates from `SIGTERM` to `SIGKILL`. Set `stop_timeout_secs` in gateway config, or `OPENSHELL_STOP_TIMEOUT` for the standalone driver, when a local runtime needs a different teardown window.
Stop stops the existing Podman container while retaining its named workspace
volume and driver-owned secrets. Start starts the same container. Delete is
the operation that removes the container and named volume. Graceful gateway
shutdown stops running-intent Podman containers through the driver RPC without
recording an explicit sandbox stop. On startup, the gateway reconciles that
retained intent with an idempotent start request while leaving explicitly
stopped sandboxes alone.
For proxy-required networks, the Podman driver also accepts the corporate egress proxy keys `https_proxy`, `no_proxy`, `proxy_auth_file`, `proxy_auth_allow_insecure`, and `proxy_connect_by_hostname`. The supervisor chains policy-approved TLS tunnels through the proxy with HTTP CONNECT instead of dialing destinations directly. See the [Gateway Configuration File](./gateway-config) reference for the full contract, including the cleartext-credential acknowledgement and the validated-IP CONNECT behavior.
On macOS with `podman machine`, the driver uses gvproxy's host-loopback IP, `192.168.127.254`, for sandbox host aliases by default. Set `host_gateway_ip` only when your Podman machine uses a non-standard host-loopback address. On Linux, an empty `host_gateway_ip` keeps Podman's `host-gateway` resolver behavior. Direct local callbacks from rootless Podman require Podman to report the pasta network helper. Slirp4netns, other helpers, and Podman versions that do not report their helper require an explicitly remote `grpc_endpoint`; otherwise the gateway fails startup rather than leaving sandbox callbacks unreachable. Rootful Podman continues to use the configured network's bridge gateway address.
### Podman Driver Config Mounts
Podman driver config accepts user-supplied `volume`, `tmpfs`, and `image`
mounts. It also accepts `bind` mounts when `[openshell.drivers.podman]` sets
`enable_bind_mounts = true` in `gateway.toml`. Podman local-driver named
volumes created with bind options also expose gateway-host paths, so
OpenShell treats them like bind mounts and requires `enable_bind_mounts = true`.
Host bind mounts expose gateway host paths to sandbox requests, so they are
disabled by default.
Use a `volume` mount for existing Podman named volumes:
```shell
podman volume create openshell-work
openshell sandbox create \
--driver-config-json '{"podman":{"mounts":[{"type":"volume","source":"openshell-work","target":"/sandbox/work","read_only":false}]}}' \
-- claude
```
<Warning>
Bind mounts share gateway-host filesystem resources with the sandbox. They may
be considered insecure because they can negate OpenShell controls such as
workspace isolation and filesystem policy. Use them only when you understand and
accept that loss of isolation.
</Warning>
Use a `bind` mount only after enabling it in the Podman driver table:
```toml
[openshell.drivers.podman]
enable_bind_mounts = true
```
```shell
openshell sandbox create \
--driver-config-json '{"podman":{"mounts":[{"type":"bind","source":"/srv/openshell/work","target":"/sandbox/work","read_only":false}]}}' \
-- claude
```
Podman mount schema:
| Type | Fields |
|---|---|
| `bind` | `source`, `target`, optional `read_only` (`true` by default), optional `selinux_label` (`shared` for `:z` or `private` for `:Z`). `source` must be an absolute host path. Requires `enable_bind_mounts = true`. |
| `volume` | `source`, `target`, optional `read_only` (`true` by default). The named volume must already exist. Podman local-driver bind-backed volumes require `enable_bind_mounts = true`. |
| `tmpfs` | `target`, optional `options`, optional `size_bytes`, optional `mode`. |
| `image` | `source`, `target`, optional `read_only` (`true` by default). |
Podman `volume` and `image` mounts do not support `subpath` in OpenShell driver
config, and OpenShell rejects `subpath` for those mount types. OpenShell rejects
mount `source` and `target` values with surrounding whitespace. OpenShell also
rejects mount targets that replace the workspace root, container root, supervisor
files, `/etc/openshell`, `/etc/openshell-tls`, authentication material, or
network namespace paths. These checks do not make host bind mounts safe.
## MicroVM Driver
MicroVM-backed sandboxes run inside VM-backed isolation instead of a container boundary. Use MicroVM when workloads need a VM boundary instead of a local container boundary.
The gateway uses the VM compute driver to create VM-backed sandboxes. MicroVM requires host virtualization support. It uses [libkrun](https://github.com/containers/libkrun) with Apple's [Hypervisor framework](https://developer.apple.com/documentation/hypervisor) on macOS, KVM on Linux, and [QEMU](https://www.qemu.org/) for GPU-backed sandboxes on Linux.
The VM driver boots a cached immutable bootstrap ext4 root disk. When the requested sandbox image differs from the bootstrap image, the driver stages the registry image as an OCI layout, unpacks it inside a bootstrap VM with `umoci`, and caches the prepared image disk by image identity. Each sandbox receives that prepared disk read-only plus its own writable `overlay.ext4` disk for `/`, including `/sandbox` writes and runtime TLS material. The overlay persists for the sandbox lifetime and is deleted with the sandbox state directory.
VM sandbox creation follows the same progress model as Kubernetes-backed sandboxes. The gateway accepts the sandbox, then the VM driver publishes watch events while it resolves the image, prepares or reuses the bootstrap and prepared image caches, creates the writable overlay, and starts the VM launcher.
On graceful gateway shutdown, the gateway stops running-intent VMs through the driver RPC while retaining their launch records and writable overlays. On restart, the gateway starts a fresh VM driver process and reconciles the retained intent through the same idempotent start request used by the local container drivers. Running-intent VMs restart with their existing `overlay.ext4`, while explicitly stopped VMs remain stopped.
Stopped VM state directories contain a marker that prevents startup from
launching the VM. The driver retains `sandbox.pb`, `overlay.ext4`, and extension
state, then removes the marker and restores the same overlay on start.
For maintainer-level implementation details, refer to the [VM driver README](https://github.com/NVIDIA/OpenShell/blob/main/crates/openshell-driver-vm/README.md).
### Enable the VM Driver
The VM driver is opt-in. Release packages can install `openshell-driver-vm`, but the gateway does not select it unless you configure the driver explicitly.
Enable VM by setting `compute_drivers = ["vm"]` in the gateway TOML file:
```toml
[openshell.gateway]
compute_drivers = ["vm"]
```
For a launch-time override, set `OPENSHELL_DRIVERS=vm` in the gateway environment and restart the service.
Configure VM driver values such as `grpc_endpoint`, `driver_dir`, `state_dir`, `default_image`, `bootstrap_image`, `vcpus`, `mem_mib`, `overlay_disk_mib`, `krun_log_level`, and `guest_tls_*` in `[openshell.drivers.vm]`. The VM `state_dir` stores overlay disks, console logs, runtime state, image-rootfs cache, and the private `run/compute-driver.sock` socket. The VM socket path is managed by the gateway and is not configurable through remote endpoint settings.
The gateway starts `openshell-driver-vm` over a private Unix socket and passes its process ID so the driver can reject unexpected local clients. The driver's standalone TCP listener is disabled unless `--allow-unauthenticated-tcp` is set for local development.
### Local image resolution
The VM driver resolves sandbox images from a local container engine before falling back to registry pulls. It tries Docker first, then falls back to the Podman socket (Docker-compatible API). On Linux with Podman, enable the API socket so the driver can find local images:
```shell
systemctl --user start podman.socket
```
### Host Firewall
The VM driver creates nftables rules on the host for each sandbox VM's TAP network interface. These rules provide NAT for VM connectivity and defense-in-depth isolation: unsolicited inbound connections to the VM are dropped, and the VM can only reach the gateway port on the host. Primary security enforcement (proxy-only egress and bypass detection) is handled by the sandbox supervisor inside the VM guest.
On hosts with restrictive firewalls (e.g. firewalld), the host firewall may additionally block VM traffic that the driver's rules accept. If VM sandboxes cannot reach the network, verify that the host firewall allows forwarding and input for `vmtap-*` interfaces. See the [VM driver README](https://github.com/NVIDIA/OpenShell/blob/main/crates/openshell-driver-vm/README.md#host-side-nftables-rules) for details.
## Kubernetes Driver
Kubernetes-backed sandboxes run as pods in the configured sandbox namespace. Use Kubernetes for shared clusters, remote compute, GPU scheduling, and operator-managed environments.
<Warning>
Kubernetes workspace namespaces are an administrative trust boundary. In
shared and managed modes, only the OpenShell gateway and its trusted Agent
Sandbox controller may administer Sandbox CRs, sandbox pods, or the configured
sandbox ServiceAccount in those namespaces. In operator mode, allowlist only
namespaces where the platform operator preserves that exclusive control.
Untrusted principals must not be able to create sandbox pods with fabricated
owner references or use the sandbox ServiceAccount. The operator namespace
allowlist is a trust grant, not a tenant isolation mechanism.
</Warning>
Helm deployments set Kubernetes driver values through the chart.
For maintainer-level implementation details, refer to the [Kubernetes driver README](https://github.com/NVIDIA/OpenShell/blob/main/crates/openshell-driver-kubernetes/README.md).
| Gateway configuration | Helm value | Description |
|---|---|---|
| `compute_drivers = ["kubernetes"]` | Not applicable | Select the Kubernetes compute driver. |
| `[openshell.drivers.kubernetes].namespace` | `server.sandboxNamespace` | Set the namespace for sandbox resources. The Helm chart defaults to the release namespace when left empty. |
| `service_account_name` | `sandboxServiceAccount.name` | Set the Kubernetes service account assigned to sandbox pods and accepted by the gateway TokenReview bootstrap path. The Helm chart creates a dedicated sandbox service account by default. |
| `default_image` | `server.sandboxImage` | Set the default sandbox image. |
| `image_pull_policy` | `server.sandboxImagePullPolicy` | Set the Kubernetes image pull policy for sandbox pods. |
| `image_pull_secrets` | `server.sandboxImagePullSecrets` | Attach Kubernetes image-pull Secrets to sandbox pods. Managed mode copies these explicitly named Secrets from the configured source namespace into each workspace namespace. In shared and operator modes, the Secrets must already exist in the sandbox namespace. |
| `[managed_ssh_ingress]` | `networkPolicy.enabled` | In managed mode, create an SSH ingress policy in every workspace namespace. Helm configures the gateway namespace and pod selector automatically. Operator mode leaves namespace policy management to the platform operator. |
| `grpc_endpoint` | `server.grpcEndpoint` | Set the gateway callback endpoint reachable from sandbox pods. |
| `client_tls_secret_name` | `server.tls.clientTlsSecretName` | Mount sandbox client TLS materials from a Kubernetes secret. |
| `supervisor_image` | `supervisor.image.repository` / `supervisor.image.tag` | Override the supervisor image that provides the `openshell-sandbox` binary. The default repository with an empty tag uses the version-pinned image built into the gateway. Changing the repository uses the effective gateway image tag, while setting a tag pins that version explicitly. |
| `supervisor_image_pull_policy` | `supervisor.image.pullPolicy` | Set the Kubernetes image pull policy for the supervisor image. |
| `supervisor_sideload_method` | `supervisor.sideloadMethod` | How the supervisor binary is delivered into sandbox pods. Leave empty to auto-detect from cluster version. Set to `image-volume` to mount the supervisor OCI image directly as a volume (requires Kubernetes 1.33+ with the ImageVolume feature gate; GA in 1.36), or `init-container` to copy it through an init container on older clusters. |
| `topology` | `supervisor.topology` | Set `combined` for the default single supervisor path, or `sidecar` to move pod-level network enforcement and the gateway session into a dedicated sidecar. |
| `https_proxy` | `upstreamProxy.url` | Set the operator-owned `http://host:port` corporate forward proxy used for policy-approved TLS CONNECT egress. |
| `no_proxy` | `upstreamProxy.noProxy` | Set destinations that bypass only the corporate proxy. OpenShell policy evaluation still applies. |
| `proxy_auth_secret_name` | `upstreamProxy.authSecret.name` | Set the existing Secret name in the sandbox namespace that contains the proxy credential. Requires `sidecar` topology. |
| `proxy_auth_secret_key` | `upstreamProxy.authSecret.key` | Set the Secret key containing the `user:pass` credential. Requires `sidecar` topology. |
| `proxy_auth_allow_insecure` | `upstreamProxy.authAllowInsecure` | Set `true` to acknowledge that Basic authentication to an HTTP proxy is cleartext. Required with a proxy credential Secret. |
| `proxy_connect_by_hostname` | `upstreamProxy.connectByHostname` | Send hostnames rather than validated IPs in CONNECT requests. Use only when proxy ACLs require hostname targets. |
| `sidecar.proxy_uid` | `supervisor.sidecar.proxyUid` | Dedicated UID of at least `1000` used by the relaxed sidecar when process/binary-aware network policy is disabled. It must not match the workload UID. The default binary-aware sidecar runs as UID 0. The network init container exempts the effective sidecar UID from proxy redirection. |
| `sidecar.process_binary_aware_network_policy` | `supervisor.sidecar.processBinaryAwareNetworkPolicy` | Keep process/binary-aware network policy enabled in `sidecar` topology. The default runs the sidecar as UID 0 with `SYS_PTRACE` and `DAC_READ_SEARCH`. Set false to run as `proxy_uid`, drop both capabilities, and enforce endpoint/L7 policy without matching `policy.binaries`. |
| `app_armor_profile` | `server.appArmorProfile` | Set the sandbox agent container's AppArmor profile. Helm defaults this to `Unconfined` so AppArmor-enabled nodes do not block supervisor network namespace setup. Set the Helm value to an empty string to omit the field, or use `RuntimeDefault` or `Localhost/<profile-name>` for operator-managed profiles. |
| `workspace_default_storage_size` | `server.workspaceDefaultStorageSize` | Set the default workspace PVC size for new sandboxes. |
| `workspace_storage_class` | `server.workspaceStorageClass` | Set the `StorageClass` for the workspace PVC. Empty (default) omits `storageClassName` and uses the cluster's default `StorageClass`. Set this on clusters with no default `StorageClass`, otherwise the workspace PVC stays `Pending` and the sandbox never starts. |
| `sa_token_ttl_secs` | `server.sandboxJwt.k8sSaTokenTtlSecs` | Set the projected ServiceAccount token TTL used for the bootstrap token exchange. |
Managed-mode Secret copying requires the gateway ServiceAccount to create
Secrets. Kubernetes RBAC cannot restrict Secret `create` by resource name, so
the Helm chart grants cluster-wide Secret `create`; Secret `get` and `patch`
remain limited to the explicitly configured TLS and image-pull Secret names.
The driver creates copies only in gateway-owned managed namespaces. Do not
reuse the gateway ServiceAccount for unrelated workloads.
In `combined` topology, the agent container carries the Linux capabilities
needed by the supervisor for network namespace setup, Landlock filesystem
policy, process privilege changes, and network policy enforcement. In `sidecar`
topology, the agent container runs as the resolved sandbox UID/GID with no added
Linux capabilities. A root init container performs the nftables setup, and the
long-running binary-aware sidecar runs as UID 0, drops default capabilities,
and adds `SYS_PTRACE` plus `DAC_READ_SEARCH` for workload process identity
resolution through shared `/proc`. The
`sidecar.process_binary_aware_network_policy = false` setting runs it as the
configured non-root `proxy_uid`, removes both capabilities, and relaxes network
policy to endpoint/L7 matching only. The
network sidecar owns gateway authentication and writes local policy/provider
state to the process supervisor over a local control socket, so the agent
container does not mount the sandbox bootstrap token or client TLS secret in
the default sidecar path. The provider environment is refreshed by the network
sidecar after settings polls and streamed to the process supervisor so future
child processes can see updated provider env without gateway access in the
agent container.
Sidecar mode keeps gateway session and SSH behavior. The process supervisor
applies Landlock filesystem policy and child seccomp filters where supported,
but it does not perform root-to-sandbox privilege dropping or supervisor
identity mount isolation. Network policy still runs in the sidecar, and sidecar
pods set `shareProcessNamespace: true` so the network sidecar can resolve
process/binary identity through `/proc/<entrypoint-pid>`.
The Kubernetes driver creates namespaced `agents.x-k8s.io` `Sandbox` resources from the Kubernetes SIG Apps [agent-sandbox](https://github.com/kubernetes-sigs/agent-sandbox) project. It detects the served Sandbox API at runtime, caches the selected API version for the gateway process, and uses `v1beta1` when available before falling back to `v1alpha1`, so supported Agent Sandbox installations work without version-specific operator configuration. The Agent Sandbox controller turns those resources into sandbox pods and related storage.
Stop patches the existing resource rather than deleting it. For `v1beta1`,
the driver sets `spec.operatingMode` to `Suspended` or `Running`. For
`v1alpha1`, it sets `spec.replicas` to `0` or `1`. The Sandbox resource and its
workspace PVC keep their identity across both operations.
If Agent Sandbox is upgraded in place, restart the OpenShell gateway after the controller and CRD rollout completes so the gateway can detect the served API versions again.
<Note>
`Sandbox.spec.volumeClaimTemplates` is immutable after creation. To change storage configuration, delete the sandbox and create a new one with the updated spec.
</Note>
### Kubernetes Driver Config PVC Mounts
Kubernetes driver config can mount existing PersistentVolumeClaims into the
agent container. Use this when storage is provisioned outside OpenShell and a
sandbox should mount selected PVC subpaths instead of using the default
OpenShell-created `/sandbox` workspace PVC.
```shell
openshell sandbox create \
--driver-config-json '{
"kubernetes": {
"volumes": [{
"name": "user-data",
"persistent_volume_claim": {
"claim_name": "pvc-user-data-123",
"read_only": false
}
}],
"containers": {
"agent": {
"volume_mounts": [
{
"name": "user-data",
"mount_path": "/sandbox/.openshell/workspace",
"sub_path": "workspace",
"read_only": false
},
{
"name": "user-data",
"mount_path": "/sandbox/.openshell/memory",
"sub_path": "memory",
"read_only": false
}
]
}
}
}
}' \
-- claude
```
Kubernetes PVC mount schema:
| Field | Description |
|---|---|
| `volumes[].name` | Pod volume name. It must be a DNS-1123 label, unique, and not use OpenShell-managed volume names. |
| `volumes[].persistent_volume_claim.claim_name` | Existing PVC name in the sandbox namespace. It must be a DNS-1123 subdomain name. |
| `volumes[].persistent_volume_claim.read_only` | Optional. Defaults to `true`. Set `false` to allow read-write mounts. |
| `containers.agent.volume_mounts[].name` | References a volume declared in `volumes`. |
| `containers.agent.volume_mounts[].mount_path` | Absolute, normalized container path for the agent mount. |
| `containers.agent.volume_mounts[].sub_path` | Optional relative PVC subpath. Absolute paths and `..` are rejected. |
| `containers.agent.volume_mounts[].read_only` | Optional. Defaults to `true`. It cannot be `false` when the PVC volume is read-only. |
OpenShell rejects duplicate volume names, mounts that reference unknown volumes,
protected mount targets, and mounts that replace OpenShell TLS, supervisor,
ServiceAccount token, or SPIFFE paths. Read-write PVC access requires
`read_only: false` on both the PVC volume and each writable mount.
Any driver-config mount under `/sandbox` disables the default `/sandbox`
workspace PVC injection for that sandbox. Only the explicit mount paths persist
through the external PVC; other `/sandbox` paths come from the current sandbox
image.
## Sandbox User Identity
The policy can set `process.run_as_user` and `process.run_as_group`
independently. Each explicit field wins. The active compute driver supplies the
identity for omitted fields.
Explicit numeric values may use any non-root Linux UID/GID from `1` through
`4294967294`. OpenShell rejects `0` as root and `4294967295` as the invalid
identity sentinel. Low numeric identities can inherit permissions from matching
accounts, files, volumes, or devices, so choose them with the same care as any
other runtime identity.
### Docker / Podman
Docker and Podman inspect the final image and use its OCI `USER` declaration as
a per-field fallback. Supported forms include `app`, `app:staff`, a numeric UID
whose passwd entry supplies its primary GID, and an accountless numeric pair
such as `1234:1235`.
The driver pins container creation to the immutable image ID it inspected. The
supervisor validates any required names inside that image and preserves the
declared name or numeric components for both direct and SSH children. When
`USER` omits the group, the supervisor uses the user's numeric primary GID. It
does not modify `/etc/passwd` or `/etc/group`.
Docker also inspects OCI `WorkingDir`. An absolute value becomes the
agent workspace; an empty, root (`/`), or explicit `/sandbox` value uses the
managed `/sandbox` compatibility workspace.
OpenShell creates and owns that compatibility workspace. Any other workdir must
already exist in the immutable image without symlink components. The completed
UID/GID and supplementary groups must already be able to traverse every parent
and write and enter the workdir. OpenShell does not change that directory's
ownership or mode. A one-shot validator drops to that identity and uses kernel
effective-access checks, including POSIX ACL grants and LSM denials. It rejects
workdirs that overlap the OCI runtime namespaces under `/proc`, `/sys`, or
`/dev`, and rejects overlap with actual OpenShell control paths. Docker checks
the original image filesystem in the final supervisor and rejects image
`VOLUME` declarations that would mask the workdir or one of its parents before
validation. The resolved workspace is the cwd and `HOME` for direct and SSH
children. The supervisor itself starts from `/`, so a missing or invalid
workspace is handled during readiness instead of preventing the container
runtime from starting it.
Sandbox creation fails before readiness if a required `USER` component is
missing, malformed, unknown, ambiguous, or resolves to UID/GID 0. An image
without `USER` therefore works only when policy explicitly provides both
identity fields.
### Kubernetes / OpenShift
The Kubernetes driver auto-detects the sandbox UID from OpenShift SCC namespace annotations:
- `openshift.io/sa.scc.uid-range` (format: `<start>/<size>`, e.g. `1000000000/10000`) provides the UID.
- `openshift.io/sa.scc.supplemental-groups` provides the GID when present; otherwise the resolved UID is used as the GID.
- On non-OpenShift clusters, or when annotations are absent, the driver falls back to `1000`.
You can override autodetection with explicit `sandbox_uid` / `sandbox_gid` config in `[openshell.drivers.kubernetes]`. When set, the driver skips namespace annotation lookup entirely.
The resolved UID/GID appear in:
- Supervisor container environment variables (`OPENSHELL_SANDBOX_UID`, `OPENSHELL_SANDBOX_GID`) for direct kernel-level privilege dropping without `/etc/passwd` lookups.
- PVC init container `securityContext.runAsUser/runAsGroup/fsGroup` for workspace ownership operations.
### VM Driver
The VM driver injects the sandbox UID into the rootfs guest's `/etc/passwd`, `/etc/group`, and `/etc/gshadow` during rootfs preparation. Default UID is `10001`; configure `sandbox_uid` in `[openshell.drivers.vm]` to use a different value.
### Custom Images
Docker and Podman custom images do not need a baked-in `"sandbox"` user. Declare
a non-root OCI `USER`, or set both process identity fields explicitly in policy.
Named image users require matching account entries; a numeric `UID:GID` pair
does not. For Docker, declare an absolute OCI `WORKDIR` to select the workspace.
Images with no working directory, `WORKDIR /`, or `WORKDIR /sandbox` use
OpenShell's managed `/sandbox` compatibility workspace. For any other Docker
path, create the directory in the image and grant the final process identity
write and execute permission in the Dockerfile. Podman, Kubernetes/OpenShift,
and VM sandboxes continue to use `/sandbox`.
@@ -0,0 +1,98 @@
---
# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
title: "Support Matrix"
description: ""
position: 5
---
This page lists the host platform, compute driver, software, runtime, and kernel requirements for running OpenShell.
## Supported Platforms
OpenShell publishes multi-architecture gateway images for `linux/amd64` and `linux/arm64`. The CLI, package-managed gateway, and standalone gateway binary are supported on the following host platforms:
| Platform | Architecture | Status |
| -------------------------------- | --------------------- | --------- |
| Linux (Debian/Ubuntu) | x86_64 (amd64) | Supported |
| Linux (Debian/Ubuntu) | aarch64 (arm64) | Supported |
| macOS (Docker Desktop) | Apple Silicon (arm64) | Supported |
| Windows (WSL 2 + Docker Desktop) | x86_64 | Experimental |
On Linux, the `openshell` CLI is a static musl binary and does not require glibc at runtime.
## Standalone Gateway Binary
OpenShell publishes standalone `openshell-gateway` release assets for manual download on these platforms:
| Platform | Artifact pattern |
| --------------------- | ---------------------------------------------- |
| Linux x86_64 (amd64) | `openshell-gateway-x86_64-unknown-linux-gnu` |
| Linux aarch64 (arm64) | `openshell-gateway-aarch64-unknown-linux-gnu` |
| macOS Apple Silicon | `openshell-gateway-aarch64-apple-darwin` |
These artifacts are attached to GitHub releases. Kubernetes deployments should use the Helm chart and the published gateway image.
On Linux, `openshell-gateway` requires glibc 2.28 or newer. Compatible systems include, for example, Ubuntu 20.04+, RHEL 8+, Rocky Linux 8+, Amazon Linux 2023+, and Fedora 32+.
## Compute Drivers
The gateway can manage sandboxes through several compute drivers.
| Compute Driver | Status | Notes |
|---|---|---|
| Docker | Supported for local development and single-machine gateways. | Requires Docker Desktop or Docker Engine on the gateway host. |
| Podman | Supported for rootless local and workstation workflows. | Requires a Podman-compatible socket and rootless networking setup. |
| Kubernetes | Supported through the [OpenShell Helm chart](https://github.com/NVIDIA/OpenShell/blob/main/deploy/helm/openshell/README.md). | Requires a Kubernetes cluster supplied by the operator. |
| MicroVM | Supported for VM-backed sandboxes. | Uses the VM compute driver and libkrun-based runtime. |
## Software Prerequisites
Install the software for the compute driver you use:
| Component | Minimum Version | Notes |
|---|---|---|
| Docker Desktop or Docker Engine | 28.0 | Required for Docker-backed gateways, local image builds, and Docker development workflows. |
| Podman | 5.x | Required for Podman-backed gateways. |
| Kubernetes | 1.29 | Required for Helm deployments and Kubernetes sandbox scheduling. |
| Helm | 3.x | Required to install `deploy/helm/openshell`. |
| kubectl | Compatible with your cluster | Required for Kubernetes operational inspection and secret creation. |
| Host virtualization | Host dependent | Required for MicroVM-backed gateways. MicroVM uses Hypervisor.framework on macOS and KVM on Linux. |
## Sandbox Runtime Versions
Sandbox container images are maintained in the [OpenShell Community](https://github.com/NVIDIA/OpenShell-Community) repository. Refer to that repository for the current list of installed components and their versions.
## Container Images
OpenShell publishes the gateway image for `linux/amd64` and `linux/arm64`.
| Image | Reference | Pulled When |
|---|---|---|
| Gateway | `ghcr.io/nvidia/openshell/gateway:latest` | Helm chart install or upgrade, or standalone container deployment |
The Helm chart in `deploy/helm/openshell` deploys the gateway workload, service account, service, optional persistent storage, and network policy for Kubernetes. It defaults to a StatefulSet for SQLite-backed installs and can render a Deployment for external database-backed installs.
Sandbox images are maintained separately in the [OpenShell Community](https://github.com/NVIDIA/OpenShell-Community) repository.
To override the default image references, use Helm values:
| Helm value | Purpose |
|---|---|
| `image.repository` / `image.tag` | Override the gateway image reference. |
| `server.sandboxImage` | Override the default sandbox image. |
## Kernel Requirements
OpenShell enforces sandbox isolation through two Linux kernel security modules:
| Module | Requirement | Details |
| -------------------------------------------------------------- | ----------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| [Landlock LSM](https://docs.kernel.org/security/landlock.html) | Recommended | Enforces filesystem access restrictions at the kernel level. The `best_effort` compatibility mode uses the highest Landlock ABI the host kernel supports. The `hard_requirement` mode fails sandbox creation if the required ABI is unavailable. |
| seccomp | Required | Filters dangerous system calls. Available on all modern Linux kernels (3.17+). |
On macOS, these kernel modules run inside the Docker Desktop Linux VM, not on the host kernel.
## Agent Compatibility
For the full list of supported agents and their default policy coverage, refer to the [Supported Agents](/about/supported-agents) page.
+201
View File
@@ -0,0 +1,201 @@
---
# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
title: "License"
description: "NVIDIA OpenShell is licensed under the Apache License, Version 2.0."
keywords: "Legal, License, Apache 2.0"
position: 1
---
NVIDIA OpenShell is licensed under the [Apache License, Version 2.0](https://www.apache.org/licenses/LICENSE-2.0).
Apache License
Version 2.0, January 2004
http://www.apache.org/licenses/
TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION
1. Definitions.
"License" shall mean the terms and conditions for use, reproduction,
and distribution as defined by Sections 1 through 9 of this document.
"Licensor" shall mean the copyright owner or entity authorized by
the copyright owner that is granting the License.
"Legal Entity" shall mean the union of the acting entity and all
other entities that control, are controlled by, or are under common
control with that entity. For the purposes of this definition,
"control" means (i) the power, direct or indirect, to cause the
direction or management of such entity, whether by contract or
otherwise, or (ii) ownership of fifty percent (50%) or more of the
outstanding shares, or (iii) beneficial ownership of such entity.
"You" (or "Your") shall mean an individual or Legal Entity
exercising permissions granted by this License.
"Source" form shall mean the preferred form for making modifications,
including but not limited to software source code, documentation
source, and configuration files.
"Object" form shall mean any form resulting from mechanical
transformation or translation of a Source form, including but
not limited to compiled object code, generated documentation,
and conversions to other media types.
"Work" shall mean the work of authorship, whether in Source or
Object form, made available under the License, as indicated by a
copyright notice that is included in or attached to the work
(an example is provided in the Appendix below).
"Derivative Works" shall mean any work, whether in Source or Object
form, that is based on (or derived from) the Work and for which the
editorial revisions, annotations, elaborations, or other modifications
represent, as a whole, an original work of authorship. For the purposes
of this License, Derivative Works shall not include works that remain
separable from, or merely link (or bind by name) to the interfaces of,
the Work and Derivative Works thereof.
"Contribution" shall mean any work of authorship, including
the original version of the Work and any modifications or additions
to that Work or Derivative Works thereof, that is intentionally
submitted to the Licensor for inclusion in the Work by the copyright owner
or by an individual or Legal Entity authorized to submit on behalf of
the copyright owner. For the purposes of this definition, "submitted"
means any form of electronic, verbal, or written communication sent
to the Licensor or its representatives, including but not limited to
communication on electronic mailing lists, source code control systems,
and issue tracking systems that are managed by, or on behalf of, the
Licensor for the purpose of discussing and improving the Work, but
excluding communication that is conspicuously marked or otherwise
designated in writing by the copyright owner as "Not a Contribution."
"Contributor" shall mean Licensor and any individual or Legal Entity
on behalf of whom a Contribution has been received by the Licensor and
subsequently incorporated within the Work.
2. Grant of Copyright License. Subject to the terms and conditions of
this License, each Contributor hereby grants to You a perpetual,
worldwide, non-exclusive, no-charge, royalty-free, irrevocable
copyright license to reproduce, prepare Derivative Works of,
publicly display, publicly perform, sublicense, and distribute the
Work and such Derivative Works in Source or Object form.
3. Grant of Patent License. Subject to the terms and conditions of
this License, each Contributor hereby grants to You a perpetual,
worldwide, non-exclusive, no-charge, royalty-free, irrevocable
(except as stated in this section) patent license to make, have made,
use, offer to sell, sell, import, and otherwise transfer the Work,
where such license applies only to those patent claims licensable
by such Contributor that are necessarily infringed by their
Contribution(s) alone or by combination of their Contribution(s)
with the Work to which such Contribution(s) was submitted. If You
institute patent litigation against any entity (including a
cross-claim or counterclaim in a lawsuit) alleging that the Work
or a Contribution incorporated within the Work constitutes direct
or contributory patent infringement, then any patent licenses
granted to You under this License for that Work shall terminate
as of the date such litigation is filed.
4. Redistribution. You may reproduce and distribute copies of the
Work or Derivative Works thereof in any medium, with or without
modifications, and in Source or Object form, provided that You
meet the following conditions:
(a) You must give any other recipients of the Work or
Derivative Works a copy of this License; and
(b) You must cause any modified files to carry prominent notices
stating that You changed the files; and
(c) You must retain, in the Source form of any Derivative Works
that You distribute, all copyright, patent, trademark, and
attribution notices from the Source form of the Work,
excluding those notices that do not pertain to any part of
the Derivative Works; and
(d) If the Work includes a "NOTICE" text file as part of its
distribution, then any Derivative Works that You distribute must
include a readable copy of the attribution notices contained
within such NOTICE file, excluding any notices that do not
pertain to any part of the Derivative Works, in at least one
of the following places: within a NOTICE text file distributed
as part of the Derivative Works; within the Source form or
documentation, if provided along with the Derivative Works; or,
within a display generated by the Derivative Works, if and
wherever such third-party notices normally appear. The contents
of the NOTICE file are for informational purposes only and
do not modify the License. You may add Your own attribution
notices within Derivative Works that You distribute, alongside
or as an addendum to the NOTICE text from the Work, provided
that such additional attribution notices cannot be construed
as modifying the License.
You may add Your own copyright statement to Your modifications and
may provide additional or different license terms and conditions
for use, reproduction, or distribution of Your modifications, or
for any such Derivative Works as a whole, provided Your use,
reproduction, and distribution of the Work otherwise complies with
the conditions stated in this License.
5. Submission of Contributions. Unless You explicitly state otherwise,
any Contribution intentionally submitted for inclusion in the Work
by You to the Licensor shall be under the terms and conditions of
this License, without any additional terms or conditions.
Notwithstanding the above, nothing herein shall supersede or modify
the terms of any separate license agreement you may have executed
with Licensor regarding such Contributions.
6. Trademarks. This License does not grant permission to use the trade
names, trademarks, service marks, or product names of the Licensor,
except as required for reasonable and customary use in describing the
origin of the Work and reproducing the content of the NOTICE file.
7. Disclaimer of Warranty. Unless required by applicable law or
agreed to in writing, Licensor provides the Work (and each
Contributor provides its Contributions) on an "AS IS" BASIS,
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or
implied, including, without limitation, any warranties or conditions
of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A
PARTICULAR PURPOSE. You are solely responsible for determining the
appropriateness of using or redistributing the Work and assume any
risks associated with Your exercise of permissions under this License.
8. Limitation of Liability. In no event and under no legal theory,
whether in tort (including negligence), contract, or otherwise,
unless required by applicable law (such as deliberate and grossly
negligent acts) or agreed to in writing, shall any Contributor be
liable to You for damages, including any direct, indirect, special,
incidental, or consequential damages of any character arising as a
result of this License or out of the use or inability to use the
Work (including but not limited to damages for loss of goodwill,
work stoppage, computer failure or malfunction, or any and all
other commercial damages or losses), even if such Contributor
has been advised of the possibility of such damages.
9. Accepting Warranty or Additional Liability. While redistributing
the Work or Derivative Works thereof, You may choose to offer,
and charge a fee for, acceptance of support, warranty, indemnity,
or other liability obligations and/or rights consistent with this
License. However, in accepting such obligations, You may act only
on Your own behalf and on Your sole responsibility, not on behalf
of any other Contributor, and only if You agree to indemnify,
defend, and hold each Contributor harmless for any liability
incurred by, or claims asserted against, such Contributor by reason
of your accepting any such warranty or additional liability.
END OF TERMS AND CONDITIONS
Copyright 2025-2026 NVIDIA CORPORATION & AFFILIATES
Licensed under the Apache License, Version 2.0 (the "License");
you may not use this file except in compliance with the License.
You may obtain a copy of the License at
http://www.apache.org/licenses/LICENSE-2.0
Unless required by applicable law or agreed to in writing, software
distributed under the License is distributed on an "AS IS" BASIS,
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
See the License for the specific language governing permissions and
limitations under the License.
@@ -0,0 +1,339 @@
---
# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
title: "Inference Routing"
sidebar-title: "Inference Routing"
description: "Understand and configure OpenShell inference routing through inference.local and external endpoints."
keywords: "Generative AI, Cybersecurity, Inference Routing, Configuration, Privacy, LLM, Provider"
position: 8
---
OpenShell handles inference traffic through two paths: external endpoints and `inference.local`.
| Path | How it works |
|---|---|
| External endpoints | Traffic to hosts like `api.openai.com` or `api.anthropic.com` is treated like any other outbound request, allowed or denied by `network_policies`. Refer to [Policies](/sandboxes/policies). |
| `inference.local` | A sandbox-local HTTPS endpoint that routes model requests through the gateway. The privacy router strips sandbox-supplied credentials, forwards only approved inference headers, injects the configured backend credentials, and forwards to the managed model endpoint. |
## How `inference.local` Works
When code inside a sandbox calls `https://inference.local`, the privacy router routes the request to the configured backend for that gateway. The configured model is applied to generation requests, provider credentials come from OpenShell rather than from code inside the sandbox, and only approved inference headers are forwarded upstream.
If code calls an external inference host directly, OpenShell evaluates that traffic only through `network_policies`.
| Property | Detail |
|---|---|
| Credentials | No sandbox API keys needed. Credentials come from the configured provider record. The router strips caller-supplied `Authorization` before forwarding the request. |
| Header forwarding | `inference.local` forwards only a per-provider header allowlist. OpenAI routes allow `openai-organization` and `x-model-id`. Anthropic routes allow `anthropic-version` and `anthropic-beta`. Vertex Claude rawPredict routes strip `anthropic-beta` and do not forward `anthropic-version` as a header because the router injects `anthropic_version` into the Vertex request body. NVIDIA routes allow `x-model-id`. AWS Bedrock routes have no passthrough headers today. All other caller headers are stripped. |
| Configuration | One provider and one model define sandbox inference for the active gateway. Every sandbox on that gateway sees the same `inference.local` backend. |
| Provider support | NVIDIA, Anthropic, Google Vertex AI, AWS Bedrock (via a translating bridge — direct AWS with SigV4 signing is a separate follow-up), and any OpenAI-compatible provider all work through the same endpoint. Vertex routes Claude models through `/v1/messages` and non-Anthropic models through `/v1/chat/completions`. The gateway resolves the upstream Vertex host from the provider config, including regional, global, and supported multi-region endpoints. |
| Streaming reliability | The router tolerates idle gaps of up to 120 seconds between streamed chunks so long reasoning responses are not cut off mid-stream. |
| Hot refresh | OpenShell picks up provider credential changes and inference updates without recreating sandboxes. Changes propagate within about 5 seconds by default. |
## Supported API Patterns
Supported request patterns depend on the provider configured for `inference.local`.
<Tabs>
<Tab title="OpenAI-compatible">
| Pattern | Method | Path |
|---|---|---|
| Chat Completions | `POST` | `/v1/chat/completions` |
| Completions | `POST` | `/v1/completions` |
| Responses | `POST` | `/v1/responses` |
| Embeddings | `POST` | `/v1/embeddings` |
| Model Discovery | `GET` | `/v1/models` |
| Model Discovery | `GET` | `/v1/models/*` |
</Tab>
<Tab title="Anthropic-compatible">
| Pattern | Method | Path |
|---|---|---|
| Messages | `POST` | `/v1/messages` |
</Tab>
<Tab title="AWS Bedrock">
| Pattern | Method | Path |
|---|---|---|
| InvokeModel | `POST` | `/model/{modelId}/invoke` |
The `{modelId}` segment is constrained to a single non-empty path segment to avoid path-traversal liabilities. `/model//invoke` and `/model/a/b/invoke` both no-match.
<Note>
Today the `aws-bedrock` provider type is bridge-fronted only. The router does not inject any auth header on outbound requests; the configured `BEDROCK_BASE_URL` is expected to point at a translating bridge or Bedrock-compatible proxy whose own pod holds operator-side credentials. SigV4 signing for direct AWS Bedrock is deferred to a follow-up release.
`InvokeModelWithResponseStream` is intentionally not advertised yet. The streaming path emits AWS event-stream framing, which our protocol-aware error path does not yet model; surfacing it without that work risks shipping responses the sandbox cannot interpret on failure. It will land alongside the streaming-error work in a follow-up.
</Note>
</Tab>
</Tabs>
Requests to `inference.local` that do not match the configured provider's supported patterns are denied.
Google Vertex AI does not expose every OpenAI-compatible path through `inference.local`. Vertex routes for Gemini and other non-Anthropic models currently support Chat Completions. Vertex routes for Claude models use the Anthropic Messages pattern. Base URL overrides are only supported for non-Anthropic Vertex routes.
## Configure Inference Routing
The managed local inference endpoint uses three values:
| Value | Description |
|---|---|
| Provider record | The credential backend OpenShell uses to authenticate with the upstream model host. |
| Model ID | The model to use for generation requests. |
| Timeout | Per-request timeout in seconds for upstream inference calls. Defaults to 60 seconds. |
For tested providers and base URLs, refer to [Supported Inference Providers](/sandboxes/manage-providers#supported-inference-providers).
## Create a Provider
Create a provider that holds the backend credentials you want OpenShell to use.
<Tabs>
<Tab title="NVIDIA API Catalog">
```shell
openshell provider create --name nvidia-prod --type nvidia --from-existing
```
This reads `NVIDIA_API_KEY` from your environment.
</Tab>
<Tab title="OpenAI-compatible Provider">
Any cloud provider that exposes an OpenAI-compatible API works with the `openai` provider type. You need three values from the provider: the base URL, an API key, and a model name.
```shell
openshell provider create \
--name my-cloud-provider \
--type openai \
--credential OPENAI_API_KEY=<your_api_key> \
--config OPENAI_BASE_URL=https://api.example.com/v1
```
Replace the base URL and API key with the values from your provider. For supported providers out of the box, refer to [Supported Inference Providers](/sandboxes/manage-providers#supported-inference-providers). For other providers, refer to your provider's documentation for the correct base URL, available models, and API key setup.
</Tab>
<Tab title="Google Vertex AI">
```shell
openshell provider create \
--name vertex-local \
--type google-vertex-ai \
--from-gcloud-adc \
--config VERTEX_AI_PROJECT_ID=my-gcp-project \
--config VERTEX_AI_REGION=us-central1
```
Use [Google Vertex AI](/providers/google-vertex-ai) for the full auth flows, including the production service-account refresh path, ADC-backed providers that mint `GOOGLE_VERTEX_AI_TOKEN`, and `--from-existing` support.
</Tab>
<Tab title="Local Endpoint">
```shell
openshell provider create \
--name my-local-model \
--type openai \
--credential OPENAI_API_KEY=empty-if-not-required \
--config OPENAI_BASE_URL=http://host.openshell.internal:11434/v1
```
Use `--config OPENAI_BASE_URL` to point to any OpenAI-compatible server running where the gateway runs. For host-backed local inference, use `host.openshell.internal` or the host's LAN IP. Avoid `127.0.0.1` and `localhost`. Set `OPENAI_API_KEY` to a dummy value if the server does not require authentication.
<Tip>
For a self-contained setup, the Ollama sandbox bundles Ollama inside the sandbox itself, so no host-level provider is needed. Refer to [Inference Ollama](/get-started/tutorials/inference-ollama) for details.
</Tip>
Ollama also supports cloud-hosted models using the `:cloud` tag suffix, for example `qwen3.5:cloud`.
</Tab>
<Tab title="Anthropic">
```shell
openshell provider create --name anthropic-prod --type anthropic --from-existing
```
This reads `ANTHROPIC_API_KEY` from your environment.
</Tab>
<Tab title="AWS Bedrock">
```shell
openshell provider create \
--name bedrock-bridge \
--type aws-bedrock \
--credential AWS_ACCESS_KEY_ID=unused-bridge-fronted-shape \
--config BEDROCK_BASE_URL=http://your-bedrock-bridge.your-ns.svc.cluster.local:8080
```
Then set the inference route, passing `--no-verify` because the validation probe does not yet support Bedrock protocols:
```shell
openshell inference set \
--provider bedrock-bridge \
--model anthropic.claude-3-5-sonnet-20241022-v2:0 \
--no-verify
```
**Why a placeholder credential?** `provider create` requires a non-empty `credentials` map even when the upstream auth scheme is `AuthHeader::None` — `aws-bedrock` falls into that bucket today because the router never injects a credential header on outbound requests; the bridge holds operator-side auth in its own pod. Any non-empty string value satisfies the structural requirement; `unused-bridge-fronted-shape` makes the intent obvious in `openshell provider get` output. The same pattern applies to any standalone-router profile that registers `AuthHeader::None`. When the SigV4 follow-up lands and the router begins signing requests itself, this becomes a real key.
**About the bridge-fronted shape.** The router does not inject any auth header on outbound requests. Point `BEDROCK_BASE_URL` at a translating bridge or Bedrock-compatible proxy that handles authentication in its own pod. The bridge is expected to accept Bedrock InvokeModel requests on the patterns listed above and forward to the operator's real upstream.
**About `--no-verify`.** The default validation probe does not yet recognize the `aws_bedrock_invoke` protocol, so without `--no-verify` the `inference set` call would fail before it could mint a route. The first sandbox round-trip is the real verification today.
**For direct AWS Bedrock**, refer to a future release that adds the SigV4 router-side signer. Until then, a `BEDROCK_BASE_URL` is required at provider-create time — the core profile sets `default_base_url: ""`, so route resolution rejects providers without it rather than silently forwarding prompts to AWS with no usable auth.
</Tab>
</Tabs>
## Set Inference Routing
Point `inference.local` at that provider and choose the model to use:
```shell
openshell inference set \
--provider nvidia-prod \
--model nvidia/nemotron-3-nano-30b-a3b
```
To override the default 60-second per-request timeout, add `--timeout`:
```shell
openshell inference set \
--provider nvidia-prod \
--model nvidia/nemotron-3-nano-30b-a3b \
--timeout 300
```
The value is in seconds. When `--timeout` is omitted or set to `0`, the default of 60 seconds applies. Increase `--timeout` when you expect extended thinking phases so the full response completes before the request deadline.
## Inspect and Update the Config
Confirm that the provider and model are set correctly:
```shell
openshell inference get
Gateway inference:
Provider: nvidia-prod
Model: nvidia/nemotron-3-nano-30b-a3b
Timeout: 300s
Version: 1
```
Use `update` when you want to change only one field:
```shell
openshell inference update --model nvidia/nemotron-3-nano-30b-a3b
openshell inference update --provider openai-prod
openshell inference update --timeout 120
```
## Use the Local Endpoint from a Sandbox
After inference is configured, code inside any sandbox can call `https://inference.local` directly. The client-supplied `model` and `api_key` values are not sent upstream — the privacy router injects the real credentials from the configured provider and rewrites the model before forwarding. Some SDKs require a non-empty API key even though `inference.local` does not use the sandbox-provided value; pass any placeholder such as `unused`.
<Tabs>
<Tab title="Claude Code">
```shell
ANTHROPIC_BASE_URL="https://inference.local" ANTHROPIC_API_KEY=unused claude --bare
```
`--bare` skips the OAuth login flow and uses `ANTHROPIC_API_KEY` directly. The key is stripped by the proxy and never reaches the upstream provider.
<Note>
Claude Code appends `/v1/messages` to `ANTHROPIC_BASE_URL`, so omit the `/v1` suffix from the base URL.
</Note>
</Tab>
<Tab title="OpenCode">
```shell
ANTHROPIC_BASE_URL="https://inference.local/v1" ANTHROPIC_API_KEY=unused opencode
```
<Note>
OpenCode appends `/messages` directly to `ANTHROPIC_BASE_URL`. Include the `/v1` suffix so the full path becomes `/v1/messages`, which matches the inference pattern.
</Note>
</Tab>
<Tab title="Python (OpenAI SDK)">
```python
from openai import OpenAI
client = OpenAI(base_url="https://inference.local/v1", api_key="unused")
response = client.chat.completions.create(
model="anything",
messages=[{"role": "user", "content": "Hello"}],
)
```
</Tab>
<Tab title="Python (Anthropic SDK)">
```python
import anthropic
client = anthropic.Anthropic(
base_url="https://inference.local",
api_key="unused",
)
message = client.messages.create(
model="anything",
max_tokens=1024,
messages=[{"role": "user", "content": "Hello"}],
)
```
</Tab>
</Tabs>
Use `inference.local` when inference should stay private and credentials should not be exposed inside the sandbox. External providers reached directly belong in `network_policies` instead.
When the upstream runs on the same machine as the gateway, bind it to `0.0.0.0` and point the provider at `host.openshell.internal` or the host's LAN IP. `127.0.0.1` and `localhost` usually fail because the request originates from the gateway or sandbox runtime, not from your shell.
If the gateway runs on a remote host or behind a cloud deployment, `host.openshell.internal` points to that remote machine, not to your laptop. A locally running Ollama or vLLM process is not reachable from a remote gateway unless you add your own tunnel or shared network path.
## Verify from a Sandbox
`openshell inference set` and `openshell inference update` verify the resolved upstream endpoint by default before saving the configuration. If the endpoint is not live yet, retry with `--no-verify` to persist the route without the probe.
To confirm end-to-end connectivity from a sandbox, run:
```shell
curl https://inference.local/v1/responses \
-H "Content-Type: application/json" \
-d '{
"instructions": "You are a helpful assistant.",
"input": "Hello!"
}'
```
A successful response confirms the privacy router can reach the configured backend and the model is serving requests.
- Gateway-scoped: Every sandbox using the active gateway sees the same `inference.local` backend.
- HTTPS only: `inference.local` is intercepted only for HTTPS traffic.
- Hot reload: Provider, model, and timeout changes are picked up by running sandboxes within about 5 seconds by default. No sandbox recreation is required.
## Next Steps
Explore related topics:
- To follow a complete Ollama-based local setup, refer to [Inference Ollama](/get-started/tutorials/inference-ollama).
- To follow a complete LM Studio-based local setup, refer to [Local Inference LM Studio](/get-started/tutorials/local-inference-lmstudio).
- To control external endpoints, refer to [Policies](/sandboxes/policies).
- To manage provider records, refer to [Providers](/sandboxes/manage-providers).
@@ -0,0 +1,199 @@
---
# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
title: "Manage Gateways"
sidebar-title: "Gateways"
description: "Register OpenShell gateways, switch between environments, inspect gateway status, and troubleshoot gateway access."
keywords: "Generative AI, Cybersecurity, Gateway, Docker, Podman, Kubernetes, MicroVM, CLI"
position: 2
---
The gateway is the control plane for OpenShell. All control-plane traffic between the CLI and running sandboxes flows through the gateway.
The gateway is responsible for:
- Provisioning and managing sandboxes, including creation, deletion, and status monitoring.
- Storing provider credentials and delivering them to sandboxes at startup.
- Delivering network and filesystem policies to sandboxes. Policy enforcement itself happens inside each sandbox through the proxy, OPA, Landlock, and seccomp.
- Managing inference configuration and serving inference bundles so sandboxes can route requests to the correct backend.
- Providing the SSH tunnel endpoint so you can connect to sandboxes without exposing them directly.
OpenShell separates gateway access from the compute driver that runs sandboxes. Use [Installation](/about/installation) to install OpenShell, choose a compute driver, and start a gateway. This page covers working with gateway entries after a gateway exists.
## Gateway Compute Drivers
A gateway provisions sandboxes through the compute driver configured for that gateway.
| Compute Driver | Where sandboxes run | Best for |
|---|---|---|
| Docker | Containers on the gateway host. | Solo development, quick iteration, and single-machine gateways. |
| Podman | Rootless containers on the gateway host. | Workstations that avoid a rootful Docker daemon. |
| Kubernetes | Pods in an operator-managed cluster. | Shared clusters and cloud environments. |
| MicroVM | VM-backed sandboxes. | Workflows that need VM-backed isolation. |
All compute drivers expose the same gateway API surface. Sandboxes, policies, and providers work the same after the CLI registers the gateway endpoint. The difference is how the gateway creates sandbox workloads and how operators expose the gateway to users.
<Tip>
For driver setup, including Docker, Podman, MicroVM, and Kubernetes paths, refer to [Installation](/about/installation).
</Tip>
## Configure Service Forwarding
Sandbox service routing is enabled for gateways by default. Users expose long-running sandbox services with `openshell service expose`, and the gateway routes browser traffic from the printed service URL to the loopback port inside the sandbox.
Loopback gateways use `openshell.localhost` service URLs. OpenShell prints a URL in the form `http://<sandbox>.openshell.localhost:<port>/` or `http://<sandbox>--<service>.openshell.localhost:<port>/` when a service name is provided.
Browser traffic enters the same gateway listener as mTLS-protected gRPC, but plaintext HTTP is accepted only from loopback clients and only for sandbox service hostnames. Gateway APIs, auth routes, health endpoints, and non-service hostnames remain unavailable over plaintext HTTP. Cross-origin and sibling-subdomain browser requests are rejected before reaching the sandbox service.
Disable the local browser path with `--enable-loopback-service-http=false` or `OPENSHELL_ENABLE_LOOPBACK_SERVICE_HTTP=false`.
Custom HTTPS service domains use the gateway server SAN configuration. Add a wildcard DNS SAN such as `*.apps.example.com` to the gateway certificate and pass the same SAN to the gateway with `--server-san` or `OPENSHELL_SERVER_SAN`.
For remote or non-loopback gateways, browser service URLs remain HTTPS and require normal gateway authentication.
## Register an Existing Gateway
Use `openshell gateway add` to register any reachable gateway endpoint so the CLI can target it.
Register a plaintext local endpoint, such as a trusted port-forward:
```shell
openshell gateway add http://127.0.0.1:8080 --local --name local
```
Register a gateway behind an authenticated reverse proxy:
```shell
openshell gateway add https://gateway.example.com --name production
```
This opens your browser for the proxy's login flow when the gateway uses edge authentication. If the token expires later, re-authenticate with:
```shell
openshell gateway login production
```
For direct mTLS endpoints, place the CLI client certificate bundle in the gateway credential directory described in [Gateway Authentication](/reference/gateway-auth), then register or select that gateway name.
## Manage Multiple Gateways
One gateway is always the active gateway. All CLI commands target it by default. `gateway add` sets the new gateway as active.
The active gateway is the persisted default. The `-g` flag and the `OPENSHELL_GATEWAY` environment variable override it when commands resolve a gateway. If `OPENSHELL_GATEWAY` is set to a different gateway, `openshell gateway select <name>` still saves the new default and warns that the current shell continues to use the environment value until you unset or update it.
Installers can seed read-only gateway entries for package-managed local services. By default the CLI reads these from `/etc/openshell`, using the same `active_gateway` plus `gateways/<name>/metadata.json` layout as per-user config. Packages can override that system config root with a non-empty absolute `OPENSHELL_SYSTEM_GATEWAY_DIR` when needed; empty or relative values fall back to `/etc/openshell` and log a warning. These entries appear in `openshell gateway list` and can be selected like user registrations. `openshell gateway list` and `openshell term` label each gateway as `user` or `system` so you can see which config layer owns it. `openshell gateway remove` removes only per-user registrations. Register a per-user gateway with the same name when you need to shadow an installer-provided default.
List all registered gateways. The table shows each gateway's endpoint, type,
config source, authentication mode, and remote target when a remote registration
has one:
```shell
openshell gateway list
```
Use structured output when you need the complete local registration metadata,
including `remote_host` and `resolved_host` for remote registrations:
```shell
openshell gateway list -o json
```
Switch the active gateway:
```shell
openshell gateway select production
```
Override the active gateway for a single command with `-g`:
```shell
openshell status -g staging
```
## Inspect Gateway Status
Use `openshell status` for a live gateway check. It reports gateway
reachability, authentication, and the gateway version independently:
```shell
openshell status
```
`Status: Connected` means the public health endpoint responded.
`Authentication: Authenticated` confirms the configured credentials also
passed the gateway authentication layer. If authentication fails while status
remains connected, run `openshell gateway login <name>` to refresh the stored
credentials.
For automation or scripting, use `--output json` or `--output yaml` to get machine-readable output:
```shell
openshell status --output json
```
Use `openshell gateway info` when you need elevated runtime details such as
initialized compute drivers and driver-reported capability versions:
```shell
openshell gateway info
```
Use structured output when scripting against live runtime info:
```shell
openshell gateway info -o json
```
Remove a local CLI registration without stopping the gateway service:
```shell
openshell gateway remove production
```
## Troubleshoot
Check gateway health:
```shell
openshell status
openshell gateway info
```
For Docker-backed local gateways, inspect Docker and the gateway process or container started by your local workflow:
```shell
openshell doctor check
openshell gateway list
```
For Kubernetes gateways, inspect the gateway workload and cluster events:
```shell
kubectl -n openshell get deployment,statefulset,pods
kubectl -n openshell logs deployment/openshell -c openshell-gateway --tail=100
kubectl -n openshell logs statefulset/openshell -c openshell-gateway --tail=100
kubectl -n openshell get events --sort-by=.lastTimestamp
```
For Podman or MicroVM gateways managed by systemd, inspect the user service and logs:
```shell
systemctl --user status openshell-gateway
journalctl --user -u openshell-gateway --no-pager -n 50
```
For sandbox startup failures, inspect the selected compute driver:
| Compute Driver | What to check |
|---|---|
| Docker | Docker daemon health, image availability, gateway logs, and sandbox container state. |
| Podman | Podman socket availability, rootless networking, image availability, and sandbox container state. |
| Kubernetes | Events and sandbox pods in the namespace configured by `server.sandboxNamespace`. |
| MicroVM | VM driver logs, rootfs availability, and gateway logs. |
## Next Steps
- To install OpenShell and choose a compute driver, refer to [Installation](/about/installation).
- To configure workspace membership and roles, refer to [Manage Workspaces and Access](/sandboxes/manage-workspaces).
- To create a sandbox using the gateway, refer to [Manage Sandboxes](/sandboxes/manage-sandboxes).
@@ -0,0 +1,460 @@
---
# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
title: "Providers"
sidebar-title: "Providers"
description: "Create and manage credential providers that inject API keys and tokens into OpenShell sandboxes."
keywords: "Generative AI, Cybersecurity, Providers, Credentials, API Keys, Sandbox, Security"
position: 4
---
AI agents typically need credentials to access external services: an API key for the AI model provider, a token for GitHub or GitLab, and so on. OpenShell manages these credentials as first-class entities called *providers*.
Create and manage providers that supply credentials to sandboxes.
<Info>
Providers v2 is available for profile-backed provider policy, provider-owned network rules, and gateway-managed credential refresh. This page remains the credential-focused provider command reference. For the new workflow, see [Providers v2](/sandboxes/providers-v2).
</Info>
Provider profiles include metadata for known endpoints and binaries. View
the available profiles before creating a provider:
```shell
openshell provider list-profiles
```
## Create a Provider
Providers can be created from local environment variables or with explicit credential values.
For refresh-backed providers such as `google-vertex-ai --from-gcloud-adc`, `openshell provider create` now waits for the gateway to configure refresh metadata and mint the initial access token before it reports success.
### From Local Credentials
The fastest way to create a provider is to let the CLI discover credentials from
your shell environment:
```shell
openshell provider create --name my-claude --type claude --from-existing
```
This reads `ANTHROPIC_API_KEY` or `CLAUDE_API_KEY` from your current environment
and stores them in the provider.
### With Explicit Credentials
Supply a credential value directly:
```shell
openshell provider create --name my-nvidia --type nvidia --credential NVIDIA_API_KEY=nvapi-example
```
### Bare Key Form
Pass a key name without a value to read the value from the environment variable
of that name:
```shell
openshell provider create --name my-nvidia --type nvidia --credential NVIDIA_API_KEY
```
This looks up the current value of `$NVIDIA_API_KEY` in your shell and stores it.
### From the Current OIDC Login
Profile-backed token-exchange providers can store the current gateway OIDC
access token as their subject credential. The subject credential stays
gateway-only and is not emitted into sandbox environment material or static
credential bindings:
```shell
openshell provider create \
--name custom-api \
--type custom-api \
--from-oidc-token
```
OpenShell infers the destination credential from the provider profile when the
profile has exactly one `token_grant.subject_token.credential`. If the profile
has more than one, pass `--credential <key>`. Refresh the stored subject token
later with:
```shell
openshell provider update custom-api --from-oidc-token
```
This copies the current OIDC access token and expiry from the active gateway
login. This requires an active named gateway that was registered for OIDC. If
the stored gateway access token is expired and a refresh token is available, the
CLI refreshes it before storing the provider credential. It does not store the
OIDC refresh token in the provider.
Provider profile metadata is available for known provider types. Provider profile
network policy is gateway opt-in:
```shell
openshell settings set --global --key providers_v2_enabled --value true
```
Without `providers_v2_enabled=true`, attached provider profiles do not contribute
network policy to the sandbox. Static credential endpoint binding remains active
in either mode.
When `providers_v2_enabled=true`, `--from-existing` uses profile-backed
discovery instead of the legacy provider registry. The requested `--type` must
have a built-in or imported provider profile with a `discovery` section. If no
matching profile exists, the CLI returns an error instead of falling back to
legacy discovery.
Update a custom provider profile after exporting it, editing its endpoints,
binaries, or credential metadata, and preserving the exported `resource_version`:
```shell
openshell provider profile export my-api -o yaml > my-api-profile.yaml
openshell provider profile update my-api -f my-api-profile.yaml
```
Import remains create-only and fails if the profile ID already exists. Use
`provider profile update <id>` for existing custom profiles. Built-in profiles are
read-only. The target ID must match the profile ID in the file. Update accepts
one file at a time and rejects stale resource versions. When
`providers_v2_enabled=true`, updated profile policy applies to all provider
instances of that type on the next sandbox config sync.
<Warning>
Static credentials require a built-in or imported provider profile. Profiles
normally define at least one endpoint. An endpointless profile requires an
explicit `credential_binding.provider` on a sandbox policy endpoint. OpenShell
does not activate static credentials from a profileless provider because it
cannot determine which credential definition applies.
</Warning>
## Manage Providers
List, inspect, update, and delete providers from the active gateway.
List all providers:
```shell
openshell provider list
```
Use `-o json` or `-o yaml` for machine-readable output:
```shell
openshell provider list -o json
openshell provider list -o yaml
```
Structured output includes provider metadata (`id`, `name`, `type`), credential key names, config key names, labels, creation timestamp, resource version, and credential expiration times. Only credential and config keys are exposed, never values, preventing accidental credential leakage in logs or output. For details on how credentials are injected into sandboxes, refer to [Credential Injection](#how-credential-injection-works).
Inspect a provider:
```shell
openshell provider get my-claude
```
Update a provider's credentials:
```shell
openshell provider update my-claude --from-existing
```
Set or clear a credential expiry timestamp:
```shell
openshell provider update my-graph \
--credential MS_GRAPH_ACCESS_TOKEN="$MS_GRAPH_ACCESS_TOKEN" \
--credential-expires-at MS_GRAPH_ACCESS_TOKEN=1767225600000
```
Use `0` as the timestamp to clear expiry for a credential key.
## Credential Refresh
Provider refresh stores non-injectable refresh material separately from the
provider's current credential values. The gateway can mint OAuth2 refresh-token
tokens, OAuth2 client credentials tokens, Google service account JWT tokens, and
AWS STS temporary credentials, then write the current access token back to the
provider record for sandbox injection.
Configure refresh metadata for one injectable credential key:
```shell
openshell provider refresh configure my-graph \
--credential-key MS_GRAPH_ACCESS_TOKEN \
--strategy oauth2-client-credentials \
--material tenant_id="$TENANT_ID" \
--material client_id="$CLIENT_ID" \
--secret-material-env client_secret=CLIENT_SECRET \
--credential-expires-at 1767225600000
```
Pass secret material with `--secret-material-env KEY[=ENVVAR]` (`ENVVAR`
defaults to `KEY`): the CLI reads the value from its own environment, so the
secret never appears in the host process table the way an expanded
`--material KEY="$VALUE"` argument would, and `KEY` is automatically marked
secret. Keep non-secret material on `--material`; `--secret-material-key KEY`
still marks a key supplied through `--material` as secret.
Check refresh status:
```shell
openshell provider refresh status my-graph
```
Delete refresh metadata for a credential:
```shell
openshell provider refresh delete my-graph \
--credential-key MS_GRAPH_ACCESS_TOKEN
```
Force a gateway-managed refresh for one credential:
```shell
openshell provider refresh rotate my-graph --credential-key MS_GRAPH_ACCESS_TOKEN
```
### AWS STS
AWS STS refresh requires `providers_v2_enabled=true`. The gateway calls
`sts:AssumeRole` and writes three short-lived credentials (`AWS_ACCESS_KEY_ID`,
`AWS_SECRET_ACCESS_KEY`, `AWS_SESSION_TOKEN`) to the provider record atomically.
The `aws` and `aws-s3` profiles declare `AWS_SECRET_ACCESS_KEY` and
`AWS_SESSION_TOKEN` as `additional_outputs` of the `AWS_ACCESS_KEY_ID` refresh,
so all three credentials are gateway-minted. Create the provider with
`--runtime-credentials` — no placeholder credential is needed.
```shell
openshell settings set --global --key providers_v2_enabled --value true --yes
openshell provider create --name my-aws --type aws-s3 --runtime-credentials
openshell provider refresh configure my-aws \
--credential-key AWS_ACCESS_KEY_ID \
--strategy aws-sts-assume-role \
--material role_arn="arn:aws:iam::123456789012:role/SandboxS3Writer" \
--material session_name="openshell-sandbox"
openshell provider refresh rotate my-aws --credential-key AWS_ACCESS_KEY_ID
```
The gateway resolves its own AWS credentials using the default credential chain
(instance role, IRSA, ECS task role, environment variables). For gateways not
running on AWS, provide explicit IAM keys via refresh material:
```shell
openshell provider refresh configure my-aws \
--credential-key AWS_ACCESS_KEY_ID \
--strategy aws-sts-assume-role \
--material role_arn="arn:aws:iam::123456789012:role/SandboxS3Writer" \
--material aws_access_key_id="$AWS_ACCESS_KEY_ID" \
--secret-material-env aws_secret_access_key=AWS_SECRET_ACCESS_KEY
```
Pass the long-lived secret with `--secret-material-env` so the CLI reads it
from its own environment instead of expanding it into a process argument,
where it would be visible in the host process table.
For temporary source credentials (AWS SSO or a prior `AssumeRole`), also pass a
session token with `--secret-material-env aws_session_token=AWS_SESSION_TOKEN`.
It requires the `aws_access_key_id` and `aws_secret_access_key` pair.
Use the generic `aws` profile type for multi-service access and scope endpoints
via sandbox network policy. Use `aws-s3` for S3-specific endpoint rules.
The proxy re-signs requests using SigV4 before forwarding to AWS. Both curl and
Python boto3 are supported.
External refresh systems should continue to push new current credentials through
`openshell provider update`. The `--credential-expires-at` option works for
static credentials, externally refreshed credentials, and gateway-managed
refresh strategies.
Delete a provider:
```shell
openshell provider delete my-claude
```
## Attach Providers to Sandboxes
Pass one or more `--provider` flags when creating a sandbox:
```shell
openshell sandbox create --provider my-claude --provider my-github -- claude
```
Each `--provider` flag attaches one provider. The sandbox receives eligible
credentials from every attached provider as placeholders at runtime. Each
static credential resolves only for endpoints in that provider's profile, or
for explicitly bound sandbox policy endpoints when the profile is endpointless.
Profile-managed providers also contribute provider-generated network policy
entries when `providers_v2_enabled` is enabled at the gateway. When the setting
is disabled, endpoint binding still applies, but provider-generated policy does
not.
<Warning>
Legacy provider attachment is fixed at sandbox creation time. Providers v2 adds
`openshell sandbox provider attach` and `openshell sandbox provider detach` for
running sandboxes. See [Providers v2](/sandboxes/providers-v2#attach-and-detach-providers)
for runtime attach and detach behavior.
</Warning>
### Auto-Discovery Shortcut
When `providers_v2_enabled=false` and the trailing command in
`openshell sandbox create` is a recognized tool name (`claude`, `codex`, or
`opencode`), the CLI auto-creates the required provider from your local
credentials if one does not already exist. You do not need to create the
provider separately:
```shell
openshell sandbox create -- claude
```
This detects `claude` as a known tool, finds your `ANTHROPIC_API_KEY`, creates
a provider, attaches it to the sandbox, and launches Claude Code.
Providers v2 disables command-derived provider inference. When
`providers_v2_enabled=true`, create or import the provider profile, create the
provider instance, and pass `--provider <name>` explicitly.
## How Credential Injection Works
The agent process inside the sandbox never sees real credential values. At
startup, OpenShell replaces each credential with an opaque placeholder token in
the agent's environment. When the agent sends an HTTP request containing a
placeholder, the proxy resolves it immediately before forwarding the request.
Static credential resolution has two independent authorization boundaries:
1. Network policy must allow the calling binary and request destination.
2. The credential binding must include the request host, port, and path. Profile
endpoints provide the binding by default. An endpointless profile can use a
sandbox policy endpoint that names the attached provider instance through
`credential_binding.provider`.
Both checks must pass. A provider profile endpoint does not grant network access
unless provider policy composition or the sandbox's own policy allows the
request. A plain sandbox policy endpoint does not grant credential use unless
the profile already covers it or the endpoint explicitly binds an endpointless
provider.
Endpoint binding applies whether `providers_v2_enabled` is enabled or disabled.
The setting controls provider policy composition only. Every static credential
declared by a provider receives the complete endpoint set from that provider's
profile, or the explicitly bound sandbox policy endpoints for an endpointless
profile.
Credential resolution requires the proxy to handle the request as HTTP. Raw
`tls: skip` and non-HTTP tunnels remain opaque and do not support credential
rewrite. Refer to [Providers v2](/sandboxes/providers-v2#understand-static-credential-endpoint-binding)
for endpoint matching and migration guidance.
### Supported injection locations
The proxy resolves credential placeholders in the following parts of an HTTP request:
| Location | How the agent uses it | Example |
|---|---|---|
| Header value | Agent reads `$API_KEY` from env and places it in a header. | `Authorization: Bearer <placeholder>` |
| Header value (Basic auth) | Agent base64-encodes `user:<placeholder>` in an `Authorization: Basic` header. The proxy decodes, resolves, and re-encodes. | `Authorization: Basic <base64>` |
| Query parameter value | Agent places the placeholder in a URL query parameter. | `GET /api?key=<placeholder>` |
| URL path segment | Agent builds a URL with the placeholder in the path. Supports concatenated patterns. | `POST /bot<placeholder>/sendMessage` |
| Supported request body | An inspected REST endpoint opts in with `request_body_credential_rewrite: true`. | `{"api_key":"<placeholder>"}` |
| WebSocket text message | A REST or WebSocket endpoint opts in with `websocket_credential_rewrite: true`. | `{"token":"<placeholder>"}` |
| AWS SigV4 signing | An endpoint configures `credential_signing`, and the proxy signs with the endpoint-bound AWS credentials. | `credential_signing: sigv4` |
The proxy does not rewrite cookies, response content, unsupported request
bodies, or WebSocket binary frames.
### Fail-closed behavior
If policy allows a request but the credential binding does not include the
request endpoint, the proxy rejects the request with HTTP 403 and the
`credential_endpoint_mismatch` reason. It emits a denied activity event and a
security finding without recording the secret, placeholder, environment key, or
query string.
Unknown, malformed, expired, or otherwise unresolved placeholders also fail
closed instead of being forwarded to the upstream service.
### Inspect an Endpoint Binding
Export the profile used by a provider to inspect its credential boundary:
```shell
openshell provider profile export github -o yaml
```
For the built-in GitHub profile, `GITHUB_TOKEN` and `GH_TOKEN` can resolve for
the profile's `api.github.com:443` and `github.com:443` endpoints. Even if a
sandbox policy allows `uploads.example.com:443`, sending either placeholder
there returns `credential_endpoint_mismatch`.
For a custom service, define the credential in a custom provider profile. Put
stable endpoints in the profile, or leave the profile endpointless and bind
each concrete provider instance from sandbox policy. Refer to [Provider
Profiles](/sandboxes/providers-v2#provider-profiles) for the profile workflow
and schema.
## Supported Provider Types
The following provider types are supported.
| Type | Environment Variables Injected | Typical Use |
|---|---|---|
| `anthropic` | `ANTHROPIC_API_KEY` | Anthropic API |
| `aws-bedrock` | `AWS_ACCESS_KEY_ID`, `AWS_SECRET_ACCESS_KEY`, `AWS_SESSION_TOKEN`, `AWS_REGION` | AWS Bedrock InvokeModel via a translating bridge. Today the router does not inject any auth header; the configured `BEDROCK_BASE_URL` upstream is expected to handle auth itself. Refer to [Inference Routing](/sandboxes/inference-routing). |
| `claude` | `ANTHROPIC_API_KEY`, `CLAUDE_API_KEY` | Claude Code, Anthropic API |
| `codex` | `OPENAI_API_KEY` | OpenAI Codex |
| `copilot` | `COPILOT_GITHUB_TOKEN`, `GH_TOKEN`, `GITHUB_TOKEN` | GitHub Copilot CLI |
| `deepinfra` | `DEEPINFRA_API_KEY` | DeepInfra inference API |
| `generic` | User-defined | Legacy credential storage. Import an endpoint-bearing custom profile for credentials attached to a sandbox. |
| `github` | `GITHUB_TOKEN`, `GH_TOKEN` | GitHub API and `gh` CLI. Refer to [GitHub Sandbox](/get-started/tutorials/github-sandbox). |
| `gitlab` | `GITLAB_TOKEN`, `GLAB_TOKEN`, `CI_JOB_TOKEN` | GitLab API, `glab` CLI |
| `nvidia` | `NVIDIA_API_KEY` | NVIDIA API Catalog |
| `openai` | `OPENAI_API_KEY` | Any OpenAI-compatible endpoint. Set `--config OPENAI_BASE_URL` to point to the provider. Refer to [Inference Routing](/sandboxes/inference-routing). |
| `opencode` | `OPENCODE_API_KEY`, `OPENROUTER_API_KEY`, `OPENAI_API_KEY` | OpenCode |
<Note>
`ANTHROPIC_API_KEY` is an API key from [console.anthropic.com](https://console.anthropic.com), not a subscription token. Subscription users must generate a separate API key from the Anthropic Console.
</Note>
<Tip>
For a service not listed above, import a custom provider profile that declares
its credential environment variables and endpoints. Use the imported profile
ID as the provider `--type`.
</Tip>
## Supported Inference Providers
The following providers have been tested with `inference.local`. Any provider that exposes an OpenAI-compatible API works with the `openai` type. Set `--config OPENAI_BASE_URL` to the provider's base URL and `--credential OPENAI_API_KEY` to your API key.
| Provider | Name | Type | Base URL | API Key Variable |
|---|---|---|---|---|
| AWS Bedrock (via bridge) | `bedrock-bridge` | `aws-bedrock` | Operator-supplied `BEDROCK_BASE_URL` | None at router level (bridge holds creds) |
| NVIDIA API Catalog | `nvidia-prod` | `nvidia` | `https://integrate.api.nvidia.com/v1` | `NVIDIA_API_KEY` |
| Anthropic | `anthropic-prod` | `anthropic` | `https://api.anthropic.com` | `ANTHROPIC_API_KEY` |
| Google Vertex AI | `vertex-prod` | `google-vertex-ai` | Regional, global, or multi-region Vertex endpoint | `GOOGLE_VERTEX_AI_TOKEN` or `GOOGLE_VERTEX_AI_SERVICE_ACCOUNT_TOKEN` |
| Baseten | `baseten` | `openai` | `https://inference.baseten.co/v1` | `OPENAI_API_KEY` |
| Bitdeer AI | `bitdeer` | `openai` | `https://api-inference.bitdeer.ai/v1` | `OPENAI_API_KEY` |
| DeepInfra | `deepinfra` | `deepinfra` | `https://api.deepinfra.com/v1/openai` | `DEEPINFRA_API_KEY` |
| Groq | `groq` | `openai` | `https://api.groq.com/openai/v1` | `OPENAI_API_KEY` |
| Ollama (local) | `ollama` | `openai` | `http://host.openshell.internal:11434/v1` | `OPENAI_API_KEY` |
| LM Studio (local) | `lmstudio` | `openai` | `http://host.openshell.internal:1234/v1` | `OPENAI_API_KEY` |
Refer to your provider's documentation for the correct base URL, available models, and API key setup. For the Vertex-specific auth flows and config keys, refer to [Google Vertex AI](/providers/google-vertex-ai). To configure inference routing, refer to [Inference Routing](/sandboxes/inference-routing).
## Next Steps
Explore related topics:
- To manage workspace access for providers, refer to [Manage Workspaces and Access](/sandboxes/manage-workspaces).
- To control what the agent can access, refer to [Policies](/sandboxes/policies).
- To use the base sandbox container, refer to [Sandboxes](/sandboxes/manage-sandboxes#base-sandbox-container).
- To view the complete field reference for the policy YAML, refer to the [Policy Schema Reference](/reference/policy-schema).
@@ -0,0 +1,542 @@
---
# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
title: "Manage Sandboxes"
sidebar-title: "Sandboxes"
description: "Create sandboxes, understand sandbox isolation, and manage the full sandbox lifecycle."
keywords: "Generative AI, Cybersecurity, Gateway, Sandboxing, AI Agents, Sandbox Management, CLI"
position: 1
---
A sandbox is the OpenShell data plane: a safe, private execution environment where an AI agent runs. Each sandbox combines runtime isolation with OpenShell policy controls that prevent unauthorized data access, credential exposure, and network exfiltration.
<Info>You need an active gateway before creating a sandbox.</Info>
## Create a Sandbox
Create a sandbox with a single command. For example, to create a sandbox with Claude, run:
```shell
openshell sandbox create -- claude
```
The trailing command is the sandbox's canonical main process. OpenShell starts
it once and attaches your terminal to it. With no trailing command, OpenShell
starts `/bin/bash -l` in a retained pseudo-terminal. Add `--detach` to create
the sandbox without attaching:
```shell
openshell sandbox create --name worker --detach -- ./worker
```
`--upload` cannot yet be combined with a trailing main command because uploads
finish after the canonical process starts. Create a scratch sandbox, upload the
files, then launch the workload with `sandbox exec`, or build the files into the
sandbox image.
For automation, use `--output json` or `--output yaml` to get machine-readable sandbox metadata after creation:
```shell
openshell sandbox create --output json
```
Every sandbox requires a gateway. Register or select one before running sandbox commands:
```shell
openshell gateway add http://127.0.0.1:18080 --local --name local
openshell gateway select local
```
### CPU and Memory
Set per-sandbox CPU and memory amounts with `--cpu` and `--memory`:
```shell
openshell sandbox create --cpu 2 --memory 4Gi -- claude
```
CPU values use Kubernetes-style quantities such as `500m`, `1`, or `2.5`.
Memory values use byte quantities such as `512Mi`, `4Gi`, or `8G`.
Docker and Podman apply these values as runtime limits. Kubernetes applies each
value as both the request and the limit so the scheduler reserves the same
amount the sandbox can use. The VM driver currently accepts these flags but
does not change VM allocation.
### Driver-Specific Configuration
Pass experimental driver-owned settings with `--driver-config-json`. The value
must be a JSON object keyed by driver name. The gateway forwards only the block
for its configured compute driver:
Nested keys inside each driver block use snake_case. The top-level envelope keys
are driver names, such as `kubernetes`, and are not part of the nested schema.
```shell
openshell sandbox create \
--driver-config-json '{"kubernetes":{"pod":{"runtime_class_name":"kata-containers","node_selector":{"pool":"gpu"}}}}' \
-- claude
```
Use this only for driver-specific fields that do not have a stable CLI flag.
Prefer stable flags such as `--cpu`, `--memory`, and `--gpu` when they cover
the same behavior.
### GPU Resources
To request GPU resources, add `--gpu`:
```shell
openshell sandbox create --gpu -- claude
```
Request a specific number of GPUs by passing a count to `--gpu`:
```shell
openshell sandbox create --gpu 2 -- claude
```
When you omit the count, OpenShell treats the request as `--gpu 1`.
Kubernetes honors counted `--gpu` requests by setting the `nvidia.com/gpu`
limit. Docker and Podman select the requested number of default NVIDIA CDI
devices in round-robin order. VM gateways accept only one GPU, either through
`--gpu` or `--gpu 1`; a single `gpu_device_ids` entry works with either form.
For Docker-backed sandboxes, GPU injection uses Docker CDI. If you enable Docker
CDI after the gateway starts, restart the gateway so OpenShell can detect the
updated Docker daemon capability.
Docker and Podman refresh the CDI inventory before validating or creating a
default GPU request, so CDI devices added or removed after the driver starts can
be reflected in later sandbox creates. On WSL2 all-only runtimes, the default
can fall back to `nvidia.com/gpu=all`; that fallback counts as one selectable
device.
Exact GPU device selection is driver-specific and still requires `--gpu`. For
Docker or Podman, pass CDI IDs through `cdi_devices`. The top-level key must
match the active driver; replace `docker` with `podman` when using Podman. CDI
IDs are treated as opaque strings. The list must not contain duplicate IDs, and
its length must match the effective GPU count:
```shell
openshell sandbox create \
--gpu \
--driver-config-json '{"docker":{"cdi_devices":["nvidia.com/gpu=0"]}}' \
-- claude
```
### Custom Containers
Use `--from` to create a sandbox from the base image, another pre-built sandbox name, a local directory, or a container image:
```shell
openshell sandbox create --from base
openshell sandbox create --from ollama
openshell sandbox create --from ./my-sandbox-dir
openshell sandbox create --from my-registry.example.com/my-image:latest
```
Bare names such as `base` and `ollama` resolve to images under `ghcr.io/nvidia/openshell-community/sandboxes`. Set `OPENSHELL_COMMUNITY_REGISTRY` when you need to use an internal mirror.
Local directories and Dockerfiles require a local gateway because the CLI builds
through the local Docker daemon. Use a registry image reference for remote
gateways.
## Base Sandbox Container
The `base` sandbox container is the default runtime image for standard OpenShell sandboxes unless the gateway overrides its default sandbox image. It is published as `ghcr.io/nvidia/openshell-community/sandboxes/base:latest` and maintained in the [OpenShell Community](https://github.com/NVIDIA/OpenShell-Community/tree/main/sandboxes/base) repository.
The base container includes common development tooling, supported agent CLIs, and the default sandbox policy. Use it when you want a general-purpose agent environment without a workflow-specific image:
```shell
openshell sandbox create --from base
```
For default policy coverage by agent, refer to [Default Policy](/reference/default-policy). For the supported agent list, refer to [Supported Agents](/about/supported-agents).
## Connect to a Sandbox
Attach to the canonical main process in a running sandbox:
```shell
openshell sandbox connect my-sandbox
```
Disconnecting does not stop the process or close its stdin. A later `connect`
attaches to the same process instance and replays up to 1 MiB of recent
output. One attachment owns stdin at a time. Use `sandbox exec --tty --
/bin/bash -l` when you want a new independent shell instead.
Press `Ctrl-P`, then `Ctrl-Q` to disconnect without terminating the main
process. `Ctrl-C` retains its normal terminal behavior and interrupts the
foreground process.
Launch VS Code or Cursor directly into the sandbox workspace:
```shell
openshell sandbox create --editor vscode --name my-sandbox
openshell sandbox connect my-sandbox --editor cursor
```
When `--editor` is used, OpenShell keeps the sandbox alive and installs an
OpenShell-managed SSH include file instead of cluttering your main
`~/.ssh/config` with generated host blocks.
## Execute a Command in a Sandbox
Run a one-shot command inside a running sandbox without opening an interactive shell:
```shell
openshell sandbox exec -n my-sandbox -- ls -la /workspace
```
Pipe stdin into the command:
```shell
echo "hello" | openshell sandbox exec -n my-sandbox -- cat
```
The command's exit code is propagated to the CLI, so `exec` works in scripts that check return codes.
Run an interactive shell with a TTY:
```shell
openshell sandbox exec -n my-sandbox --tty -- /bin/bash
```
OpenShell allocates a TTY automatically when both stdin and stdout are terminals. Force the behavior with `--tty` or disable it with `--no-tty`.
| Flag | Purpose |
| -------------- | -------------------------------------------------------- |
| `-n`, `--name` | Sandbox to target. |
| `--workdir` | Working directory for the command inside the sandbox. |
| `--timeout` | Command timeout in seconds. `0` disables the timeout. |
| `--tty` | Force TTY allocation. |
| `--no-tty` | Disable TTY allocation even when attached to a terminal. |
| `--env` | Set an environment variable for the command (`KEY=VALUE`, repeatable). |
## Set Environment Variables
Inject environment variables into the sandbox at creation time:
```shell
openshell sandbox create --env API_KEY=sk-test --env DEBUG=1 -- my-agent
```
Variables set with `--env` are available to all processes in the sandbox, including interactive shells and exec commands.
When an `--env` key looks like a credential — a known provider variable, or a name whose underscore-separated segments include a credential word such as `TOKEN`, `SECRET`, `PASSWORD`, `CREDENTIAL`, `API_KEY`, `ACCESS_KEY`, or `SECRET_KEY` (for example `DB_TOKEN` or `MY_ACCESS_KEY`) — `sandbox create` prints a non-blocking warning. Matching is on whole segments, so unrelated names like `TOKENIZERS_PARALLELISM` or `PASSWORDLESS_LOGIN` do not warn. The agent inside the sandbox can read plain environment values directly, so to hide a secret from the agent, attach it through a [provider](/sandboxes/providers-v2) with `--provider` instead. Suppress the warning with `--no-credential-warnings`. Detection uses the key name only; values are never inspected or printed.
You can also set per-command environment variables with `sandbox exec`:
```shell
openshell sandbox exec -n my-sandbox --env MY_VAR=hello -- printenv MY_VAR
```
Per-command variables override the sandbox-level environment for that command only.
Environment variable names starting with `OPENSHELL_` are reserved. Keys must match `[A-Za-z_][A-Za-z0-9_]*`.
## Label a Sandbox
Attach labels when you create a sandbox to track ownership, environment, or workflow grouping:
```shell
openshell sandbox create --label env=dev --label team=platform -- claude
```
List only the sandboxes that match a label selector:
```shell
openshell sandbox list --selector env=dev
openshell sandbox list --selector env=dev,team=platform
```
The Python SDK accepts the same gateway labels and selectors. Labels passed to
`create` are stored on the gateway sandbox object and are returned on
`SandboxRef.labels`, so selector-based listing finds Python-created sandboxes:
```python
from openshell import SandboxClient
with SandboxClient.from_active_cluster() as client:
sandbox = client.create(workspace="default", name="deep-research-1", labels={"env": "dev", "team": "platform"})
assert sandbox.labels["team"] == "platform"
matches = client.list(workspace="default", label_selector="env=dev,team=platform")
assert sandbox.id in {s.id for s in matches}
```
For non-interactive automation, pass a renewable client-credentials provider.
Omitted issuer, client ID, audience, and scopes are read from the active
gateway's metadata. The client requires TLS for non-loopback gateways:
```python
from openshell import ClientCredentialsAuth, SandboxClient
auth = ClientCredentialsAuth(client_secret=lambda: load_secret())
with SandboxClient.from_active_cluster(client_credentials=auth) as client:
sandboxes = client.list(workspace="default")
```
## Expose Long Running Services
Service forwarding makes a long-running process inside a sandbox reachable through a gateway-managed URL. Use it for development servers, notebooks, dashboards, or other services that keep listening after the sandbox starts. Run the service on loopback inside the sandbox, expose its port, then open the URL printed by OpenShell.
Expose a service that listens on loopback inside the sandbox:
```shell
openshell service expose my-sandbox 8080
```
Pass an optional service name to create a named service URL:
```shell
openshell service expose my-sandbox 8080 web
```
List exposed endpoints:
```shell
openshell service list
```
List endpoints for one sandbox:
```shell
openshell service list my-sandbox
```
Show or delete one endpoint:
```shell
openshell service get my-sandbox web
openshell service delete my-sandbox web
```
Omit the service name to manage the unnamed endpoint:
```shell
openshell service get my-sandbox
openshell service delete my-sandbox
```
<Note>
Loopback gateways return local `openshell.localhost` URLs. Remote gateways return HTTPS URLs that require normal gateway authentication. For gateway service-domain configuration, refer to [Manage Gateways](/sandboxes/manage-gateways#configure-service-forwarding).
</Note>
## Monitor and Debug
List all sandboxes:
```shell
openshell sandbox list
```
Filter the list by labels when you want a narrower view:
```shell
openshell sandbox list --selector team=platform
```
Use `-o json` or `-o yaml` for machine-readable output:
```shell
openshell sandbox list -o json
openshell sandbox list -o yaml
```
Get detailed information about a specific sandbox. The output lists **Policy source** (`sandbox` or `global`), **Revision** (the active policy’s row version for that source), and the formatted active policy YAML:
```shell
openshell sandbox get my-sandbox
```
For automation, use `--output json` or `--output yaml` to get machine-readable sandbox details:
```shell
openshell sandbox get my-sandbox --output json
```
Print only that policy YAML for scripting (same effective policy, no metadata):
```shell
openshell sandbox get my-sandbox --policy-only
```
Stream sandbox logs to monitor agent activity and diagnose policy decisions:
```shell
openshell logs my-sandbox
```
| Flag | Purpose | Example |
| ---------- | ---------------------------- | ---------------------------------- |
| `--tail` | Stream logs in real time | `openshell logs my-sandbox --tail` |
| `--source` | Filter by log source | `--source sandbox` |
| `--level` | Filter by severity | `--level warn` |
| `--since` | Show logs from a time window | `--since 5m` |
OpenShell Terminal combines sandbox status and live logs in a single real-time dashboard:
```shell
openshell term
```
Use the terminal to spot blocked connections marked `action=deny` and inference-related proxy activity. If a connection is blocked unexpectedly, add the host to your network policy. Refer to [Policies](/sandboxes/policies) for the workflow.
The dashboard has three panels stacked vertically: Gateways, Providers (or Global Settings), and Sandboxes. Navigate within a panel with `Up`/`Down` or `j`/`k`. At a list boundary the cursor overflows into the adjacent panel, skipping empty panels. Use `Tab`/`Shift+Tab` to cycle panels directly. Press `h`/`l` or `Left`/`Right` in the middle panel to switch between the Providers and Global Settings tabs.
## Port Forwarding
Forward a local port to a running sandbox to access services inside it, such as a web server or database:
```shell
openshell forward start 8000 my-sandbox
openshell forward start 8000 my-sandbox -d # run in background
```
OpenShell prints the local URL only after the forward listener is reachable. Background forwards must be tracked locally so `openshell forward list` and `openshell forward stop` can manage them.
List and stop active forwards:
```shell
openshell forward list
openshell forward stop 8000 my-sandbox
```
<Tip>
You can also forward a port at creation time with `--forward`:
```shell
openshell sandbox create --forward 8000 -- claude
```
</Tip>
## SSH Config
Generate an SSH config entry for a sandbox so tools like VS Code Remote-SSH can connect directly:
```shell
openshell sandbox ssh-config my-sandbox
```
Append the output to `~/.ssh/config` or use `--editor` on `sandbox create`/`sandbox connect` for automatic setup.
## Transfer Files
Upload files from your host into the sandbox:
```shell
openshell sandbox upload my-sandbox ./src
```
When you omit the destination, OpenShell discovers the sandbox's working
directory and uploads there. For a named local directory, OpenShell preserves
the basename, matching `scp -r` and `cp -r`. If that directory already exists,
the upload merges into it and overwrites matching entries without deleting
unrelated entries.
OpenShell preserves symlinks during upload. A symlink arrives in the sandbox as a symlink with the same target path instead of an expanded copy of the target file or directory. Dangling symlinks are also preserved.
Download files from the sandbox to your host:
```shell
openshell sandbox download my-sandbox output ./local
```
When the sandbox-side source is a single file, the destination follows `cp`-style placement: if the destination already exists as a directory or ends with `/`, the file lands inside it as `<dest>/<basename>`; otherwise the file is written at the exact destination path.
The CLI discovers the sandbox's canonical working directory and only allows
sandbox-side sources that resolve inside it. Paths that escape lexically, such
as `/etc/passwd` or `/sandbox/../etc/passwd`, and paths that escape through a
symlink are refused before any data is transferred. Relative sources are
resolved from the working directory; absolute sources within the same canonical
directory are also accepted.
<Note>
You can also upload files at creation time with the `--upload` flag on
`openshell sandbox create`. Pass `--upload` multiple times to upload
several paths in a single command:
```shell
openshell sandbox create --upload ./src:/workspace/src --upload ./config:/workspace/config -- claude
```
</Note>
By default, uploads inside a Git repository respect `.gitignore` rules so that
build artifacts, dependency caches, and other ignored files are not transferred.
If `.gitignore` filtering excludes every file in the upload path, the CLI falls
back to an unfiltered upload and prints a warning. Pass `--no-git-ignore` to
opt into unfiltered uploads explicitly, upload a path outside the Git work
tree, or force-add the intended files if they should remain Git-aware.
## Stop and Start Sandboxes
Stop compute when you want to retain a sandbox and its persistent workspace
without keeping its container, pod, or VM running:
```shell
openshell sandbox stop my-sandbox
openshell sandbox start my-sandbox
```
The name is optional and defaults to the last-used sandbox. Stop stops local
background forwards and waits for the `Stopped` phase. Start waits until the
same sandbox returns to `Ready`. While stopped, you cannot connect, execute
commands, transfer files, forward ports, or reach exposed services. Policies,
provider attachments, settings, service definitions, and persistent workspace
data remain associated with the sandbox.
Stop and start are idempotent. Delete a stopped sandbox normally when you
no longer need its retained state.
## Delete Sandboxes
Deleting a sandbox stops all processes, releases resources, and purges injected credentials.
```shell
openshell sandbox delete my-sandbox
```
## Sandbox Lifecycle
Every sandbox moves through a defined set of phases:
| Phase | Description |
| ------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Provisioning | The runtime is setting up the sandbox environment, or the gateway is waiting for the sandbox supervisor to establish its authenticated control session. |
| Ready | The sandbox is running and its supervisor control session is connected. You can connect, execute commands, sync files, and view logs. |
| Stopping | The gateway accepted a stop request and is stopping compute while retaining persistent state. |
| Stopped | Compute is stopped and access is unavailable. The sandbox record and driver-owned persistent workspace remain. |
| Starting | Compute is starting. The sandbox becomes usable only after a fresh supervisor session connects. |
| Error | Provisioning failed or the canonical main process exited unexpectedly. Main-process exit is terminal even with exit code 0. Check logs with `openshell logs`. |
| Deleting | The sandbox is being torn down. The system releases resources and purges credentials. |
The compute backend can become ready before the sandbox supervisor connects to
the gateway. During this interval, the sandbox remains in `Provisioning` and
reports a `Ready=False` condition with the reason `SupervisorNotConnected`.
After a gateway restart, an existing sandbox can return to `Provisioning`
temporarily while its supervisor reconnects. Wait for the phase to return to
`Ready` before you connect to the sandbox or execute commands.
The gateway records a canonical main-process exit as `Ready=False` with reason
`MainProcessExited`. It also sets `status.exit_code`; signal exits use the
standard `128 + signal` convention. Compute runtimes do not automatically
restart that process.
## Sandbox Compute Drivers
The gateway's configured compute driver determines how OpenShell creates each sandbox. The CLI workflow stays the same across drivers: you create, connect to, inspect, and delete sandboxes through the gateway API.
For Docker, Podman, MicroVM, and Kubernetes behavior, refer to [Sandbox Compute Drivers](/reference/sandbox-compute-drivers).
## Next Steps
- To follow a complete end-to-end example, refer to the [GitHub Sandbox](/get-started/tutorials/github-sandbox) tutorial.
- To select a workspace or understand access roles, refer to [Manage Workspaces and Access](/sandboxes/manage-workspaces).
- To supply API keys or tokens, refer to [Manage Providers](/sandboxes/manage-providers).
- To control what the agent can access, refer to [Policies](/sandboxes/policies).
- To use the default runtime image, refer to [Base Sandbox Container](#base-sandbox-container).
@@ -0,0 +1,196 @@
---
# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
title: "Manage Workspaces and Access"
sidebar-title: "Workspaces and Access"
description: "Create OpenShell workspaces, assign members, and understand platform and workspace roles."
keywords: "Generative AI, Cybersecurity, Workspaces, Access Control, RBAC, OIDC, Membership, CLI"
position: 3
---
An OpenShell workspace is an access and resource isolation boundary. Sandboxes,
providers, services, policies, settings, and inference routes belong to a
workspace and are not visible to members of other workspaces.
The CLI targets the `default` workspace unless you set `--workspace` or
`OPENSHELL_WORKSPACE`. The logical OpenShell workspace described here is
separate from the `/sandbox` filesystem directory inside a sandbox.
## Understand the Role Model
OpenShell combines an identity-provider role with a membership record for each
workspace.
| Role | Assignment | Access |
|---|---|---|
| Platform Admin | The OIDC role configured as `admin_role`. | Manages platform-scoped configuration and every workspace. Platform Admins bypass workspace membership checks. |
| Workspace Admin | An `admin` membership stored by the gateway for one workspace. | Manages providers, provider profiles, policies, settings, and members in that workspace. |
| Workspace User | A `user` membership stored by the gateway for one workspace. | Creates and uses sandboxes and services, reads providers, and uses provider attachments in that workspace. |
OIDC users also need the role configured as `user_role` for ordinary workspace
operations. The configured Platform Admin role satisfies this requirement.
Membership does not grant access to another workspace, and a Workspace Admin
cannot perform platform-scoped or cross-workspace operations.
When the gateway enables scope enforcement through `scopes_claim`, the token
must also contain the scope required by the operation. Common scopes include
`workspace:read`, `workspace:write`, `sandbox:read`, `sandbox:write`,
`provider:read`, `provider:write`, `config:read`, and `config:write`.
`openshell:all` satisfies every scope requirement. For OIDC and scope
configuration, refer to [Gateway Authentication](/reference/gateway-auth).
The following table summarizes common operations.
| Operation | Platform Admin | Workspace Admin | Workspace User |
|---|---|---|---|
| Create or delete a workspace | Any workspace | No | No |
| View a workspace or list workspaces | All workspaces | Assigned workspaces | Assigned workspaces |
| List workspace members | Any workspace | Assigned workspace | Assigned workspace |
| Add Workspace Users or remove members | Any workspace | Assigned workspace | No |
| Assign the Workspace Admin role | Any workspace | No | No |
| Create, use, or delete sandboxes and services | Any workspace | Assigned workspace | Assigned workspace |
| Create, update, or delete providers | Any workspace | Assigned workspace | No |
| Change workspace policy or settings | Any workspace | Assigned workspace | No |
| Manage platform profiles or global configuration | Yes | No | No |
| List resources across workspaces | Yes | No | No |
<Note>
Local gateways without OIDC role configuration treat authenticated users as
Platform Admins. Configure OIDC roles and workspace membership for shared
gateways.
</Note>
## Inspect Your Identity
Use the identity validated by the gateway when an administrator needs your
membership subject.
```shell
openshell whoami
openshell whoami --output json
```
The `subject` field is the stable identity used in workspace membership
records. Each user can run this command even when they do not belong to a
workspace.
## Create a Workspace and Add Members
A Platform Admin creates workspaces and assigns the first Workspace Admin.
The gateway creates the `default` workspace automatically, but it does not add
OIDC users to that workspace automatically.
Create a workspace:
```shell
openshell workspace create --name team-ml
```
Ask the intended Workspace Admin to run `openshell whoami`, then add the
reported subject:
```shell
openshell workspace member add \
--workspace team-ml \
--subject 'oidc-subject-for-admin' \
--role admin
```
The Workspace Admin can add Workspace Users:
```shell
openshell workspace member add \
--workspace team-ml \
--subject 'oidc-subject-for-user' \
--role user
```
Only a Platform Admin can assign the `admin` membership role. A Workspace
Admin can add `user` members and remove members in their assigned workspace.
## List and Remove Members
All members can inspect membership in their workspace. Workspace Admins and
Platform Admins can remove members.
```shell
openshell workspace member list --workspace team-ml
openshell workspace member remove \
--workspace team-ml \
--subject 'oidc-subject-for-user'
```
To change a member's role, remove the existing membership and add it again
with the new role. A Platform Admin must perform any change to `admin`.
## Target a Workspace
Pass `--workspace` to scope a resource operation. The flag is global, so it
can appear before or after the subcommand.
```shell
openshell sandbox list --workspace team-ml
openshell provider list --workspace team-ml
openshell sandbox create --workspace team-ml --name research -- bash
```
Set a default for the current shell with `OPENSHELL_WORKSPACE`:
```shell
export OPENSHELL_WORKSPACE=team-ml
openshell sandbox list
```
An empty workspace value resolves to `default`. It never means all
workspaces.
Platform Admins can opt into cross-workspace list operations:
```shell
openshell sandbox list --all-workspaces
openshell provider list --all-workspaces
openshell service list --all-workspaces
```
Provider profiles and policy also have explicit `--global` operations. Those
operations target platform scope and require Platform Admin access. A
Workspace Admin should use `--workspace` for workspace-scoped profiles and
configuration.
## Diagnose Access Denials
If `openshell workspace list` returns no rows, the authenticated subject has no
workspace memberships. Run `openshell whoami` and send the `subject` value to
a Platform Admin.
Workspace authorization errors include a copyable membership command. A
non-member denial suggests `--role user`. If an operation requires Workspace
Admin access, the denial suggests `--role admin`; only a Platform Admin can run
that assignment successfully.
If the membership is correct but the request is still denied, inspect `roles`
and `scopes` with `openshell whoami --output json`. Confirm that the token has
the configured OIDC user role and, when scope enforcement is enabled, the
scope required by the operation.
## Delete a Workspace
Only a Platform Admin can delete a workspace. The `default` workspace cannot
be deleted.
```shell
openshell workspace delete team-ml
```
A custom workspace must not contain sandboxes, providers, provider profiles,
services, SSH sessions, settings, policies, draft policy chunks, or credential
refresh state. Remove those resources before retrying deletion. OpenShell
removes membership records and inference routes as part of successful
workspace deletion.
## Next Steps
- To configure OIDC roles and scopes, refer to [Gateway Authentication](/reference/gateway-auth).
- To create resources in a workspace, refer to [Manage Sandboxes](/sandboxes/manage-sandboxes).
- To manage workspace credentials, refer to [Providers](/sandboxes/manage-providers).
+865
View File
@@ -0,0 +1,865 @@
---
# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
title: "Customize Sandbox Policies"
sidebar-title: "Policies"
description: "Apply, iterate, and debug sandbox network policies with hot-reload on running OpenShell sandboxes."
keywords: "Generative AI, Cybersecurity, Policy, Network Policy, Sandbox, Security, Hot Reload"
position: 6
---
Use this page to apply and iterate policy changes on running sandboxes. For a full field-by-field YAML definition, use the [Policy Schema Reference](/reference/policy-schema).
## Policy Structure
A policy has static sections `filesystem_policy`, `landlock`, and `process` that are locked at sandbox creation, and dynamic `network_policies` and `network_middlewares` sections that are hot-reloadable on a running sandbox.
```yaml wordWrap showLineNumbers={false}
version: 1
# Static: locked at sandbox creation. Paths the agent can read vs read/write.
filesystem_policy:
include_workdir: true
read_only: [/usr, /lib, /etc]
read_write: [/tmp]
# Static: Landlock LSM kernel enforcement. best_effort uses highest ABI the host supports.
landlock:
compatibility: best_effort
# Static, optional: override the identity selected by the compute driver.
# process:
# run_as_user: "1500"
# run_as_group: "1500"
# Dynamic: hot-reloadable. Named blocks of endpoints + binaries allowed to reach them.
network_policies:
my_api:
name: my-api
endpoints:
- host: api.example.com
port: 443
protocol: rest
enforcement: enforce
access: full
binaries:
- path: /usr/bin/curl
# Dynamic: ordered middleware selected independently by admitted host.
network_middlewares:
regex-redactor:
name: Redact API tokens
middleware: openshell/regex
order: 10
config:
mode: redact
on_error: fail_closed
endpoints:
include: ["api.example.com"]
exclude: []
```
Static sections are locked at sandbox creation. Changing them requires destroying and recreating the sandbox.
Dynamic sections can be updated on a running sandbox with `openshell policy update` for incremental merges or `openshell policy set` for full replacement, and take effect without restarting.
When a hot reload changes rules, the supervisor publishes a new policy generation and closes connections pinned to the previous generation. This includes HTTP keep-alive tunnels, `tls: skip`, non-HTTP payloads, HTTP upgrades such as WebSocket, and long-lived response streams such as SSE. Most clients reconnect automatically, and the next connection or request is evaluated against the current policy. A parsed WebSocket relay closes with code `1012` when its attached policy generation becomes stale. Use `protocol: websocket` when policy should stay attached to the RFC 6455 upgrade and client text messages after the allowed upgrade. Provider-credentialed endpoints reject L4-only and `tls: skip` modes by default. On a credentialed WebSocket upgrade, OpenShell keeps the connection on the parsed relay and rejects binary frames unless the endpoint explicitly sets `allow_uninspected_credentials: true`. Add `websocket_credential_rewrite: true` when the relay should rewrite credential placeholders in client-to-server WebSocket text messages. Add `request_body_credential_rewrite: true` only on inspected REST endpoints that need OpenShell to rewrite placeholders in supported text request bodies.
| Section | Type | Description |
|---|---|---|
| `filesystem_policy` | Static | Controls which directories the agent can access on disk. Paths are split into `read_only` and `read_write` lists. Any path not listed in either list is inaccessible. Set `include_workdir: true` to automatically add the agent's working directory to `read_write`. [Landlock LSM](https://docs.kernel.org/security/landlock.html) enforces these restrictions at the kernel level. |
| `landlock` | Static | Configures Landlock LSM enforcement behavior. Set `compatibility` to `best_effort` (skip individual inaccessible paths while applying remaining rules) or `hard_requirement` (fail if any path is inaccessible or the required kernel ABI is unavailable). Refer to the [Policy Schema Reference](/reference/policy-schema#landlock) for the full behavior table. |
| `process` | Static | Optionally overrides the OS-level identity for the agent process. Explicit values must be `sandbox` or numeric UID/GID values from `1` through `4294967294`; root and the invalid identity sentinel are rejected. Docker and Podman may use named identities through per-field OCI `USER` fallback; Kubernetes uses its platform-selected numeric identity. The agent also runs with seccomp filters that block dangerous system calls. |
| `network_policies` | Dynamic | Controls network access for ordinary outbound traffic from the sandbox. Each block has a name, a list of endpoints (host, port, protocol, and optional rules), and a list of binaries allowed to use those endpoints. <br />Every outbound connection except `https://inference.local` passes through the network supervisor, which queries the [policy engine](/about/how-it-works#core-components) with the destination and calling binary. A connection is allowed only when both match an entry in the same policy block. <br />For endpoints with `protocol: rest`, the proxy auto-detects TLS and terminates it so each HTTP request can be checked against that endpoint's `rules` (method and path). For endpoints with `protocol: websocket`, the proxy validates the RFC 6455 upgrade and evaluates `GET` rules for the handshake plus either `WEBSOCKET_TEXT` rules for raw client text messages or GraphQL operation rules for GraphQL-over-WebSocket messages. Set `websocket_credential_rewrite: true` only when a WebSocket or REST compatibility endpoint must keep placeholder credentials in sandbox-owned text frames and resolve them at the OpenShell relay boundary. <br />Endpoints with `protocol: tcp` allow ordinary DNS resolution and native TCP connections without inspecting payloads. Endpoints without `protocol` retain L4 passthrough through an explicit proxy. <br />If no endpoint matches, the connection is denied. Configure managed inference separately through [Inference Routing](/sandboxes/inference-routing). |
| `network_middlewares` | Dynamic | Declares keyed HTTP and WebSocket middleware configs. After network and L7 policy admit a request or upgrade, OpenShell matches each config's host selectors independently and runs matching entries by their unique ascending `order` before credential injection. WebSocket-capable entries continue on complete client text messages. |
## Supervisor Middleware
Supervisor middleware can inspect, deny, or replace admitted HTTP request bodies and client WebSocket text messages before provider credentials are injected. Middleware selection is independent of the `network_policies` rule that admitted the traffic: each keyed `network_middlewares` entry matches the destination host through `endpoints.include` and `endpoints.exclude`.
```yaml
network_middlewares:
regex-redactor:
name: Redact API tokens
middleware: openshell/regex
order: 10
config:
mode: redact
on_error: fail_closed
endpoints:
include: ["*.example.com"]
exclude: ["trusted.example.com"]
```
Matching entries run once each by ascending `order`; lower values run first, and duplicate order values are rejected. The default order is `0`, so policies with multiple entries normally set it explicitly. Map keys are structurally unique. An optional `name` provides a human-readable label, defaults to the map key, and does not change the key used as the config identity. Different keys may use the same implementation and run as distinct stages. `exclude` takes precedence over `include`.
`openshell/regex` is an example built into the supervisor. It applies fixed regular expressions to UTF-8 HTTP request bodies and complete client-to-upstream WebSocket text messages before credential injection. The initial pattern recognizes `sk-` tokens. This is a best-effort text transformation without guarantees that sensitive values will be detected or fully removed; it does not inspect binary or upstream-to-client WebSocket messages. Custom expressions are not configurable yet. Operator-run middleware must be registered by name before a policy can reference it. The gateway validates implementation-owned config before accepting the policy.
`on_error` defaults to `fail_closed`. Use `fail_open` only when skipping a selected stage that fails is acceptable. For WebSocket streams, a broken fail-open stage is disabled for the rest of that connection and OpenShell emits a state-change finding. A host-matched attachment joins only operation chains advertised by its implementation. An HTTP-only attachment may inspect the WebSocket upgrade GET, but post-upgrade messages pass with informational `binding_not_selected` coverage under either error mode. Binary messages also pass without middleware inspection and produce `unsupported_message_type` coverage for active WebSocket stages. Upstream-to-client messages remain uninspected. Policy validation rejects a fail-closed selector that can cover a `tls: skip` endpoint. An all-`fail_open` match may cover the endpoint; the supervisor bypasses the middleware and emits a detection finding.
See [Supervisor Middleware](/extensibility/supervisor-middleware) for registration, chain ordering, body limits, failure behavior, and operations.
## Baseline Filesystem Paths
When a sandbox runs in proxy mode (the default), OpenShell automatically adds baseline filesystem paths required for the sandbox child process to function: `/usr`, `/lib`, `/etc`, and `/var/log` (read-only), plus `/tmp` (read-write). When `filesystem.include_workdir` is `true`, OpenShell also adds the resolved working directory as read-write. Paths like `/app` are included in the baseline set but are only added if they exist in the container image.
For GPU sandboxes, OpenShell also adds existing GPU device nodes as read-write paths. CUDA workloads require write access to procfs for thread metadata, so GPU baseline enrichment moves `/proc` from read-only to read-write when GPU devices are present.
This filtering prevents a missing baseline path from degrading Landlock enforcement. Without it, a single missing path could cause the entire Landlock ruleset to fail, leaving the sandbox with no filesystem restrictions at all.
User-specified paths in your policy YAML are not pre-filtered. If you list a path that does not exist:
- In `best_effort` mode, the path is skipped with a warning and remaining rules are still applied.
- In `hard_requirement` mode, sandbox startup fails immediately.
This distinction means baseline system paths degrade gracefully while user-specified paths surface configuration errors.
## Allow Native TCP Connections
Use `protocol: tcp` when an application needs to resolve a policy-approved hostname and open a normal TCP connection without configuring an HTTP proxy. The endpoint remains L4-only, so OpenShell authorizes the hostname, port, and calling binary but does not inspect application payloads.
```yaml showLineNumbers={false}
network_policies:
postgres:
name: postgres
endpoints:
- host: db.internal.example
port: 5432
protocol: tcp
binaries:
- path: /usr/bin/psql
```
OpenShell answers DNS only for hostnames eligible under an active `protocol: tcp` endpoint. It validates upstream answers against destination and SSRF controls, returns a supervisor-owned synthetic address, and records the validated real addresses. When the application connects to the synthetic address, OpenShell recovers the hostname and port, evaluates the calling process against the current policy generation, and dials only an address pinned by that DNS result.
Applications must honor the returned DNS TTL and resolve the hostname again before reconnecting after that TTL expires. A client that caches the synthetic address indefinitely can receive a connection failure after the mapping expires. Docker and Podman currently advertise only IPv4 egress for this feature, so OpenShell returns an empty successful answer for AAAA queries and lets dual-stack clients use the working A record.
DNS resolution does not authorize a connection by itself. Unknown names, wrong ports, stale mappings, disallowed destination addresses, and binaries outside the matching policy fail closed. Applications cannot inherit access by connecting directly to a real IP returned by an upstream resolver.
Do not combine `protocol: tcp` with L7-only fields such as `path`, `enforcement`, `access`, `rules`, `deny_rules`, request rewriting, or credential signing. Docker and Podman sandboxes support policy DNS and transparent TCP capture. Other compute drivers reject policies containing `protocol: tcp` until they provide the required runtime capability.
Prefer exact hostnames for TCP endpoints. A wildcard authorizes DNS queries for every matching name, which means a compromised process can encode data into matching DNS labels even when the lookup does not return an address. OpenShell records failed eligible lookups for operators, but logging does not remove that exfiltration channel.
The first TCP endpoint is a deliberate exception to ordinary dynamic network policy updates. A sandbox created without any `protocol: tcp` endpoint does not install the DNS and transparent-capture substrate. A hot reload that introduces the first TCP endpoint is rejected atomically and leaves the previous policy active; recreate the sandbox with a TCP endpoint to install the substrate. A sandbox that started with TCP support can remove and later re-add TCP endpoints through normal policy reloads.
## Apply a Custom Policy
Pass a policy YAML file when creating the sandbox:
```shell
openshell sandbox create --policy ./my-policy.yaml -- claude
```
The trailing command is the sandbox's canonical main process. If it exits, the
sandbox enters `Error`; use `sandbox exec` for one-shot commands that should not
define sandbox health.
To avoid passing `--policy` every time, set a default policy with an environment variable:
```shell
export OPENSHELL_SANDBOX_POLICY=./my-policy.yaml
openshell sandbox create -- claude
```
The CLI uses the policy from `OPENSHELL_SANDBOX_POLICY` whenever `--policy` is not explicitly provided.
## Iterate on a Running Sandbox
To change what the sandbox can access, pull the current policy, edit the YAML, and push the update. The workflow is iterative: create the sandbox, monitor logs for denied actions, pull the policy, modify it, push, and verify.
```mermaid
flowchart TD
A["1. Create sandbox with initial policy"] --> B["2. Monitor logs for denied actions"]
B --> C["3. Pull current policy"]
C --> D["4. Modify the policy YAML"]
D --> E["5. Push updated policy"]
E --> F["6. Verify the new revision loaded"]
F --> B
style A fill:#76b900,stroke:#000000,color:#000000
style B fill:#76b900,stroke:#000000,color:#000000
style C fill:#76b900,stroke:#000000,color:#000000
style D fill:#ffffff,stroke:#000000,color:#000000
style E fill:#76b900,stroke:#000000,color:#000000
style F fill:#76b900,stroke:#000000,color:#000000
linkStyle default stroke:#76b900,stroke-width:2px
```
The following steps outline the hot-reload policy update workflow.
1. Create the sandbox with your initial policy by following [Apply a Custom Policy](#apply-a-custom-policy) above (or set `OPENSHELL_SANDBOX_POLICY`).
2. Monitor denials. Each log entry shows host, port, binary, and reason. Alternatively, use `openshell term` for a live dashboard.
```shell
openshell logs <name> --tail --source sandbox
```
3. For additive network changes, use `openshell policy update`. This is the fastest path for adding endpoints, binaries, or REST and WebSocket allow/deny rules without replacing the full policy. The full option and format reference is in [Incremental Policy Updates](#incremental-policy-updates).
```shell
openshell policy update <name> \
--add-endpoint api.github.com:443:read-only:rest:enforce \
--binary /usr/bin/gh \
--wait
openshell policy update <name> \
--add-allow 'api.github.com:443:POST:/repos/*/issues' \
--wait
```
`--add-allow` and `--add-deny` target existing `protocol: rest` or `protocol: websocket` endpoints. If you pass multiple update flags in one command, OpenShell applies them as one atomic merge batch and persists at most one new revision.
4. For larger edits, pull the current base policy and edit the YAML directly. The base policy is the user-authored policy without provider-composed `_provider_*` entries, so it is safe to round-trip through `openshell policy set`. Before reusing the file, strip the metadata header above the `---` line.
```shell
openshell policy get <name> --base > current-policy.yaml
```
To inspect the effective policy that the sandbox enforces, including provider-composed entries, use `openshell policy get <name> --full`. To inspect a stored sandbox-authored revision instead of the current effective policy, pass `--rev <version>`.
5. Edit the YAML: add or adjust `network_policies` entries, binaries, `access`, `rules`, or protocol-specific matchers such as GraphQL operation fields, MCP `method` / `tool` rules, and generic JSON-RPC `method` rules.
6. Push the updated policy when you need a full replacement. Exit codes: 0 = loaded, 1 = validation failed, 124 = timeout.
```shell
openshell policy set <name> --policy current-policy.yaml --wait
```
7. Verify the new revision. If status is `loaded`, repeat from step 2 as needed; if `failed`, fix the policy and repeat from step 4.
```shell
openshell policy list <name>
```
### Validation failures
OpenShell validates a complete candidate policy before activating any part of it. Endpoints may overlap when their connection and request-processing metadata agree. For example, two `api.example.com:443` REST entries can contribute different allow and deny rules when they use the same TLS, destination, credential, parser, and enforcement settings. A plain L4 endpoint may overlap an L7 endpoint because it authorizes the destination without contributing request-processing metadata. A more-specific path endpoint may override request-processing metadata from a broader endpoint, such as a `/graphql` GraphQL endpoint alongside a general REST endpoint for the same host. OpenShell rejects the candidate when overlapping exact or wildcard host selectors can both contribute equally specific endpoint configuration and disagree on those fields.
When the gateway knows the affected sandbox scope, it validates the complete
effective candidate before persistence. This covers direct policy replacement,
incremental merges and proposal approvals, provider attachment, and
provider-profile updates that fan out to attached sandboxes. An ambiguity
failure returns `FAILED_PRECONDITION`; OpenShell stores no invalid policy
revision and does not partially apply a profile update. Supervisor validation
remains a defense-in-depth boundary for startup, concurrent changes, and policy
sources outside those mutation paths.
A gateway preflight rejection leaves the currently active policy unchanged
regardless of failure mode because the candidate is never persisted or
distributed. If a candidate reaches a supervisor and fails runtime validation,
the gateway's `policy_validation_failure_mode` configuration determines the
supervisor posture. Set it under `[openshell.gateway]` in `gateway.toml`. Its
default is `fail_closed`:
```toml
[openshell.gateway]
policy_validation_failure_mode = "fail_closed"
```
In `fail_closed` mode, the supervisor publishes a quarantine generation, denies new egress, and closes connections pinned to the previous generation. The previous policy is not active. A later valid policy exits quarantine automatically.
Operators that explicitly prioritize availability can retain the previous generation:
```toml
[openshell.gateway]
policy_validation_failure_mode = "retain_last_valid"
```
In `retain_last_valid` mode, the rejected candidate remains inactive and the previous valid generation remains active. If no previous valid generation exists, such as during initial startup, OpenShell still fails closed. Restart the gateway after changing `gateway.toml`; connected sandbox supervisors receive the configured posture from the restarted gateway. Individual sandboxes cannot override it.
OCSF configuration and finding events identify the rejected candidate, validation rationale, configured and effective modes, active generation, and whether the previous policy is active. When `retain_last_valid` is configured without a previous valid generation, the effective mode remains `fail_closed`. Connection denials during quarantine include the validation failure as their policy denial rationale.
## Incremental Policy Updates
Use `openshell policy update` when you want to merge network policy changes into the current live policy instead of replacing the whole YAML document. This command only updates the dynamic `network_policies` section.
`openshell policy update` is useful when you want to:
- add a new endpoint for an existing binary without touching other policy sections.
- add a few REST or WebSocket allow/deny rules after you see a blocked request in the logs.
- remove one endpoint or one named rule without rewriting the rest of the file.
- preview a merged result locally with `--dry-run` before you send it to the gateway.
Use `openshell policy set` instead when you want to replace the full policy, update static sections, or make broader edits that are easier to express in YAML. Use full YAML for GraphQL, MCP, and JSON-RPC rule shapes.
### Update Commands
The incremental update surface is split into endpoint-level operations and method/path rule-level operations for REST and WebSocket endpoints.
| Flag | What it changes | Typical use |
|---|---|---|
| `--add-endpoint <SPEC>` | Creates or merges a network rule and endpoint. | Allow a new host and port, optionally with `access`, `protocol`, `enforcement`, endpoint options, and binaries. |
| `--remove-endpoint <SPEC>` | Removes one host and port match from the current policy. | Drop a stale endpoint or remove one port from a multi-port endpoint. |
| `--remove-rule <NAME>` | Deletes a named `network_policies` entry. | Remove a whole rule by name when you no longer need it. |
| `--add-allow <SPEC>` | Appends method/path allow rules to an existing REST or WebSocket endpoint. | Permit one additional REST method/path or WebSocket `WEBSOCKET_TEXT` path on an API that is already configured. |
| `--add-deny <SPEC>` | Appends method/path deny rules to an existing REST or WebSocket endpoint. | Block a sensitive REST path or WebSocket text-message path under an endpoint that is otherwise allowed. |
| `--binary <PATH>` | Adds binaries to every `--add-endpoint` rule in the same command. | Bind a new endpoint to one or more executables. |
| `--rule-name <NAME>` | Overrides the generated rule name. | Keep a stable human-chosen rule name when adding exactly one endpoint. |
| `--dry-run` | Shows the merged policy locally and does not call the gateway. | Review the result before persisting it. |
| `--wait` | Polls until the sandbox reports that the new revision loaded. | Confirm the change took effect before continuing. |
| `--timeout <SECS>` | Sets the timeout for `--wait`. | Extend the wait window for slower sandboxes. |
`--wait` and `--dry-run` cannot be used together.
### Add Endpoint Compared to Allow and Deny
`--add-endpoint` works at the endpoint and rule level. It creates a new `network_policies` entry when needed, or merges into an existing rule that already covers the same host and port. Use it when you define where traffic can go and which binaries can send it.
`--add-allow` and `--add-deny` work at the method/path rule level. They do not create binaries, and they do not create a new endpoint. They modify an existing endpoint that already has `protocol: rest` or `protocol: websocket`.
This is the practical difference:
- Use `--add-endpoint` to say "allow this binary to reach `api.github.com:443`."
- Use `--add-allow` to say "for that existing REST endpoint, also allow `POST /repos/*/issues`."
- Use `--add-deny` to say "for that existing REST endpoint, explicitly deny `POST /admin/**`."
- Use `--add-allow` to say "for that existing WebSocket endpoint, also allow client text messages on `/v1/realtime/**`."
Current constraints:
- `--add-allow` and `--add-deny` work on `protocol: rest` and `protocol: websocket` endpoints.
- GraphQL, MCP, and JSON-RPC fine-grained rules require full policy YAML applied with `openshell policy set`.
- `--add-deny` requires the endpoint to already have an allow base, either an `access` preset or explicit allow `rules`.
- `protocol: sql` is not a practical incremental workflow today. OpenShell does not do full SQL parsing, and SQL enforcement is not meaningfully supported yet.
### Endpoint Specs
`--add-endpoint` uses this format:
```text
host:port[:access[:protocol[:enforcement[:options]]]]
```
Each segment has a fixed meaning:
| Segment | Required | Meaning |
|---|---|---|
| `host` | Yes | Destination hostname. |
| `port` | Yes | Destination port, `1` through `65535`. |
| `access` | No | Access preset for L7 endpoints: `read-only`, `read-write`, or `full`. Incremental updates expand presets into protocol-specific method/path rules for REST and WebSocket endpoints. |
| `protocol` | No | Endpoint mode accepted by `openshell policy update`: `tcp`, `rest`, `websocket`, or `sql`. Use `tcp` for native DNS and TCP without L7 inspection. `sql` is audit-only and not a recommended workflow today. Full policy YAML also supports `graphql`, `mcp`, and `json-rpc`. |
| `enforcement` | No | Enforcement mode for inspected traffic: `enforce` or `audit`. |
| `options` | No | Comma-separated endpoint options. Use `websocket-credential-rewrite` with `protocol: websocket` or REST compatibility endpoints that perform a WebSocket upgrade. Use `request-body-credential-rewrite` only with `protocol: rest`. |
Examples:
| Example | Meaning |
|---|---|
| `pypi.org:443` | Add a plain L4 endpoint. The proxy allows the TCP stream and does not inspect HTTP requests. |
| `telemetry.example.com:443::::allow-uninspected-credentials` | Explicitly allow a provider-credentialed L4 endpoint after accepting that OpenShell cannot inspect or rewrite its traffic. |
| `db.internal.example:5432::tcp` | Add an L4 endpoint for native DNS resolution and transparent TCP capture. The empty `access` segment is required before `tcp`. |
| `api.github.com:443:read-only:rest:enforce` | Add a REST endpoint with the `read-only` preset expanded by the policy engine into GET, HEAD, and OPTIONS access. |
| `api.example.com:443:read-write:rest:enforce:request-body-credential-rewrite` | Add a REST endpoint that rewrites credential placeholders in supported text request bodies. |
| `realtime.example.com:443:read-write:websocket:enforce` | Add a WebSocket endpoint with the `read-write` preset expanded by the policy engine into the upgrade `GET` and client `WEBSOCKET_TEXT` access. |
| `realtime.example.com:443:read-write:websocket:enforce:websocket-credential-rewrite` | Add a WebSocket endpoint that rewrites `openshell:resolve:env:*` placeholders in client text frames after an allowed upgrade. |
If you set `protocol: rest` or `protocol: websocket`, you also need an allow shape. With incremental updates, that means you should provide an `access` preset on `--add-endpoint`, then use `--add-allow` or `--add-deny` to refine method/path rules later.
Use the `websocket-credential-rewrite` endpoint option with `protocol: websocket` when the sandbox should send credential placeholders in client text frames and have OpenShell resolve them after the allowed upgrade. The option can also be used with `protocol: rest` compatibility endpoints that perform a WebSocket upgrade. It is rejected for plain L4 or `protocol: sql` endpoints.
Use the `request-body-credential-rewrite` endpoint option with `protocol: rest` when an API expects OpenShell-managed credentials in UTF-8 JSON, form, or text request bodies. OpenShell buffers up to 256 KiB, rewrites recognized credential placeholders, updates `Content-Length`, and rejects unresolved placeholders instead of forwarding them. For chunked requests, the 256 KiB limit counts the complete wire representation, including framing, extensions, and trailers. The option is rejected for WebSocket, GraphQL, SQL, and plain L4 endpoints.
Use `allow-uninspected-credentials` only when a provider-credentialed endpoint must remain L4-only, use `tls: skip`, or carry uninspectable WebSocket traffic. Without this explicit opt-in, the gateway rejects credentialed L4-only and `tls: skip` endpoints. REST bodies without placeholders continue to work when body rewrite is disabled; a body containing an OpenShell credential placeholder fails closed.
Credential rewrite recognizes the canonical `openshell:resolve:env:KEY` placeholder form and whole-token provider-shaped aliases such as `provider-OPENSHELL-RESOLVE-ENV-API_TOKEN` when the referenced environment key exists in the configured provider credentials.
Static provider placeholders resolve only when the request host, port, and path
also match an endpoint in the provider profile. A sandbox policy allow does not
expand that binding. A mismatch returns HTTP 403 with
`credential_endpoint_mismatch`. Refer to [Static Credential Endpoint
Binding](/sandboxes/providers-v2#understand-static-credential-endpoint-binding).
For example:
- `db.internal.example:5432::tcp` is valid.
- `api.github.com:443:read-only:rest` is valid.
- `realtime.example.com:443:read-write:websocket` is valid.
- `api.github.com:443::rest` is invalid. It does not mean "allow all traffic." An L7 endpoint with `protocol` but no `access` or `rules` is rejected when the policy loads.
Endpoint options belong to the individual `--add-endpoint` spec. When you pass multiple `--add-endpoint` flags in one command, every `--binary` value applies to every added endpoint in that command. If different endpoints need different binaries, use separate `policy update` commands.
If you do not pass `--rule-name`, OpenShell generates one from the host and port, such as `allow_api_github_com_443`.
### Method/Path Rule Specs
`--add-allow` and `--add-deny` use this format:
```text
host:port:METHOD:path_glob
```
This string identifies an existing REST or WebSocket endpoint and the request pattern you want to add.
In shell commands, quote the full `SPEC` when it contains `*` or `**` so your shell passes it literally instead of expanding it as a local file glob.
| Segment | Meaning |
|---|---|
| `host` | Existing endpoint host. |
| `port` | Existing endpoint port. |
| `METHOD` | HTTP method for REST endpoints, or `GET` / `WEBSOCKET_TEXT` for WebSocket endpoints. The CLI normalizes it to uppercase. |
| `path_glob` | URL path glob. For WebSocket text messages, this still matches the upgraded request path, not message payload content. It must start with `/`, or be `**`, or start with `**/`. |
This example:
```text
api.github.com:443:POST:/repos/*/issues
```
means:
- match the endpoint `api.github.com:443`.
- match HTTP method `POST`.
- match paths like `/repos/acme/issues`.
- also match deeper paths when the surrounding literals align, because `*` may include `/`.
Path globs follow the same semantics as YAML allow and deny rules:
- `*` and `**` match zero or more characters and may cross `/` boundaries.
- `?` matches exactly one character.
- bracket classes such as `[0-9]` and negated classes such as `[!0]` are supported.
- `/repos/*/issues` matches any intervening text, including multiple path segments.
- `/repos/**` matches everything under `/repos/`.
The rule-level commands only modify method and path constraints. They do not change binaries, hostnames, ports, protocol settings, or WebSocket message payload matching.
### Common Workflows
Use these patterns as starting points when you decide whether to update an endpoint or append REST/WebSocket rules.
#### Add a new L4 endpoint
Use `--add-endpoint` when you need a new host and port and do not need REST inspection.
```shell
openshell policy update demo \
--add-endpoint pypi.org:443 \
--add-endpoint files.pythonhosted.org:443 \
--binary /usr/bin/pip \
--binary /usr/local/bin/uv \
--wait
```
This creates or merges endpoint entries and binds them to the listed binaries. It does not create inspected method/path rules.
#### Create a REST endpoint with a base allow set
Use `--add-endpoint` first when the endpoint does not exist yet.
```shell
openshell policy update demo \
--add-endpoint api.github.com:443:read-only:rest:enforce \
--binary /usr/bin/gh \
--wait
```
This creates a REST endpoint and sets its base allow behavior through the `read-only` access preset.
#### Add one more REST allow rule
Use `--add-allow` after the REST endpoint already exists.
```shell
openshell policy update demo \
--add-allow 'api.github.com:443:POST:/repos/*/issues' \
--wait
```
This keeps the existing endpoint definition and appends one new allow rule. It does not add binaries or change the endpoint host and port.
#### Add a REST deny rule under an allowed endpoint
Use `--add-deny` when you want to carve out a blocked subtree under an existing REST endpoint.
```shell
openshell policy update demo \
--add-deny 'api.github.com:443:POST:/admin/**' \
--wait
```
This adds a deny rule to the existing REST endpoint. The endpoint must already have an allow base.
#### Create a WebSocket endpoint with a base allow set
Use `--add-endpoint` with `protocol: websocket` when the destination is an RFC 6455 WebSocket API.
```shell
openshell policy update demo \
--add-endpoint realtime.example.com:443:read-write:websocket:enforce:websocket-credential-rewrite \
--binary /usr/bin/node \
--wait
```
This creates a WebSocket endpoint and sets its base allow behavior through the `read-write` access preset. For WebSocket endpoints, `read-write` expands to the upgrade `GET` and client `WEBSOCKET_TEXT` messages on the upgraded request path. The rewrite option lets the sandbox send `openshell:resolve:env:*` placeholders in client text frames; OpenShell resolves them before forwarding to the upstream service.
#### Add a WebSocket text-message deny rule
Use `WEBSOCKET_TEXT` when you want to refine client-to-server text-frame policy without matching message payload content.
```shell
openshell policy update demo \
--add-deny 'realtime.example.com:443:WEBSOCKET_TEXT:/v1/admin/**' \
--wait
```
This adds a deny rule to the existing WebSocket endpoint. The path glob matches the WebSocket upgrade path.
#### Remove one endpoint or rule
Use `--remove-endpoint` to remove one host and port pair, or `--remove-rule` to delete the whole named rule.
```shell
openshell policy update demo --remove-endpoint pypi.org:443 --wait
openshell policy update demo --remove-rule github_repos --wait
```
If the target endpoint is part of a multi-port endpoint, `--remove-endpoint` removes only the specified port and keeps the rest.
### Merge Semantics
OpenShell applies all update flags from one `openshell policy update` command as one merge batch. The gateway validates the full merged result and persists at most one new policy revision.
This means:
- one command is atomic at the revision level.
- multiple flags in one command succeed or fail together.
- concurrent writers do not partially interleave one batch with another.
When two updates race, the gateway uses optimistic retry. It fetches the latest revision, reapplies the full batch, validates the result again, and retries the write. This preserves the intent of each individual command while still allowing concurrent sandbox policy updates.
### Preview and Validation
Use `--dry-run` when you want to inspect the merged YAML before you send it to the gateway.
```shell
openshell policy update demo \
--add-allow 'api.github.com:443:GET:/repos/**' \
--dry-run
```
The CLI validates the argument shapes before it sends the request. The gateway then validates the merged policy against the current live policy and returns clear errors when:
- a required segment is missing.
- a port is outside `1` through `65535`.
- `--add-allow` or `--add-deny` points at an endpoint that does not exist.
- `--add-allow` or `--add-deny` targets an endpoint that is neither REST nor WebSocket.
- `--add-allow` or `--add-deny` targets a host and port that resolves to more than one endpoint.
- `--add-deny` targets an endpoint that has no base allow set.
- an update names an existing rule and adds a binary to it without declaring every endpoint and port that rule already authorizes.
- an update names an existing rule and adds or changes an endpoint on it without declaring every binary that rule already authorizes.
- an update widens a rule to any binary without declaring every endpoint that rule already authorizes.
- an update changes an endpoint without declaring every port that endpoint carries.
- an update puts an MCP endpoint on the same host and port as a differently inspected endpoint, or gives one host and port two different MCP inspection contracts, including through separate rules or different paths.
A rule authorizes every listed binary to reach every listed endpoint and port, so merging a binary and an endpoint into the same rule authorizes that pair too. The gateway rejects the whole batch rather than granting a pair the update did not ask for. The error names the binary scope and the ports involved, and lists the binaries you still need to declare. An empty binary list means any binary, so widening a rule to any binary is subject to the same requirement.
You can declare the scope across several `--add-endpoint` arguments. The update is complete as long as every binary-to-port pair the merged rule ends up authorizing appears somewhere in the update.
An endpoint's allow rules, deny rules, and allowed IPs apply to every port that endpoint carries, so an update that changes any of them has to name every one of those ports. Declaring `api.example.com:443` alone on an endpoint that also serves `8443` is rejected, because the change would reach `8443` as well. Declaring every existing binary does not lift this requirement; the two are separate axes of the same product.
The sandbox picks the parser for a request by most-specific path, but it authorizes the request against every endpoint that matches it. A broad REST endpoint and a narrower GraphQL endpoint on one host and port share the same method-and-path rule vocabulary, so that combination stays supported. MCP does not: its rules address JSON-RPC methods and tool names, so a plain REST rule on an overlapping path could authorize a tool call the MCP endpoint denies. An MCP endpoint therefore cannot share a host and port with a differently inspected endpoint, and two MCP endpoints there must agree on strict-tool-name, method-profile, and body-limit settings, even under different paths or in separate rules. An update creating either situation is rejected. A policy that already contains one is left alone so unrelated updates still apply, but it should be repaired with full YAML replacement.
To grant one binary access to only part of an existing rule's endpoints, send it under its own `--rule-name`. The gateway normally folds an update into an existing rule that shares an endpoint, but it keeps your rule name whenever folding would grant authorization you did not declare, so the narrow grant lands as its own rule authorizing exactly what you asked for. The update reports that it kept your rule name and names the rule it would otherwise have folded into. An MCP contract conflict is the exception: one host and port carry a single MCP contract regardless of which rule holds them, so a conflicting update is rejected rather than moved to a separate rule.
Once a host and port appears in more than one rule, `--add-allow` and `--add-deny` can no longer target it. They select an endpoint by host and port alone, so they cannot say which rule's binary scope to widen, and the gateway rejects the update rather than guessing. The same applies when one rule carries two endpoints on that host and port under different paths. Use full YAML replacement to change L7 rules on an endpoint that appears more than once.
## Global Policy Override
Use a global policy when you want one policy payload to apply to every sandbox.
```shell
openshell policy set --global --policy ./global-policy.yaml
```
When a global policy is configured:
- The global payload is applied in full for all sandboxes.
- Sandbox-level policy updates are rejected until the global policy is removed.
To restore sandbox-level policy control, delete the global policy setting:
```shell
openshell policy delete --global
```
You can inspect a sandbox's effective settings and policy source with:
```shell
openshell settings get <name>
```
## Debug Denied Requests
Check `openshell logs <name> --tail --source sandbox` for the denied host, path, and binary.
For agent-authored draft updates on running sandboxes, enable [Policy Advisor](/sandboxes/policy-advisor). Policy advisor lets the sandboxed agent submit a narrow proposal through `policy.local` while a developer still approves or rejects the structured rule from outside the sandbox.
When triaging denied requests, check:
- Destination host and port to confirm which endpoint is missing.
- Calling binary path to confirm which `binaries` entry needs to be added or adjusted.
- HTTP method and path for REST endpoints, or `GET` / `WEBSOCKET_TEXT` and the upgraded request path for WebSocket endpoints, to confirm which `rules` entry needs to be added or adjusted.
- `credential_endpoint_mismatch` in sandbox logs to confirm that policy admitted the request but the attached provider profile did not authorize its credential for that host, port, and path.
- `request_authority_mismatch` in the response or sandbox logs to confirm that the HTTP request authority differs from the authorized tunnel endpoint. For a CONNECT tunnel to `api.example.com:8443`, send `Host: api.example.com:8443`; omitting the non-default port makes the request authority use the transport default and OpenShell rejects it. Absolute-form request targets must use the same host and port.
Then push the updated policy as described above.
Do not fix `credential_endpoint_mismatch` by widening sandbox policy. Export the
provider profile with `openshell provider profile export <profile-id> -o yaml`.
Update the custom provider profile only when the destination is an intended
credential recipient. Refer to [Static Credential Endpoint
Binding](/sandboxes/providers-v2#understand-static-credential-endpoint-binding)
for the complete authorization model.
For small changes, prefer `openshell policy update` over rewriting the full YAML:
```shell
openshell policy update <name> --add-allow 'api.github.com:443:GET:/repos/**' --wait
```
## Examples
Add these blocks to the `network_policies` section of your sandbox policy. Apply simple endpoints and REST/WebSocket rule additions with `openshell policy update`, or apply any complete YAML block with `openshell policy set <name> --policy <file> --wait`.
Use **Simple endpoint** for host-level allowlists and **Granular rules** for method/path control.
<Tabs>
<Tab title="Simple endpoint">
Allow `pip install` and `uv pip install` to reach PyPI:
```yaml showLineNumbers={false}
pypi:
name: pypi
endpoints:
- host: pypi.org
port: 443
- host: files.pythonhosted.org
port: 443
binaries:
- { path: /usr/bin/pip }
- { path: /usr/local/bin/uv }
```
Endpoints without `protocol` use explicit-proxy TCP passthrough, where OpenShell allows the stream without inspecting payloads. Use `protocol: tcp` when the application needs ordinary DNS resolution and native TCP connections through transparent capture. Provider-credentialed endpoints cannot use either L4 shape unless `allow_uninspected_credentials: true` records the exception. If an explicit-proxy stream is HTTP and TLS is auto-terminated, the proxy can still rewrite configured credential placeholders and closes keep-alive passthrough tunnels on policy reload before forwarding another request. WebSocket text-frame policy requires an explicit `protocol: websocket` endpoint. WebSocket payload credential rewrite can also be enabled on a `protocol: rest` compatibility endpoint with `websocket_credential_rewrite: true`. REST request body credential rewrite requires an inspected `protocol: rest` endpoint with `request_body_credential_rewrite: true`.
</Tab>
<Tab title="Granular rules">
Allow Claude and the GitHub CLI to reach `api.github.com` with separate REST and GraphQL endpoint scopes: read-only REST for general API paths, GraphQL operation inspection on `/graphql`, full REST write access for `alpha-repo`, and create/edit issues only for `bravo-repo`. Replace `<org_name>` with your GitHub org or username.
<Tip>
For an end-to-end walkthrough that combines this policy with a GitHub credential provider and sandbox creation, refer to [GitHub Sandbox](/get-started/tutorials/github-sandbox).
</Tip>
```yaml showLineNumbers={false}
github_repos:
name: github_repos
endpoints:
- host: api.github.com
port: 443
path: "/**"
protocol: rest
enforcement: enforce
rules:
- allow:
method: GET
path: "/**"
- allow:
method: HEAD
path: "/**"
- allow:
method: OPTIONS
path: "/**"
- allow:
method: "*"
path: "/repos/<org_name>/alpha-repo/**"
- allow:
method: POST
path: "/repos/<org_name>/bravo-repo/issues"
- allow:
method: PATCH
path: "/repos/<org_name>/bravo-repo/issues/*"
- host: api.github.com
port: 443
path: "/graphql"
protocol: graphql
enforcement: enforce
rules:
- allow:
operation_type: query
- allow:
operation_type: mutation
fields: [createIssue, updateIssue, addComment]
deny_rules:
- operation_type: mutation
fields: [deleteRepository, deleteRef, updateBranchProtectionRule]
binaries:
- { path: /usr/local/bin/claude }
- { path: /usr/bin/gh }
```
Endpoints with `protocol: rest` enable HTTP request inspection and can opt in to supported text request body credential rewrite. Endpoints with `protocol: websocket` validate WebSocket upgrades and inspect client text messages on the upgraded request path. WebSocket endpoints can also classify GraphQL-over-WebSocket operation messages with the same operation rules used by GraphQL-over-HTTP. Endpoints with `protocol: graphql` parse GraphQL-over-HTTP payloads before evaluating rules. Endpoints with `protocol: mcp` parse MCP Streamable HTTP request bodies and evaluate `method`, optional `tool`, and supported params rules. Endpoints with `protocol: json-rpc` parse JSON-RPC-over-HTTP request bodies and evaluate `method` rules. The endpoint-level `path` field lets these protocols share `api.github.com:443` without treating GraphQL payloads as plain REST `POST /graphql` requests.
</Tab>
</Tabs>
### Query parameter matching
REST rules can also constrain query parameter values:
```yaml showLineNumbers={false}
download_api:
name: download_api
endpoints:
- host: api.example.com
port: 443
protocol: rest
enforcement: enforce
rules:
- allow:
method: GET
path: "/api/v1/download"
query:
slug: "skill-*"
version:
any: ["1.*", "2.*"]
binaries:
- { path: /usr/bin/curl }
```
`query` matchers are case-sensitive and run on decoded values. If a request has duplicate keys (for example, `tag=a&tag=b`), every value for that key must match the configured glob(s).
### MCP and JSON-RPC matching
MCP endpoints use `protocol: mcp`. The proxy parses sandbox-to-server MCP Streamable HTTP request bodies, validates known MCP request and notification params, can evaluate the MCP method against rule `method`, and can match tool calls with the `tool` alias. Unknown extension methods stay addressable as literal method strings. Until OpenShell exposes MCP version profiles, `mcp.allow_all_known_mcp_methods` defaults to `false`, so endpoints require explicit MCP method rules. Set `mcp.allow_all_known_mcp_methods: true` to enable the endpoint method profile; in that mode, rules can omit `method`, and tool selectors are normalized to `tools/call` internally. By default, MCP `tools/call` tool names must match `^[A-Za-z0-9_.-]{1,128}$`; set `mcp.strict_tool_names: false` on that endpoint only when a server intentionally uses names outside the MCP-recommended pattern. Wildcard `tool` matchers require `mcp.strict_tool_names` to remain enabled. Generic JSON-RPC endpoints use `protocol: json-rpc` and evaluate `method`.
MCP endpoints must declare a concrete destination with `host` and `port` or `ports`. A policy entry that only sets `protocol: mcp` is invalid and is not treated as a wildcard MCP authorization. Use `path: /mcp` when the server's MCP endpoint is path-scoped; omitting `path` matches every HTTP path on that host and port.
MCP policy enforcement is directional. It applies to HTTP request bodies sent by the sandboxed process to the configured endpoint. JSON-RPC responses and server-to-client MCP messages carried on response bodies or SSE streams are relayed but are not currently parsed for policy enforcement.
MCP and JSON-RPC endpoint policies currently require full policy YAML applied with `openshell policy set`; the incremental `openshell policy update --add-endpoint` parser does not accept `mcp` or `json-rpc` as protocols.
An MCP client first sends `initialize`. After the server returns a successful response, the client sends `notifications/initialized`. After initialization completes and the server advertises the `tools` capability, the client can call an advertised tool. The response does not need an allow rule because these rules inspect messages sent from the client to the server. This example adds both client initialization messages to the existing tool rules. It omits `tools/list` because it assumes the client already knows the tool names; add that method when the client performs discovery.
```yaml showLineNumbers={false}
mcp_server:
name: mcp_server
endpoints:
- host: mcp.example.com
port: 443
path: /mcp
protocol: mcp
enforcement: enforce
mcp:
max_body_bytes: 131072
rules:
- allow:
method: initialize
- allow:
method: notifications/initialized
- allow:
method: tools/call
tool: read_status
- allow:
method: tools/call
tool:
any: [submit_report, list_reports]
deny_rules:
- method: tools/call
tool: delete_resource
binaries:
- { path: /usr/bin/python3 }
```
`mcp.max_body_bytes` controls how many MCP-over-HTTP request body bytes OpenShell buffers for inspection. It defaults to `65536`. `mcp.strict_tool_names` defaults to `true` for each MCP endpoint. `mcp.allow_all_known_mcp_methods` defaults to `false`; when it is unset or `false`, the endpoint must define explicit MCP method rules. If an MCP endpoint sets `mcp.allow_all_known_mcp_methods: true` and omits `rules`, OpenShell allows all MCP-family methods and all tools, then applies any `deny_rules`. A broad allow or deny rule whose method matcher includes `tools/call` cannot be combined with tool-specific allow rules because it would bypass or erase the tool filter; add `tool` or `params.name` to scope `tools/call`, or remove the tool-specific rules.
Use `protocol: json-rpc` and `method` when you need generic JSON-RPC 2.0 matching for a non-MCP server. Generic JSON-RPC method rules accept exact method names, or `method: "*"` as the all-method sentinel; other wildcard or glob patterns are rejected. `json_rpc.max_body_bytes` controls the generic JSON-RPC inspection buffer.
Generic JSON-RPC policy `params` matchers are not supported. Generic JSON-RPC policy rules match only the JSON-RPC method. For batch requests, OpenShell evaluates each JSON-RPC call independently and denies the whole batch if any call is denied.
For MCP, `tool` accepts a string glob or `{ any: [...] }` matcher for `tools/call` `params.name`. Rules that use `tool` or lower-level `params.name` must set `method: tools/call` unless `mcp.allow_all_known_mcp_methods: true` enables the endpoint method profile. MCP method globs are accepted only for the `tools/` method family, such as `tools/*`; omit `method` instead of writing `method: "*"` only when the endpoint method profile should allow all MCP methods. Omit `tool` to allow all tools for a `tools/call` method rule. OpenShell does not support MCP tool argument matching yet; allowed tools accept all argument payloads by default. Other MCP `params` keys are rejected. For batch requests, OpenShell evaluates each JSON-RPC call independently and denies the whole batch if any call is denied.
### GraphQL matching
GraphQL endpoints use `protocol: graphql`. The proxy parses GraphQL-over-HTTP `GET` and `POST` requests, classifies each operation, and evaluates rules against the operation type, optional operation name, and selected root fields.
GraphQL endpoint policies currently require full policy YAML applied with `openshell policy set`; the incremental `openshell policy update --add-endpoint` parser does not accept `graphql` as a protocol.
```yaml showLineNumbers={false}
github_graphql:
name: github_graphql
endpoints:
- host: api.github.com
port: 443
path: "/graphql"
protocol: graphql
enforcement: enforce
rules:
- allow:
operation_type: query
fields: [viewer, repository]
- allow:
operation_type: mutation
operation_name: Issue*
fields: [createIssue]
deny_rules:
- operation_type: mutation
fields: [deleteRepository]
binaries:
- { path: /usr/bin/gh }
```
For allow rules, every selected root field in an operation must match one of the configured `fields` globs. For deny rules, one matching root field blocks the request. Batched GraphQL requests are fail-closed: if any operation is malformed, denied, or unregistered, the whole HTTP request is denied.
Hash-only persisted queries cannot be classified from the request alone. OpenShell denies them unless the endpoint uses `persisted_queries: allow_registered` and provides a trusted `graphql_persisted_queries` entry keyed by hash or saved-query ID.
### GraphQL-over-WebSocket matching
Some APIs carry GraphQL operations over RFC 6455 WebSockets, commonly for subscriptions and realtime updates. Configure these as `protocol: websocket`, allow the upgrade with a normal `GET` rule, then add GraphQL operation rules for client operation messages. OpenShell recognizes modern `graphql-transport-ws` `subscribe` messages and legacy `graphql-ws` `start` messages.
```yaml showLineNumbers={false}
realtime_graphql:
name: realtime_graphql
endpoints:
- host: realtime.example.com
port: 443
path: "/graphql"
protocol: websocket
enforcement: enforce
rules:
- allow:
method: GET
path: "/graphql"
- allow:
operation_type: subscription
fields: [messageAdded]
- allow:
operation_type: query
fields: [viewer]
websocket_credential_rewrite: true
binaries:
- { path: /usr/bin/node }
```
When a WebSocket endpoint has GraphQL operation policy, client operation messages are fail-closed on malformed JSON, unsupported message types, parse errors, unregistered hash-only persisted queries, or unallowed operations. Use GraphQL operation rules for client messages rather than a raw `WEBSOCKET_TEXT` allow rule. Protocol lifecycle messages such as `connection_init`, `ping`, `pong`, and `complete` are allowed without payload logging; if `websocket_credential_rewrite: true` is set, placeholders inside those text messages are resolved before forwarding.
### GraphQL service policy shapes
GraphQL field names are application-specific, so treat these as starting shapes to review against the actual app schema:
| Service | Endpoint shape | Starting policy |
|---|---|---|
| Railway | `backboard.railway.app/graphql/v2` | Allow `query`; allow only reviewed deployment mutations; deny `volumeDelete`, `projectDelete`, `*Delete`, `*Destroy`. |
| GitHub | `api.github.com/graphql` | Allow `query`; optionally allow low-risk mutations such as reactions; deny broad destructive/admin roots like `deleteRef`, `deleteRepository`, `updateBranchProtectionRule`, and `delete*`. |
| GitLab | `/api/graphql` | Prefer read-only token scopes where possible; allow `query`; deny mutations by default or allow only reviewed workflow roots. |
| Shopify Admin | `*.myshopify.com/admin/api/**/graphql.json` | Allow `query`; allow app-specific mutations only; deny `*Delete`, `bulkOperationRunMutation`, and high-impact inventory/order/customer roots unless approved. |
| monday.com | `api.monday.com/v2` | Allow board/item reads; allow tightly scoped create/update mutations only where needed; deny delete/archive roots. |
| Salesforce GraphQL | Salesforce GraphQL endpoint | Allow `query`; deny record create/update/delete mutations unless the sandbox is intended to modify CRM data. |
| Hygraph | Project content API endpoint | Allow content reads; deny generated destructive content roots such as `delete*`, `deleteMany*`, `unpublish*`, and batch mutations unless a publishing workflow requires them. |
| Atlassian GraphQL Gateway | `api.atlassian.com/graphql` | Allow reads by default; require explicit mutation allowlists because the gateway spans Jira, Confluence, Bitbucket, and admin surfaces. |
## Next Steps
Explore related topics:
- To learn about the built-in sandbox policy, refer to [Default Policy](/reference/default-policy).
- To view the full field-by-field YAML definition, refer to the [Policy Schema Reference](/reference/policy-schema).
- To review the default policy breakdown, refer to [Default Policy](/reference/default-policy).
@@ -0,0 +1,266 @@
---
# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
title: "Use Policy Advisor"
sidebar-title: "Policy Advisor"
description: "Let sandboxed agents propose narrow policy changes through policy.local while keeping developer approval in the loop."
keywords: "Generative AI, Cybersecurity, Policy Advisor, Policy, Sandbox, policy.local, Agent Policy"
position: 7
---
Policy advisor lets a running sandboxed agent ask for a narrow network policy change after OpenShell denies a request. The agent submits a draft through `policy.local`, a developer approves or rejects it from outside the sandbox, and approved network policy hot-reloads into the same sandbox.
Policy advisor preserves OpenShell's default-deny posture. The structured rule is the approval contract, and the agent's rationale is supporting context. By default every accepted proposal lands in the draft inbox for human review. Opt-in [auto mode](#approval-modes) approves a proposal without a reviewer only when its [prover delta](#what-auto-approval-checks) is empty and the current draft rule produces no security notes. A prover finding or security note keeps the proposal pending for human review.
## Enable Policy Advisor
Policy advisor is disabled by default. Enable it globally when you want every sandbox on the selected gateway to expose the agent proposal surface:
```shell
openshell settings set --global \
--key agent_policy_proposals_enabled \
--value true \
--yes
```
You can also enable it for one sandbox, unless the key is managed globally:
```shell
openshell settings set <sandbox-name> \
--key agent_policy_proposals_enabled \
--value true
```
Check the effective setting for a sandbox:
```shell
openshell settings get <sandbox-name>
```
The output shows whether `agent_policy_proposals_enabled` is `global`, `sandbox`, or `unset`. A global value overrides sandbox-scoped values. To return control to sandbox-scoped settings, delete the global key:
```shell
openshell settings delete --global \
--key agent_policy_proposals_enabled \
--yes
```
Set the value before creating a sandbox when you want the first denied request to include policy advisor guidance. Running sandboxes poll settings and can enable the surface after startup, but startup enablement gives the agent the clearest first-denial path.
## Approval Modes
Every proposal, mechanistic or agent-authored, is routed through the [policy prover](#what-auto-approval-checks). The gateway also recalculates security notes from the current draft rule before auto-approval. The `proposal_approval_mode` setting decides whether proposals that pass both checks require human review.
| Mode | When unset / `manual` | `auto` |
|---|---|---|
| Empty prover delta and no security notes | Lands in the draft inbox for human review. | Approved automatically. The sandbox hot-reloads the new rule and the agent retries. |
| Any prover finding or security note | Lands in the draft inbox. | Remains pending for human review. |
`manual` is the default. Auto mode is an explicit opt-in; OpenShell's default-deny posture is preserved unless you choose otherwise.
Enable auto mode at gateway scope when you want every sandbox on this gateway to auto-approve eligible proposals:
```shell
openshell settings set --global \
--key proposal_approval_mode \
--value auto \
--yes
```
Enable it for one sandbox when no global value is set:
```shell
openshell settings set <sandbox-name> \
--key proposal_approval_mode \
--value auto
```
The shorthand at create time writes the sandbox-scoped setting for you:
```shell
openshell sandbox create --approval-mode auto <name>
```
Only `manual` and `auto` are accepted; typos like `autom` are rejected at configure time. Stale or unknown values found in storage are still treated as `manual` at runtime as a defense-in-depth measure.
**Precedence.** Gateway scope wins over sandbox scope. A reviewer can pin `manual` for a fleet by setting it globally; per-sandbox overrides only apply when no global value is set.
**Audit trail.** Every auto-approval emits a `CONFIG:APPROVED` event with `auto=true`, `source=<mechanistic|agent_authored>`, `prover_delta=empty`, and `resolved_from=<gateway|sandbox|default>` so operators can reconstruct why a given approval ran without human review.
## How It Works
When policy advisor is enabled, the sandbox supervisor turns on three agent-facing surfaces:
- It installs `/etc/openshell/skills/policy_advisor.md` inside the sandbox.
- It also installs `/etc/openshell/skills/policy-advisor/SKILL.md` as a short Codex/generic-agent pointer, and writes a root `/AGENTS.md` pointer only when the image does not already provide one.
- It serves `http://policy.local` from inside the sandbox.
- It adds `agent_guidance` and `next_steps` to L7 `policy_denied` response bodies so the agent can find the skill and local API.
The loop has seven steps:
1. A sandboxed process attempts a network request that policy denies.
2. For inspected REST traffic, OpenShell returns a structured `403` body with fields such as `layer`, `host`, `port`, `binary`, `method`, `path`, `rule_missing`, `agent_guidance`, and `next_steps`.
3. The agent reads the policy advisor skill, inspects the current policy, and optionally reads recent denial log lines.
4. The agent submits one or more `addRule` proposals to `http://policy.local/v1/proposals`.
5. The gateway turns the proposal into the exact effective-policy candidate it would apply. It preserves any existing L7 or provider-owned endpoint contract, adds the proposed binary as a sandbox overlay, validates the full merge, and runs the [policy prover](#what-auto-approval-checks) against that candidate.
6. The gateway stores the candidate, its prover result, any application error, and a review token tied to the live policy, provider rules, and credential metadata. Provider rules remain immutable inputs.
7. Before approval, the gateway cheaply recomputes the candidate token from live inputs. An unchanged token reuses the stored prover result. A changed token leaves the proposal pending with a refreshed candidate and requires a fresh review. Under `auto` mode, an unchanged candidate is approved only when the prover delta and security notes are empty. Under `manual` mode, every valid proposal lands in the draft inbox.
8. The agent waits on `/v1/proposals/{chunk_id}/wait` until a decision is available. Approved proposals hot-reload into the sandbox; rejected proposals return `rejection_reason` and `validation_result` so the agent can revise.
```mermaid
flowchart TD
A["Denied request creates a narrow proposal"] --> B["Gateway builds the exact effective-policy candidate"]
B --> C["Validate merge, L7 contract, providers, and credentials"]
C -->|Invalid| D["Keep pending and show application error"]
C -->|Valid| E["Run prover once and store candidate plus review token"]
E --> F["Reviewer approves using that token"]
F --> G["Recompute token from live inputs"]
G -->|Unchanged| H["Reuse stored prover result and apply candidate"]
G -->|Changed| I["Refresh candidate and token; require fresh review"]
I --> F
```
When a proposal is approved, `/wait` reports `policy_reloaded: true` only after the local sandbox policy covers the approved rule. At that point the agent can retry the original denied action once. If a proposal is rejected, `/wait` returns `rejection_reason` and `validation_result` so the agent can revise or stop. `validation_result` carries the categorical prover findings — `link_local_reach`, `l7_bypass_credentialed`, `credential_reach_expansion`, `capability_expansion` — so the agent can narrow the next attempt to the specific concern the prover flagged.
## What Gets Proposed
OpenShell has two proposal paths:
| Path | Source | Typical rule shape |
|---|---|---|
| Mechanistic mapper | Aggregated denial summaries from the sandbox. | Groups by host, port, and binary. If L7 request samples are available, it can draft REST method and path rules. Otherwise it drafts an L4 endpoint. |
| Agent-authored proposal | The in-sandbox agent, using `policy.local`. | Usually a REST `addRule` with exact host, port, binary, method, and path from the structured denial. It can also propose L4 rules for opaque protocols. |
For REST APIs, prefer L7 rules over broad L4 access. A good proposal allows one method and the smallest safe path:
```json
{
"intent_summary": "Allow gh to update documentation in NVIDIA/OpenShell.",
"operations": [
{
"addRule": {
"ruleName": "github_contents_docs_write",
"rule": {
"name": "github_contents_docs_write",
"endpoints": [
{
"host": "api.github.com",
"port": 443,
"protocol": "rest",
"enforcement": "enforce",
"rules": [
{
"allow": {
"method": "PUT",
"path": "/repos/NVIDIA/OpenShell/contents/docs/**"
}
}
]
}
],
"binaries": [
{
"path": "/usr/bin/gh"
}
]
}
}
}
]
}
```
The current `policy.local` JSON shape covers L4 endpoints and REST method or path rules. Use [Customize Sandbox Policies](/sandboxes/policies) or [Policy Schema Reference](/reference/policy-schema) for policy fields that are not part of the agent-authored proposal surface, such as WebSocket credential rewrite, GraphQL operation matching, endpoint path scoping, and provider-owned policy bundles.
Policy advisor proposals do not add `allowed_ips` automatically. If an advisor-proposed hostname resolves to an internal or private address, OpenShell's SSRF protections still block the connection until a developer explicitly adds the required `allowed_ips` entry. Exact hostname trust for user-declared policy endpoints does not apply to advisor-generated proposal binaries.
Private RFC 1918, CGNAT, IPv6 ULA, and other special-use destinations classified as internal produce advisory security notes when they appear as literal endpoint IPs or in `allowed_ips`. CIDR intersections are included, and hostless `allowed_ips` rules receive an additional warning because they can match any hostname resolving into the configured range.
Always-blocked destinations are not advisory. Loopback, link-local, and unspecified IPs or CIDRs, plus `localhost` and known metadata endpoint hostnames, are excluded from security notes. Submit and edit can store such a draft, but approval fails when merge validation prevents it from entering the active policy. Runtime SSRF protections continue to enforce the same boundary.
## What Auto-Approval Checks
Auto-approval requires all three conditions: the effective mode is `auto`, the prover delta is empty, and recalculating security notes from the current stored rule produces none.
The policy prover runs against mechanistic and agent-authored proposals alike and asks four formal questions about the proposed change. Each "yes" is one categorical finding. Any finding blocks auto-approval. An empty delta is necessary but not sufficient.
| Category | Triggered when |
|---|---|
| `link_local_reach` | A rule reaches `169.254.0.0/16`, `fe80::/10`, or a known metadata hostname. |
| `l7_bypass_credentialed` | A binary using a wire protocol the L7 proxy cannot inspect (`git-remote-https`, `ssh`, `nc`) gains reach to a host where a credential is in scope. |
| `credential_reach_expansion` | A binary gains credentialed reach to a `(host, port)` it could not reach before. |
| `capability_expansion` | On a `(binary, host, port)` that already had credentialed reach, the proposal adds a new HTTP method. The finding cites the specific method. |
Findings are categorical. There is no severity tier. The reviewer reads the category and the structured evidence to decide.
Before approval, the gateway rebuilds the candidate token from the live base policy, immutable provider rules, and non-secret credential metadata. When that token is unchanged, it reuses the persisted prover result instead of rerunning the prover. When it changes, the gateway evaluates and persists the refreshed candidate, leaves the chunk pending, and requires the reviewer to inspect and approve the new token. Edits and deduplicated resubmissions follow the same path. Merge, policy-shape, provider-composition, credential, or prover failures are shown as application errors and cannot be approved. Security notes flag concerns such as internal or private destinations and `allowed_ips`, wildcard hosts, hostless `allowed_ips`, ephemeral ports, and well-known database or service ports. Any prover finding or security note keeps the chunk pending in auto mode.
The full reasoning model lives in [`crates/openshell-prover/README.md`](https://github.com/NVIDIA/OpenShell/blob/main/crates/openshell-prover/README.md). Provider profiles composed in via [Providers v2](/sandboxes/providers-v2) are part of the effective policy the prover reasons over.
## Review Proposals
Review pending chunks from the host:
```shell
openshell rule get <sandbox-name> --status pending
```
Under `auto` mode, proposals with a prover finding or any recalculated security note remain pending for human review. Proposals that pass both checks are visible under `--status approved` with the auto-approval audit fields described in [Approval Modes](#approval-modes). Under `manual` mode, every accepted proposal shows up as pending regardless of the prover verdict or security notes.
The output shows the chunk ID, status, rationale, binary, endpoint summary, prover result, application error (if any), and candidate hash. For L7 proposals, the endpoint summary includes the protocol, method, and path:
```text
Endpoints: api.github.com:443 [L7 rest, allow PUT /repos/NVIDIA/OpenShell/contents/docs/**]
```
Approve only when the structured rule matches the access you intend to grant:
```shell
openshell rule approve <sandbox-name> --chunk-id <chunk-id>
```
The CLI fetches the current review token and submits it with the approval. If policy, provider, or credential inputs changed after the proposal was displayed, the gateway leaves it pending and asks you to review the refreshed candidate. Run `rule get` again before retrying. Bulk approval binds each selected chunk to its own review token in the same way.
Reject with guidance when the rule is too broad or points at the wrong target:
```shell
openshell rule reject <sandbox-name> \
--chunk-id <chunk-id> \
--reason "Scope this to docs/ paths only."
```
The rejection reason is returned to the agent through `policy.local`. The agent can use it to draft a narrower proposal.
Reviewers see the same guidance in the terminal UI. Run `openshell term`, open the sandbox's draft inbox, and select a rejected chunk to open its detail popup. The stored reason appears on a `Guidance:` line, and the list row shows a shortened copy of it. Rejection reasons have no length limit, so long guidance wraps and the popup body scrolls with `j`/`k`, `PageUp`/`PageDown`, and `g`/`G`; the approve and close controls stay pinned below the body, and the bottom border shows the scroll position.
## Agent API
`policy.local` is available only inside the sandbox and uses plain HTTP:
| Endpoint | Purpose |
|---|---|
| `GET /v1/policy/current` | Returns the current effective sandbox policy as YAML. |
| `GET /v1/denials?last=10` | Returns recent denied OCSF shorthand log lines, newest first. Query strings are redacted before lines are returned to the agent. |
| `POST /v1/proposals` | Submits `addRule` operations. The response includes `accepted_chunk_ids` and `rejection_reasons`. |
| `GET /v1/proposals/{chunk_id}` | Returns one proposal's current `pending`, `approved`, or `rejected` status. |
| `GET /v1/proposals/{chunk_id}/wait?timeout=300` | Holds one HTTP request open until the proposal is approved, rejected, or the timeout expires. |
If policy advisor is disabled, every route returns `404 feature_disabled`, the skill is not installed for new sandboxes, and L7 deny bodies do not advertise `policy.local` routes or include `agent_guidance`.
## What to Expect
Approved network rules hot-reload without restarting the sandbox. HTTP L7 keep-alive connections are closed at the reload boundary so the next parsed request uses the new policy. Raw streams remain connection-scoped, as described in [Customize Sandbox Policies](/sandboxes/policies#policy-structure).
Policy advisor emits audit events into the sandbox log. Use these lines to trace the full loop:
```shell
openshell logs <sandbox-name> --since 10m
```
Look for `HTTP:* DENIED`, `CONFIG:PROPOSED`, `CONFIG:APPROVED` or `CONFIG:REJECTED`, `CONFIG:LOADED`, and the final allowed request if the agent retries successfully. Auto-approved chunks emit `CONFIG:APPROVED` with `auto=true`, `source=<mechanistic|agent_authored>`, `prover_delta=empty`, and `resolved_from=<gateway|sandbox|default>`.
## Next Steps
- Use [Customize Sandbox Policies](/sandboxes/policies) for manual policy updates and L7 rule syntax.
- Use [Policy Schema Reference](/reference/policy-schema) for full YAML field details.
- Use [Logging](/observability/logging) to interpret OCSF shorthand log entries.
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,300 @@
---
# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
title: "OpenShell Security Best Practices for Controls, Risks, and Configuration Guidance"
sidebar-title: "Security Best Practices"
slug: "security/best-practices"
description: "A guide to every configurable security control in OpenShell: defaults, what you can change, and the risks of each choice."
keywords: "Generative AI, Cybersecurity, Security, Policy, Sandbox, Landlock, Seccomp"
position: 1
---
OpenShell enforces sandbox security across four layers: network, filesystem, process, and inference.
This page documents every configurable control, its default, what it protects, and the risk of relaxing it.
For the full policy YAML schema, refer to the [Policy Schema](/reference/policy-schema).
For the architecture of each enforcement layer, refer to [How OpenShell Works](/about/how-it-works).
<Note>
If you use [NemoClaw](https://github.com/NVIDIA/NemoClaw), its [Security Best Practices](https://docs.nvidia.com/nemoclaw/latest/security/best-practices.html) guide covers additional entrypoint-level controls, policy presets, provider trust tiers, and posture profiles specific to the NemoClaw blueprint.
</Note>
## Enforcement Layers
OpenShell applies security controls at two enforcement points.
OpenShell locks static controls at sandbox creation and requires destroying and recreating the sandbox to change them.
You can update dynamic controls on a running sandbox with `openshell policy update` or `openshell policy set`.
| Layer | What it protects | Enforcement point | Changeable at runtime |
| --- | --- | --- | --- |
| Network | Unauthorized outbound connections and data exfiltration. | CONNECT proxy + OPA policy engine | Yes. Use `openshell policy update`, `openshell policy set`, or operator approval in the TUI. |
| Filesystem | System binary tampering, credential theft, config manipulation. | Landlock LSM (kernel level) | No. Requires sandbox re-creation. |
| Process | Privilege escalation, fork bombs, dangerous syscalls. | Seccomp BPF + privilege drop (`setuid`/`setgid`) | No. Requires sandbox re-creation. |
| Inference | Credential exposure, unauthorized model access. | Proxy intercept of `inference.local` | Yes. Use `openshell inference set`. |
## Network Controls
The CONNECT proxy and OPA policy engine enforce all network controls at the gateway level.
### Deny-by-Default Egress
Every outbound connection from the sandbox goes through the CONNECT proxy.
The proxy evaluates each connection against the OPA policy engine.
If no `network_policies` entry matches the destination host, port, and calling binary, the proxy denies the connection.
| Aspect | Detail |
|---|---|
| Default | All egress denied. Only endpoints listed in `network_policies` can receive traffic. |
| What you can change | Add entries to `network_policies` in the policy YAML. Apply statically at creation (`--policy`) or dynamically (`openshell policy update` for incremental changes, `openshell policy set` for full replacement). |
| Risk if relaxed | Each allowed endpoint is a potential data exfiltration path. The agent can send workspace content, credentials, or conversation history to any reachable host. |
| Recommendation | Add only endpoints the agent needs for its task. Start with a minimal policy and use denied-request logs (`openshell logs <name> --source sandbox`) to identify missing endpoints. |
### Network Namespace Isolation
The sandbox runs in a dedicated Linux network namespace with a veth pair.
All traffic routes through the host-side veth IP (`10.200.0.1`) where the proxy listens.
Even if a process ignores proxy environment variables, it can only reach the proxy.
| Aspect | Detail |
|---|---|
| Default | Always active. The sandbox cannot bypass the proxy at the network level. |
| What you can change | This is not a user-facing knob. OpenShell always enforces it in proxy mode. |
| Risk if bypassed | Without network namespace isolation, a process could connect directly to the internet, bypassing all policy enforcement. |
| Recommendation | No action needed. OpenShell enforces this automatically. |
### User Namespace Isolation
Kubernetes user namespaces (`hostUsers: false`) map container UID 0 to an unprivileged host UID.
Capabilities like `CAP_SYS_ADMIN` become namespaced. They grant power over container-local resources only, not the host.
This provides defense-in-depth: even if a container escape vulnerability exists, the attacker lands as an unprivileged host user.
| Aspect | Detail |
|---|---|
| Default | Disabled. Set `server.enableUserNamespaces: true` in Helm values or `enable_user_namespaces = true` in the gateway config to enable cluster-wide. |
| What you can change | Enable cluster-wide through Helm or gateway config. Override per-sandbox through the `user_namespaces` field on `SandboxTemplate` in the API. |
| Prerequisites | Kubernetes 1.33+ with user namespace support available (beta through 1.35, GA in 1.36+), a container runtime that supports user namespaces (containerd 2.0+, CRI-O 1.25+), and Linux 5.12+ for ID-mapped mounts. |
| Risk if enabled with GPU | NVIDIA device plugin compatibility with user namespaces is unverified. OpenShell logs a warning when both GPU and user namespaces are active on the same sandbox. |
| Recommendation | Enable on non-GPU clusters running Kubernetes with user namespace support available (1.33+ beta, 1.36+ GA) for stronger host isolation. Test GPU workloads separately before enabling on GPU clusters. |
### Binary Identity Binding
The proxy identifies which binary initiated each connection by reading `/proc/<pid>/exe` (the kernel-trusted executable path).
It walks the process tree for ancestor binaries and parses `/proc/<pid>/cmdline` for script interpreters.
The proxy SHA256-hashes each binary on first use (trust-on-first-use). If someone replaces a binary mid-session, the hash mismatch triggers an immediate deny.
| Aspect | Detail |
|---|---|
| Default | Every `network_policies` entry requires a `binaries` list. Only listed binaries can reach the associated endpoints. Binary paths support glob patterns (`*` for one path component, `**` for recursive). |
| What you can change | Add binaries to an endpoint entry. Use glob patterns for directory-scoped access (for example, `/sandbox/.vscode-server/**`). |
| Risk if relaxed | Broad glob patterns (like `/**`) allow any binary to reach the endpoint, defeating the purpose of binary-scoped enforcement. |
| Recommendation | Scope binaries to the specific executables that need each endpoint. Use narrow globs when the exact path varies (for example, across Python virtual environments). |
### L4-Only vs L7 Inspection
The `protocol` field on an endpoint controls whether the proxy inspects individual HTTP requests inside the tunnel.
| Aspect | Detail |
|---|---|
| Default | Endpoints without a `protocol` field use L4-only enforcement: the proxy checks host, port, and binary, then relays the TCP stream without inspecting payloads. Provider-credentialed endpoints reject this mode unless an operator explicitly opts in. |
| What you can change | Add `protocol: rest` to enable per-request HTTP method/path inspection, `protocol: websocket` to inspect RFC 6455 upgrade handshakes and client text messages, or `protocol: graphql` to inspect GraphQL-over-HTTP operation type, operation name, and root fields. WebSocket endpoints can also use GraphQL operation rules for GraphQL-over-WebSocket messages. Pair inspected protocols with `rules` or access presets (`full`, `read-only`, `read-write`). REST endpoints that need credential placeholders in supported text request bodies can set `request_body_credential_rewrite: true`. Set `allow_uninspected_credentials: true` only as an explicit exception for credentialed traffic that cannot use an inspected path. |
| Risk if relaxed | L4-only endpoints allow the agent to send any data through the tunnel after the initial connection is permitted. The proxy cannot see HTTP methods, paths, or GraphQL operations. Adding `access: full` with L7 inspection enables observability but permits all inspected actions. |
| Recommendation | Use `protocol: rest` with specific `rules` for APIs where intent is encoded in method and path. Add `request_body_credential_rewrite: true` only for REST APIs that require OpenShell-managed credentials in UTF-8 JSON, form, or text request bodies. Use `protocol: graphql` for GraphQL-over-HTTP APIs where destructive operations are body-encoded. Use `protocol: websocket` for RFC 6455 endpoints, with explicit `GET` and `WEBSOCKET_TEXT` rules for raw text protocols or explicit GraphQL operation rules for GraphQL-over-WebSocket. Prefer `access: read-only` or explicit allowlists, and deny hash-only persisted queries unless you maintain a trusted registry. Omit `protocol` for non-HTTP protocols. For WebSocket endpoints that must carry placeholder credentials in client text frames, add `websocket_credential_rewrite: true`. |
### Enforcement Mode (`audit` vs `enforce`)
When L7 inspection is active, the `enforcement` field controls whether the proxy blocks or logs rule violations.
| Aspect | Detail |
|---|---|
| Default | `audit`. The proxy logs violations but forwards traffic. |
| What you can change | Set `enforcement: enforce` to block requests that do not match any `rules` entry. Denied requests receive a `403 Forbidden` response with a JSON body describing the violation. |
| Risk if relaxed | `audit` mode provides visibility but does not prevent unauthorized actions. An agent can still perform write or delete operations on an API even if the rules would deny them. |
| Recommendation | Start with `audit` to understand traffic patterns and verify that rules are correct. Switch to `enforce` after you validate that the rules match the intended access pattern. |
### TLS Handling
The proxy auto-detects TLS on every tunnel by peeking the first bytes.
When a TLS ClientHello is detected, the proxy terminates TLS transparently using a per-sandbox ephemeral CA.
This enables credential injection and L7 inspection without explicit configuration.
| Aspect | Detail |
|---|---|
| Default | Auto-detect and terminate. OpenShell generates the sandbox CA at startup and injects it into the process trust stores (`NODE_EXTRA_CA_CERTS`, `DENO_CERT`, `SSL_CERT_FILE`, `REQUESTS_CA_BUNDLE`, `CURL_CA_BUNDLE`, `GIT_SSL_CAINFO`). |
| What you can change | Set `tls: skip` on an endpoint to disable TLS detection and termination for that endpoint. Use this for client-certificate mTLS to upstream or non-standard binary protocols. |
| Risk if relaxed | `tls: skip` disables placeholder credential rewriting, dynamic token grant injection, and L7 inspection for that endpoint. The proxy relays encrypted traffic without seeing the contents. |
| Recommendation | Use auto-detect (the default) for most endpoints. Use `tls: skip` only when the upstream requires the client's own TLS certificate (mTLS) or uses a non-HTTP protocol. A provider-credentialed endpoint also requires the explicit `allow_uninspected_credentials: true` exception. |
### SSRF Protection
After OPA policy allows a connection, the proxy resolves DNS and rejects undeclared internal destinations.
| Aspect | Detail |
|---|---|
| Default | The proxy blocks private IPs for undeclared, wildcard, hostless, and policy-advisor-proposed endpoints. Exact hostnames declared in user policy may resolve to private RFC 1918 addresses. Loopback (`127.0.0.0/8`), link-local (`169.254.0.0/16`), and unspecified (`0.0.0.0`) addresses are always blocked and cannot be overridden with `allowed_ips`. |
| What you can change | Declare a known internal service as an exact `host` endpoint, or add `allowed_ips` (CIDR notation) to an endpoint to permit specific private IP ranges for wildcard or hostless policies. On Linux, the proxy consults the sandbox's `/etc/hosts` before DNS, so Kubernetes `hostAliases` can make LAN-only hostnames resolvable. Policies with `allowed_ips` entries that overlap loopback, link-local, or unspecified addresses fail to load with a clear validation error. |
| Risk if relaxed | Without SSRF protection, a misconfigured policy could allow the agent to reach cloud metadata services (`169.254.169.254`), internal databases, or other infrastructure endpoints through DNS rebinding. |
| Recommendation | Prefer exact hostname endpoints for stable internal services. Use `allowed_ips` when you need hostless or wildcard authorization, and scope the CIDR as narrowly as possible (for example, `10.0.5.20/32` for a single host). Loopback, link-local, and unspecified addresses are always blocked regardless of `allowed_ips`. `hostAliases` change resolution, not authorization: `host: searxng.local` with `/etc/hosts` mapping to `192.168.1.105` is trusted only when that exact hostname is declared by user policy. The policy advisor does not propose rules for always-blocked destinations and still requires a separate `allowed_ips` edit for private-IP endpoints. |
### Operator Approval
When the agent requests an endpoint not in the policy, OpenShell blocks it and surfaces the request in the TUI for operator review.
The system merges approved endpoints into the sandbox's policy as a new durable revision.
| Aspect | Detail |
|---|---|
| Default | Enabled. The proxy blocks unlisted endpoints and requires approval. |
| What you can change | Approved endpoints persist across sandbox restarts within the same sandbox instance. They reset when the sandbox is destroyed and recreated. |
| Risk if relaxed | Approving an endpoint permanently widens the running sandbox's policy. Review each request before approving. |
| Recommendation | Use operator approval for exploratory work. For recurring endpoints, add them to the policy YAML with appropriate binary and path restrictions. To reset all approved endpoints, destroy and recreate the sandbox. |
## Filesystem Controls
Landlock LSM restricts which paths the sandbox process can read or write at the kernel level.
### Landlock LSM
Landlock enforces filesystem access at the kernel level.
Paths listed in `read_only` receive read-only access.
Paths listed in `read_write` receive full access.
All other paths are inaccessible.
Landlock setup runs in two phases. The parent supervisor probes the kernel ABI and opens the configured path file descriptors before forking. The child then applies the ruleset with `restrict_self()` after privilege drop. At startup, OpenShell emits the selected ABI version and the applied read-only and read-write rule counts so you can confirm what the kernel accepted.
| Aspect | Detail |
|---|---|
| Default | `compatibility: best_effort`. Uses the highest kernel ABI available. Missing paths are skipped. If the kernel does not support Landlock or any configured path cannot be opened, the sandbox continues without those restrictions and emits a High-severity OCSF `DetectionFinding`. |
| What you can change | Set `compatibility: hard_requirement` to abort sandbox startup if Landlock is unavailable or any configured path cannot be opened. |
| Risk if relaxed | On kernels without Landlock (pre-5.13), or when all paths fail to open, the sandbox runs without kernel-level filesystem restrictions. The agent can access any file the process user can access. |
| Recommendation | Use `best_effort` for development. Use `hard_requirement` in environments where any gap in filesystem isolation is unacceptable. Treat High-severity Landlock findings as a signal to investigate the host kernel or the image. Run on Ubuntu 22.04+ or any kernel 5.13+ for Landlock support. |
### Read-Only vs Read-Write Paths
The policy separates filesystem paths into read-only and read-write groups.
| Aspect | Detail |
|---|---|
| Default | System paths (`/usr`, `/lib`, `/etc`, `/var/log`) are read-only. The resolved working directory and `/tmp` are read-write. `/app` is conditionally included if it exists. |
| What you can change | Add or remove paths in `filesystem_policy.read_only` and `filesystem_policy.read_write`. |
| Risk if relaxed | Making system paths writable lets the agent replace binaries, modify TLS trust stores, or change DNS resolution. Validation rejects broad read-write paths (like `/`). |
| Recommendation | Keep system paths read-only. If the agent needs additional writable space, add a specific subdirectory. |
### Path Validation
OpenShell validates policies before they take effect.
| Constraint | Behavior |
|---|---|
| Paths must be absolute (start with `/`). | Rejected with `INVALID_ARGUMENT`. |
| Paths must not contain `..` traversal. | Rejected with `INVALID_ARGUMENT`. |
| Read-write paths must not be overly broad (for example, `/` alone). | Rejected with `INVALID_ARGUMENT`. |
| Each path must not exceed 4096 characters. | Rejected with `INVALID_ARGUMENT`. |
| Combined `read_only` + `read_write` paths must not exceed 256. | Rejected with `INVALID_ARGUMENT`. |
## Process Controls
The sandbox supervisor drops privileges, applies seccomp filters, and enforces process-level restrictions during startup.
### Privilege Drop
The sandbox process runs as a non-root user after explicit privilege dropping.
| Aspect | Detail |
|---|---|
| Default | The compute driver selects a non-root identity. Docker and Podman use the image's OCI `USER` as a per-field fallback. The supervisor calls `setuid()`/`setgid()` with post-condition verification, disables core dumps with `RLIMIT_CORE=0`, and on Linux sets `PR_SET_DUMPABLE=0`. |
| What you can change | Set either or both `run_as_user` and `run_as_group` fields in the `process` section. Each explicit field takes precedence and must be `sandbox` or a numeric UID/GID from `1` through `4294967294`. Docker and Podman may use named identities through OCI `USER` fallback. Root and the invalid identity sentinel are always rejected. |
| Risk if relaxed | A numeric identity inherits every permission granted to the same UID/GID on image files, mounted volumes, or devices. Low IDs often collide with system accounts and groups. |
| Recommendation | Use a dedicated non-root image identity or explicit numeric policy identity. Do not attempt to set root. |
### Seccomp Filters
OpenShell applies seccomp in two phases. A narrow supervisor-startup prelude runs before CLI parsing and async runtime initialization, then the child process receives the broader runtime seccomp filter after privilege drop.
| Aspect | Detail |
|---|---|
| Startup prelude | After privileged bootstrap helpers complete, including network setup and provider-token SPIFFE child mount-namespace preparation, the supervisor sets `PR_SET_NO_NEW_PRIVS` and synchronizes a seccomp filter across all runtime threads that blocks `mount`, the new mount API syscalls, `pivot_root`, `umount2`, `bpf`, `perf_event_open`, `userfaultfd`, module-loading syscalls, and kexec. This closes the long-lived privileged remount and kernel-surface window while leaving required setup syscalls such as `setns` available. |
| Socket domains | The filter allows `AF_INET` and `AF_INET6` (for proxy communication) and blocks `AF_PACKET`, `AF_BLUETOOTH`, and `AF_VSOCK` with `EPERM`. `AF_NETLINK` is partially allowed: only `NETLINK_ROUTE` (protocol 0) is permitted so that `getifaddrs(3)` works; all other netlink protocols are blocked. Write operations via `NETLINK_ROUTE` still require `CAP_NET_ADMIN`, which the sandbox does not grant. |
| Runtime unconditional syscall blocks | `memfd_create`, `ptrace`, `bpf`, `process_vm_readv`, `process_vm_writev`, `pidfd_open`, `pidfd_getfd`, `pidfd_send_signal`, `io_uring_setup`, `mount`, `fsopen`, `fsconfig`, `fsmount`, `fspick`, `move_mount`, `open_tree`, `setns`, `umount2`, `pivot_root`, `userfaultfd`, `perf_event_open`. |
| Conditional syscall blocks | `execveat` with `AT_EMPTY_PATH`, `unshare` and `clone` with `CLONE_NEWUSER`, and `seccomp(SECCOMP_SET_MODE_FILTER)` are denied with `EPERM`. |
| What you can change | This is not a user-facing knob. OpenShell enforces it automatically. |
| Risk if relaxed | The blocked syscalls support container escape (`mount`, `pivot_root`, `move_mount`, namespace creation), cross-process observation (`ptrace`, `process_vm_readv`, `pidfd_*`), raw kernel bypass (`bpf`, `io_uring_setup`, `perf_event_open`), and filter evasion (`seccomp`, `userfaultfd`). |
| Recommendation | No action needed. OpenShell enforces this automatically. |
### Enforcement Application Order
The sandbox supervisor applies enforcement in a specific order during process startup.
This ordering is intentional: named network-namespace setup still relies on privileged helpers, and privilege dropping still needs `/etc/group` and `/etc/passwd`, which Landlock subsequently restricts.
1. Privileged supervisor bootstrap helpers, including network-namespace setup, provider-token SPIFFE child mount-namespace setup, and optional `nft` probes.
2. Supervisor startup prelude seccomp (`PR_SET_NO_NEW_PRIVS` plus the early syscall denylist) synchronized across runtime threads.
3. Network and child-only mount namespace entry (`setns`) in child `pre_exec`.
4. Privilege drop (`initgroups` + `setgid` + `setuid`).
5. Core-dump hardening (`RLIMIT_CORE=0`, plus `PR_SET_DUMPABLE=0` on Linux).
6. Landlock filesystem restrictions.
7. Runtime seccomp socket domain and syscall filters.
## Inference Controls
OpenShell routes all inference traffic through the gateway to isolate provider credentials from the sandbox.
### Routed Inference through `inference.local`
The proxy intercepts HTTPS CONNECT requests to `inference.local` and routes matching inference API requests through the sandbox-local router.
The agent never receives the provider API key.
| Aspect | Detail |
|---|---|
| Default | Always active. The proxy handles `inference.local` before OPA policy evaluation. The gateway injects credentials on the host side. |
| Keep-alive isolation | If a sandbox reuses a keep-alive connection that previously carried a routed inference request for a subsequent non-inference request, the proxy denies the non-inference request with `connection not allowed by policy` and closes the connection. This prevents agents from reusing an inference-authorized connection for other destinations. |
| What you can change | Configure inference routes with `openshell inference set`. |
| Risk if bypassed | If an inference provider's host is added directly to `network_policies`, the agent could reach it with a stolen or hardcoded key, bypassing credential isolation. |
| Recommendation | Do not add inference provider hosts to `network_policies`. Use OpenShell inference routing instead. |
## Gateway Security
The gateway secures communication between the CLI, sandbox workloads, and external clients with mutual TLS and token-based authentication.
### mTLS
Gateway transport uses TLS, with client certificate checks available where the deployment provides a client CA. Local single-user Docker, Podman, and VM gateways can use the verified client certificate as user authentication. Kubernetes deployments use the certificate bundle for transport and sandbox supervisor connectivity only; configure OIDC or a trusted access proxy for user authentication.
| Aspect | Detail |
|---|---|
| Default | Local TLS bundles enable mTLS user authentication for single-user local gateways. Helm deployments generate mTLS certificates for transport, while sandbox supervisors authenticate API calls with gateway-minted sandbox JWTs. TLS-enabled loopback gateways also accept plaintext HTTP for sandbox service hostnames by default. |
| What you can change | Configure OIDC or a trusted access proxy for multi-user gateways, set `OPENSHELL_ENABLE_MTLS_AUTH=true` for local single-user gateways, enable `server.auth.allowUnauthenticatedUsers=true` only for trusted local Kubernetes development or a fully trusted proxy, disable TLS only for trusted reverse-proxy setups, or disable loopback service HTTP with `--enable-loopback-service-http=false`. |
| Risk if relaxed | Disabling TLS removes transport-level protection entirely. Allowing unauthenticated users removes the gateway user-auth boundary and must not be exposed to shared or public networks. Treating transport certificates as shared user identity in Kubernetes would collapse user and sandbox trust boundaries. Loopback service HTTP is local-only and rejects cross-origin browser requests, but any local process can still reach exposed service URLs directly. |
| Recommendation | Use local mTLS user authentication only for single-user Docker, Podman, and VM gateways. Use OIDC or a trusted access proxy for Kubernetes and shared deployments. |
### SSH Tunnel Authentication
SSH connections to sandboxes travel through the gateway over the sandbox supervisor's authenticated control path. Each SSH connect call also carries a short-lived session token scoped to a specific sandbox. The sandbox never exposes an SSH port on the network. Its SSH daemon listens on a local Unix socket that only the sandbox's own supervisor process can reach.
| Aspect | Detail |
|---|---|
| Default | Session tokens expire after 24 hours. Concurrent connections are limited to 10 per token and 20 per sandbox. |
| What you can change | Configure `ssh_session_ttl_secs`. Set to 0 for no expiry. |
| Risk if relaxed | Longer TTLs or no expiry increase the window for stolen token reuse. Higher connection limits increase the blast radius of a compromised token. |
| Recommendation | Keep the 24-hour default. Monitor connection counts through the TUI. |
## Common Mistakes
The following patterns weaken security without providing meaningful benefit.
| Mistake | Why it matters | What to do instead |
|---------|---------------|-------------------|
| Omitting an inspected protocol on REST or WebSocket API endpoints | Without `protocol: rest` or `protocol: websocket`, the proxy uses L4-only enforcement. It allows the TCP stream through after checking host, port, and binary, but cannot inspect individual HTTP requests or WebSocket text messages. | Add `protocol: rest` or `protocol: websocket` with specific `rules` to enable method and path control. |
| Using `access: full` when finer rules would suffice | `access: full` with `protocol: rest` or `protocol: websocket` enables inspection but allows all methods and paths for that protocol. | Use `access: read-only` or explicit `rules` to restrict what the agent can do at the L7 level. |
| Adding endpoints permanently when operator approval would suffice | Adding endpoints to the policy YAML makes them permanently reachable across all instances. | Use operator approval. Approved endpoints persist within the sandbox instance but reset on re-creation. |
| Using broad binary globs | A glob like `/**` allows any binary to reach the endpoint, defeating binary-scoped enforcement. | Scope globs to specific directories (for example, `/sandbox/.vscode-server/**`). |
| Skipping TLS termination on HTTPS APIs | Setting `tls: skip` disables placeholder credential rewriting, dynamic token grant injection, and L7 inspection. | Use the default auto-detect behavior unless the upstream requires client-certificate mTLS. |
| Setting `enforcement: enforce` before auditing | Jumping to `enforce` without first running in `audit` mode risks breaking the agent's workflow. | Start with `audit`, review the logs, and switch to `enforce` after you validate the rules. |
| Pasting raw stack traces in bug reports | Some frameworks include the full request config — including credentials — in error objects. The sandbox does not scrub application-level output. | Inspect error output before sharing. Redact any credentials, API keys, or tokens. |
## Related Topics
- [Policies](/sandboxes/policies) for applying and iterating on sandbox policies.
- [Policy Schema](/reference/policy-schema) for the full field-by-field YAML reference.
- [Default Policy](/reference/default-policy) for the built-in default policy breakdown.
- [Gateway Auth](/reference/gateway-auth) for gateway authentication details.
- [How OpenShell Works](/about/how-it-works) for the system architecture.
- NemoClaw [Security Best Practices](https://docs.nvidia.com/nemoclaw/latest/security/best-practices.html) for entrypoint-level controls (capability drops, PATH hardening, build toolchain removal), policy presets, provider trust tiers, and posture profiles.
@@ -0,0 +1,36 @@
---
# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
title: "Verify the Contents of Published OpenShell Images"
sidebar-title: "Verify Image Contents"
slug: "security/verify-image-contents"
description: "Read the SBOM attestation attached to published gateway and supervisor images to audit what each image contains."
keywords: "Generative AI, Cybersecurity, Supply Chain, SBOM, Container Images"
position: 2
---
Published gateway and supervisor images carry one SPDX SBOM per platform as OCI attestations.
## Inspect an Image
Read a platform's document without pulling the image:
```shell
docker buildx imagetools inspect ghcr.io/nvidia/openshell/gateway:latest --format '{{ json (index .SBOM "linux/amd64").SPDX }}'
```
List the packages instead of the full document:
```shell
docker buildx imagetools inspect ghcr.io/nvidia/openshell/gateway:latest --format '{{ range (index .SBOM "linux/amd64").SPDX.packages }}{{ .name }}@{{ .versionInfo }}{{ println }}{{ end }}'
```
The same commands work for `ghcr.io/nvidia/openshell/supervisor`.
## Coverage
Every SBOM lists the base-image packages. Release Dev and Release Tag images also list the Rust crates compiled into their OpenShell binary.
<Note>
OpenShell also publishes minimal SLSA provenance. It records how BuildKit produced the image, including its source revision, build platform, and base-image materials, without the extra build parameters included by full provenance.
</Note>
+29
View File
@@ -0,0 +1,29 @@
landing-page:
page: Home
path: ../pages-v0.0.116/index.mdx
navigation:
- folder: ../pages-v0.0.116/about
title: About NVIDIA OpenShell
- section: Get Started
slug: get-started
contents:
- page: Quickstart
path: ../pages-v0.0.116/get-started/quickstart.mdx
- folder: ../pages-v0.0.116/get-started/tutorials
skip-slug: true
- folder: ../pages-v0.0.116/sandboxes
title: Manage OpenShell
- folder: ../pages-v0.0.116/providers
title: Providers
- folder: ../pages-v0.0.116/extensibility
title: Extensibility
- folder: ../pages-v0.0.116/observability
title: Observability
- folder: ../pages-v0.0.116/kubernetes
title: Kubernetes
- folder: ../pages-v0.0.116/reference
title: Reference
- folder: ../pages-v0.0.116/security
title: Security
- folder: ../pages-v0.0.116/resources
title: Resources