Max Dubrinsky e6f319c76d feat(sdk): add openshell-sdk crate (#1862)
* feat(sdk): add openshell-sdk crate

Additive extraction of the shared async gRPC client core (transport, TLS,
OIDC single-flight refresh, edge tunnel, high-level sandbox surface, raw
escape hatch) as a new workspace crate. No existing consumers yet; CLI/TUI
migration follows in a separate PR.

Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com>

* refactor(sdk,core): reuse shared JWT exp decoder for refresh deadlines

Extract the signature-unverified JWT exp decode out of
openshell-core/grpc_client.rs into openshell_core::jwt::parse_exp_secs,
and have the openshell-sdk refresh path reuse it to derive a proactive
refresh deadline from a bearer JWT when the caller does not advertise
expires_at. Addresses review feedback to reuse pre-existing logic rather
than reimplement JWT expiry handling per client.

Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com>

* fix(sdk): harden refresh single-flight and redact tokens in Debug

Addresses review feedback on the openshell-sdk refresh path.

- Make the single-flight cleanup cancellation-safe. The in-flight slot is
  now cleared by the shared refresh computation itself (epoch-guarded)
  rather than the leader's post-await code. Previously, if the leader future
  was dropped (e.g. an FFI caller cancelling its promise) after a follower
  drove the refresh to completion, the completed future was stranded in the
  slot and later refresh_now() calls re-joined it, pinning the client to a
  stale or already-rejected token. Adds a regression test that cancels the
  leader and asserts the next refresh starts a fresh attempt.

- Redact bearer secrets from Debug. RefreshedToken and the oidc
  RefreshTokenInput/RefreshTokenOutput now use manual Debug impls that omit
  the access/refresh token fields via finish_non_exhaustive, matching the
  house style (e.g. SecretResolver, SandboxJwtIssuer). Prevents a stray
  {:?} or a containing struct's derived Debug from writing tokens to logs.

Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com>

* fix(sdk): fail refresh when the new token can't be encoded as metadata

store_bearer now returns an error instead of silently keeping the previous
bearer value. The TokenSource commits the refreshed token to its state
before the client writes it into the interceptor slot, so a silent drop
left the interceptor on the old (expiring) token with no path back to a
refresh. Surfacing the error fails the call loudly instead. Adds a unit
test covering a token that can't be encoded as gRPC metadata.

Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com>

* fix(sdk): refresh OIDC tokens on raw routes and harden rotation

Raw gRPC access never triggered OIDC refresh: a client that only used
raw_grpc/raw_inference kept sending the initial bearer until it expired,
with no proactive or reactive refresh. Add raw_grpc_fresh and
raw_inference_fresh accessors that refresh before returning the client,
plus force_refresh for reactive recovery after an Unauthenticated raw
RPC.

Guard the single-flight refresh commit against a concurrent replace().
The in-flight attempt now records the generation it started from and
skips its write when an external replace() has advanced it, so timer or
callback driven rotation is no longer clobbered by a slower refresh.

Remove TokenSource::snapshot(): it returned an empty string under write
contention and had no consumer on the CLI/TUI path. Tests read committed
state directly instead.

Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com>

* docs(sdk): rewrite crate README for consumers

Recast the openshell-sdk README as a usable crate README rather than an
RFC excerpt. Drop the Responsibilities/Non-responsibilities/Consumers
scope-boundary sections and the mTLS migration rationale, folding the
useful facts (explicit token, no disk/name resolution, Refresh trait,
SdkError mapping) into the intro, a new Auth and refresh section, and
Public surface. Remove the dead relative RFC link and status-label
prose so the doc renders cleanly wherever it is published.

Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com>

* fix(sdk): address review feedback on OIDC refresh and transport

Apply Drew's review notes on PR #1862:

- Drop unused `rustls-pemfile` dependency and move `tokio-stream` to
  dev-dependencies (only used by tests).
- Guard OIDC `expires_at` against u64 overflow with `saturating_add`.
- Fix stale `#[non_exhaustive]` rationale in `AuthConfig` (the struct
  `Oidc` variant it described as future already ships).
- Stop double-wrapping refresh errors: store the bare refresh-error text
  so the single `SdkError::auth` wrap happens once at await.
- Strip stale CLI porting breadcrumbs from `build_channel` docs, keeping
  the branch table.
- Treat proactive token refresh as best-effort: a transient failure falls
  through to the request instead of failing an RPC whose current token is
  still valid, with a regression test.
- Collapse `exec`'s inline auth retry into the shared `unary` helper.

Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com>

* fix(sdk): preserve transient/terminal distinction in refresh errors

The refresh single-flight collapsed both `RefreshError::Transient` and
`RefreshError::Terminal` into a stringified `SdkError::Auth`, so consumers
(CLI, TUI, future language bindings) had no machine-readable way to tell a
retryable IdP blip from a dead session that needs re-authentication.

Carry the `RefreshError` through the shared outcome (kept `Clone` for
`Shared`) instead of its rendered text, and map it at the await site to a
new `retryable` flag on `SdkError::Auth`. Add `SdkError::auth_retryable`
and a `SdkError::retryable()` accessor; transient refresh failures report
`true`, every other error `false`. Add classification tests.

Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com>

---------

Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com>
2026-07-14 22:35:46 -07:00

NVIDIA OpenShell

License PyPI Security Policy Documentation Project Status

OpenShell is the safe, private runtime for autonomous AI agents. It provides sandboxed execution environments that protect your data, credentials, and infrastructure — governed by declarative YAML policies that prevent unauthorized file access, data exfiltration, and uncontrolled network activity.

OpenShell is built agent-first. The project ships with agent skills for everything from gateway troubleshooting to policy generation, and we expect contributors to use them.

Alpha software — single-player mode. OpenShell is proof-of-life: one developer, one environment, one gateway. We are building toward multi-tenant enterprise deployments, but the starting point is getting your own environment up and running. Expect rough edges. Bring your agent.

Quickstart

Prerequisites

  • A supported host — macOS, Windows with WSL 2, or Linux.
  • A local runtime — Docker, Podman, or host virtualization enabled for MicroVM-backed sandboxes.

Install

Binary (recommended):

curl -LsSf https://raw.githubusercontent.com/NVIDIA/OpenShell/main/install.sh | sh

From PyPI (requires uv):

uv tool install -U openshell

Both methods install the latest stable release by default. To install a specific version, set OPENSHELL_VERSION (binary) or pin the version with uv tool install openshell==<version>. A dev release is also available that tracks the latest commit on main.

Helm chart:

Experimental — the Kubernetes deployment path is under active development. Expect rough edges and breaking changes.

Deploy the OpenShell gateway into a Kubernetes cluster from the OCI chart published to GHCR:

helm install openshell oci://ghcr.io/nvidia/openshell/helm-chart

See deploy/helm/openshell/README.md for available versions, dev tag conventions, and configuration.

For deploying OpenShell on OpenShift, see deploy/helm/openshell/README.md#install-on-openshift.

Create a sandbox

openshell sandbox create -- claude  # or opencode, codex, copilot

The sandbox container includes the following tools by default:

Category Tools
Agent claude, opencode, codex, copilot
Language python (3.14), node (22)
Developer gh, git, vim, nano
Networking ping, dig, nslookup, nc, traceroute, netstat

For more details see https://github.com/NVIDIA/OpenShell-Community/tree/main/sandboxes/base.

See network policy in action

Every sandbox starts with minimal outbound access. You open additional access with a short YAML policy that the proxy enforces at the HTTP method and path level, without restarting anything.

# 1. Create a sandbox (starts with minimal outbound access)
openshell sandbox create

# 2. Inside the sandbox — blocked
sandbox$ curl -sS https://api.github.com/zen
curl: (56) Received HTTP code 403 from proxy after CONNECT

# 3. Back on the host — apply a read-only GitHub API policy
sandbox$ exit
openshell policy set demo --policy examples/sandbox-policy-quickstart/policy.yaml --wait

# 4. Reconnect — GET allowed, POST blocked by L7
openshell sandbox connect demo
sandbox$ curl -sS https://api.github.com/zen
Anything added dilutes everything else.

sandbox$ curl -sS -X POST https://api.github.com/repos/octocat/hello-world/issues -d '{"title":"oops"}'
{"error":"policy_denied","detail":"POST /repos/octocat/hello-world/issues not permitted by policy"}

See the full walkthrough or run the automated demo:

bash examples/sandbox-policy-quickstart/demo.sh

How It Works

OpenShell isolates each sandbox in its own container with policy-enforced egress routing. A lightweight gateway coordinates sandbox lifecycle, and every outbound connection is intercepted by the policy engine, which does one of three things:

  • Allows — the destination and binary match a policy block.
  • Routes for inference — strips caller credentials, injects backend credentials, and forwards to the managed model.
  • Denies — blocks the request and logs it.
Component Role
Gateway Control-plane API that coordinates sandbox lifecycle and acts as the auth boundary.
Sandbox Isolated runtime with container supervision and policy-enforced egress routing.
Policy Engine Enforces filesystem, network, and process constraints from application layer down to kernel.
Privacy Router Privacy-aware LLM routing that keeps sensitive context on sandbox compute.

OpenShell runs a gateway control plane that manages sandbox lifecycle through a configured compute driver. Supported compute platforms include Docker, Podman, MicroVM, and Kubernetes.

Protection Layers

OpenShell applies defense in depth across four policy domains:

Layer What it protects When it applies
Filesystem Prevents reads/writes outside allowed paths. Locked at sandbox creation.
Network Blocks unauthorized outbound connections. Hot-reloadable at runtime.
Process Blocks privilege escalation and dangerous syscalls. Locked at sandbox creation.
Inference Reroutes model API calls to controlled backends. Hot-reloadable at runtime.

Policies are declarative YAML files. Static sections (filesystem, process) are locked at creation; dynamic sections (network, inference) can be hot-reloaded on a running sandbox with openshell policy set.

Providers

Agents need credentials — API keys, tokens, service accounts. OpenShell manages these as providers: named credential bundles that are injected into sandboxes at creation. The CLI auto-discovers credentials for recognized agents (Claude, Codex, OpenCode, Copilot) from your shell environment, or you can create providers explicitly with openshell provider create. Credentials never leak into the sandbox filesystem; they are injected as environment variables at runtime.

GPU Support (Experimental)

Experimental — GPU passthrough works on supported hosts but is under active development. Expect rough edges and breaking changes.

OpenShell can pass host GPUs into sandboxes for local inference, fine-tuning, or any GPU workload. Add --gpu when creating a sandbox:

openshell sandbox create --gpu --from [gpu-enabled-sandbox] -- claude

Docker-backed GPU sandboxes auto-select CDI when available and otherwise fall back to Docker's NVIDIA GPU request path (--gpus all).

Requirements: NVIDIA drivers and the NVIDIA Container Toolkit must be installed on the host. The sandbox image itself must include the appropriate GPU drivers and libraries for your workload — the default base image does not. See the BYOC example for building a custom sandbox image with GPU support.

Supported Agents

Agent Source Notes
Claude Code base Works out of the box. Provider uses ANTHROPIC_API_KEY.
OpenCode base Works out of the box. Provider uses OPENAI_API_KEY or OPENROUTER_API_KEY.
Codex base Works out of the box. Provider uses OPENAI_API_KEY.
GitHub Copilot CLI base Works out of the box. Provider uses GITHUB_TOKEN or COPILOT_GITHUB_TOKEN.
OpenClaw NemoClaw Run OpenClaw more securely inside NVIDIA OpenShell with managed inference using NemoClaw.
Hermes Agent NemoClaw Run Hermes Agent more securely inside NVIDIA OpenShell with managed inference using NemoClaw.
Ollama Community Launch with openshell sandbox create --from ollama.
Pi Community Launch with openshell sandbox create --from pi.

Key Commands

Command Description
openshell sandbox create -- <agent> Create a sandbox and launch an agent.
openshell sandbox connect [name] SSH into a running sandbox.
openshell sandbox list List all sandboxes.
openshell provider create --type [type] --from-existing Create a credential provider from env vars.
openshell policy set <name> --policy file.yaml Apply or update a policy on a running sandbox.
openshell policy get <name> Show the active policy.
openshell inference set --provider <p> --model <m> Configure the inference.local endpoint.
openshell logs [name] --tail Stream sandbox logs.
openshell term Launch the real-time terminal UI for debugging.

See the full documentation for command guides, tutorials, and reference material.

Terminal UI

OpenShell includes a real-time terminal dashboard for monitoring gateways, sandboxes, and providers — inspired by k9s.

openshell term

OpenShell Terminal UI

The TUI gives you a live, keyboard-driven view of your gateway and sandboxes. Navigate with Tab to switch panels, j/k to move through lists, Enter to select, and : for command mode. Gateway health and sandbox status auto-refresh every two seconds.

Community Sandboxes and BYOC

Use --from to create sandboxes from the OpenShell Community catalog, a local directory, or a container image:

openshell sandbox create --from gemini             # community catalog
openshell sandbox create --from ./my-sandbox-dir   # local Dockerfile
openshell sandbox create --from registry.io/img:v1 # container image

See the community sandboxes catalog and the BYOC example for details.

Explore with Your Agent

Clone the repo and point your coding agent at it. The project includes agent skills that can answer questions, walk you through workflows, and diagnose problems — no issue filing required.

git clone https://github.com/NVIDIA/OpenShell.git   # or git@github.com:NVIDIA/OpenShell.git
cd OpenShell
# Point your agent here — it will discover the skills in .agents/skills/ automatically

Your agent can load skills for CLI usage (openshell-cli), gateway troubleshooting (debug-openshell-cluster), inference troubleshooting (debug-inference), policy generation (generate-sandbox-policy), and more. See CONTRIBUTING.md for the full skills table.

Built With Agents

OpenShell is developed using the same agent-driven workflows it enables. The .agents/skills/ directory contains workflow automation that powers the project's development cycle:

  • Spike and build: Investigate a problem with create-spike, then implement it with build-from-issue once a human approves.
  • Triage and route: Community issues are assessed with triage-issue, classified, and routed into the spike-build pipeline.
  • Security review: review-security-issue produces a severity assessment and remediation plan. fix-security-issue implements it.
  • Policy authoring: generate-sandbox-policy creates YAML policies from plain-language requirements or API documentation.

All implementation work is human-gated — agents propose plans, humans approve, agents build. See AGENTS.md for the full workflow chain documentation.

Getting Help

  • Questions and discussion: GitHub Discussions
  • Bug reports: GitHub Issues — use the bug report template
  • Security vulnerabilities: See SECURITY.md — do not use GitHub Issues
  • Agent-assisted help: Clone the repo and use the agent skills in .agents/skills/ for self-service diagnostics

Learn More

Contributing

OpenShell is built agent-first — your agent is your first collaborator. Before opening issues or submitting code, point your agent at the repo and let it use the skills in .agents/skills/ to investigate, diagnose, and prototype. See CONTRIBUTING.md for the full agent skills table, contribution workflow, and development setup.

Telemetry

OpenShell collects anonymous telemetry to help improve the project for developers. This data is not used to track individual user behavior. It helps us understand aggregate usage of sandbox, provider, and policy workflows so we can prioritize product improvements and share usage trends with the community.

Disable telemetry at runtime by setting OPENSHELL_TELEMETRY_ENABLED=false on the gateway deployment. OpenShell propagates this deployment setting into sandbox supervisor environments so sandbox-side telemetry collection is disabled as well.

You can also compile telemetry out entirely. Telemetry support is a default-on telemetry Cargo feature; building with --no-default-features produces binaries that contain no telemetry endpoint, no telemetry HTTP client, and no emission code. Build telemetry-free artifacts with, for example, cargo build --release -p openshell-server --no-default-features (gateway) and the equivalent for openshell-sandbox and openshell-driver-vm. With telemetry compiled out, the gateway emits nothing and reports telemetry disabled to the sandboxes it launches.

Telemetry events are limited to anonymous operational categories and counts, such as sandbox lifecycle outcomes, provider profile buckets, policy decision counts, and aggregate network activity denial categories. OpenShell telemetry does not collect sandbox names or IDs, hostnames, file paths, binary paths, prompts, credentials, provider names, model names, or user content.

Opting out applies only to telemetry emitted by OpenShell. Third-party services, model providers, inference endpoints, agents, or tools that you configure and use with OpenShell may have their own terms and privacy practices.

We publish aggregate usage trends from this telemetry every two weeks. See the community telemetry reports for the latest summary.

Notice and Disclaimer

This software automatically retrieves, accesses or interacts with external materials. Those retrieved materials are not distributed with this software and are governed solely by separate terms, conditions and licenses. You are solely responsible for finding, reviewing and complying with all applicable terms, conditions, and licenses, and for verifying the security, integrity and suitability of any retrieved materials for your specific use case. This software is provided "AS IS", without warranty of any kind. The author makes no representations or warranties regarding any retrieved materials, and assumes no liability for any losses, damages, liabilities or legal consequences from your use or inability to use this software or any retrieved materials. Use this software and the retrieved materials at your own risk.

License

This project is licensed under the Apache License 2.0.

S
Description
GitHub Trending: NVIDIA/OpenShell
Readme Apache-2.0
182 MiB
Languages
Rust 86.3%
Go 5.4%
Shell 3.6%
Python 2.6%
TypeScript 0.9%
Other 1%