* refactor(inference): remove managed inference routes
Closes#3172
Remove the inference route control plane, inference.local data path, built-in router crate, and SDK surface. Move inference workloads to explicitly imported provider profiles and native endpoints, with migration cleanup and updated tests and documentation.
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
* fix(policy): preserve alternate upstream isolation
Restore the provider policy activation guard so legacy OpenAI and Anthropic providers configured for alternate base URLs do not grant egress to the built-in public vendor endpoints.
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
---------
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
The Maturin-based wheel packaging was a historical remnant from when the local gateway launch path and OpenShell CLI were coupled in one binary. The gateway and CLI now ship as standalone artifacts, so the Python distribution should contain only the SDK.
Build a single platform-independent setuptools wheel, verify that it cannot contain native code or an openshell entry point, and simplify the release jobs and documentation for SDK-only PyPI installs.
Signed-off-by: Simon Scatton <sscatton@nvidia.com>
Previously, the openshell snap used the ssh-keys interface to get access
to the host's ssh binary, which is used for sandbox connect/exec/forward.
However, ssh-keys is a privileged interface which also grants access to
the public and private ssh keys on the host. As such, it required manual
connection in order to be used.
This weakened the security sandbox of the snap, and hurt the UX of
installing it.
This commit changes this by removing the `ssh-keys` interface and
instead vendoring the `ssh` binary within the snap.
This is safe because OpenShell always invokes the `ssh` binary with
`StrictHostKeyChecking=no`, `UserKnownHostsFile=/dev/null`, and
`GlobalKnownHostsFile=/dev/null`, and never uses any host credentials or
ssh configuration. Openshell only ever access to `~/.ssh/config` to
write OpenShell-managed aliases, and this can safely live within the
snap sandbox, rather than leaking into the host environment.
Signed-off-by: Oliver Calder <oliver.calder@canonical.com>
The existing docs omitted or misstated several requirements when running
the gateway as a container with the Docker compute driver:
- OPENSHELL_GRPC_ENDPOINT is required; the Docker driver uses only the
scheme (http/https) — host and port are substituted automatically with
host.openshell.internal and the gateway's own bind port
- Supervisor binary must be extracted to a host path before starting the
gateway; bind-mount sources are resolved by the host Docker daemon so
the path must be identical inside and outside the gateway container
- Docker socket access requires adding the docker group (UID 1000 default)
- Port binding should remain 127.0.0.1; Docker driver adds a bridge
listener automatically
- add --server-san host.openshell.internal to generate-certs for mTLS
- Complete the mTLS docker run with all Docker driver requirements
- Add deploy/docker/gateway.toml — TOML config for the Docker driver
- Add deploy/docker/docker-compose.yml referencing the TOML
- Add docs/get-started/tutorials/docker-compose.mdx tutorial page
- Remote gateway registration instructions (--remote flag)
Address reviewer feedback:
- Move Docker Compose tutorials card to the bottom of the list
- Replace inline YAML snippet in Docker Compose section with a reference
to deploy/docker/ to avoid drift
- Clarify OPENSHELL_DB_URL is safe in compose.yml (plain SQLite path,
no credentials); the TOML block targets credential-bearing DSNs
- Note that ./ in source: resolves relative to the compose file directory
- Clarify that only the scheme from OPENSHELL_GRPC_ENDPOINT matters
- Add note that the tilde volume mount resolves to the same absolute
path on both host and container
* feat(providers): add Google Vertex AI provider
Adds Vertex AI provider profiles, routing, credential refresh plumbing, CLI support, docs, and regression coverage. Keeps the related NETLINK_ROUTE seccomp allowance needed by Vertex client tooling that calls getifaddrs.
* docs: add Vertex AI sandbox usage for Claude Code and OpenCode
Cover the full end-to-end setup for running Claude Code and OpenCode
inside an OpenShell sandbox via inference.local with a Vertex AI backend:
- google-vertex-ai.mdx: add 'Use from a Sandbox' section with tabbed
examples for Claude Code (--bare flag, no /v1 suffix) and OpenCode
(/v1 suffix required). Add providers_v2_enabled prerequisite and
--no-verify note for global region. Document policy proposals table
covering metadata.google.internal (always blocked), downloads.claude.ai,
and storage.googleapis.com.
- inference-routing.mdx: expand 'Use the Local Endpoint' section with
tabbed examples for Claude Code, OpenCode, Python OpenAI SDK, and
Python Anthropic SDK. Add notes explaining the /v1 path suffix
difference between clients.
- supported-agents.mdx: update Claude Code and OpenCode rows to mention
inference.local support and correct base URL requirements.
* fix: address vertex review findings
* test(sandbox): retry on spurious Ok in fork-exec ambiguity test
On arm64 under heavy CI load, the /proc fd scan in
find_socket_inode_owners can transiently miss the parent process's
socket fd entry, returning only the child as an owner. This causes
resolve_process_identity to return Ok (single owner, no ambiguity
check fires) instead of the expected ambiguous-ownership Err.
Extend the retry loop to also handle unexpected Ok results, mirroring
the existing retry for transient Err results. 10 retries at 50ms gives
a 500ms settling window, which is sufficient for procfs to stabilize
on loaded arm64 runners.
* fix: address vertex review regressions
* docs(router): clarify stream_response semantics for Vertex rawPredict routing
Document the three call sites of prepare_backend_request and their
stream_response values in a caller table:
- send_backend_request: false → :rawPredict (unary endpoint)
- send_backend_request_streaming: true → :streamRawPredict
- verify_backend_endpoint: explicitly false to probe the unary endpoint
Cross-reference the table from build_provider_url and
is_vertex_anthropic_rawpredict_route so the stream_response=true guard
in the suffix upgrade branch is understood in full context.
Also note that is_vertex_anthropic_rawpredict_route is a structural
predicate (model_in_path + anthropic_messages + :rawPredict suffix),
not a named-provider check, so any future provider with the same route
shape inherits the transforms automatically.
Add a dedicated page documenting how to run the OpenShell gateway as a
container using docker run, docker-compose, or podman, without the
system package manager installer.
This is useful for users on immutable OS distributions (Fedora CoreOS,
bootc-based images, Silverblue) where the standard install.sh path is
not appropriate, and for container-first environments.
Covers a quick-start with TLS disabled (localhost-bound), a full mTLS
setup using the gateway's generate-certs subcommand, a docker-compose
example, and a Podman variant.
Closes the gap raised in #1285.
Signed-off-by: Eric Curtin <eric.curtin@docker.com>
* docs(kubernetes): add initial reference docs
Add a Kubernetes reference section under docs/reference/kubernetes/ covering
gateway setup, OIDC and reverse-proxy authentication, cert-manager PKI, and
Gateway API external access. Update installation.mdx to link to the new setup
page and nest the new folder under the Reference section in index.yml.
* docs(kubernetes): promote Kubernetes to top-level nav and retitle pages
Move docs/reference/kubernetes/ to docs/kubernetes/ so the section appears
as a top-level nav entry rather than nested under Reference. Retitle the
pages to broader, action-oriented names (Managing Certificates, Ingress,
Access Control) and rename the files to match. Update cross-references in
setup.mdx and the installation page.
* docs(nav): revert Reference to folder mode
The Reference section was switched to explicit section/contents to host
the nested Kubernetes subsection. With Kubernetes promoted to a top-level
folder, restore folder-mode auto-detection. Side effect: sandbox-compute-drivers
is picked up again, which was silently absent from the explicit list.
* docs(kubernetes): add OpenShift install page
Mirror the OpenShift install instructions from PR #1240 into the
Kubernetes docs section. Documents the SCC binding for sandbox pods
and the chart overrides required by OpenShift's Security Context
Constraints (disabled PKI init Job, plaintext TLS, cleared fsGroup
and runAsUser).
* docs: refresh user-facing docs for recent sandbox and inference changes
- architecture: document system CA loading for upstream TLS, `tls: skip`
as the opt-out, gateway state persistence across restarts, and OCSF
structured logging surface.
- inference: document per-provider header allowlist, Authorization
stripping, 120s streaming idle tolerance, and extended-thinking
timeout guidance.
- manage-sandboxes: add "Execute a Command in a Sandbox" section for
`openshell sandbox exec` with flag reference.
- security best practices: expand seccomp denylist (unconditional and
conditional blocks), document two-phase Landlock probe, High-severity
`landlock-unavailable` finding, and inference keep-alive closure.
- observability logging: document port in HTTP log URLs, `[reason:...]`
denial suffixes, proxy 403/502 JSON error bodies, and Landlock
CONFIG:ENABLED/CONFIG:OTHER events.
Signed-off-by: Miyoung Choi <miyoungc@nvidia.com>
* docs(architecture): tighten high-level architecture page
Trim implementation detail (CA bundle paths, deprecated TLS keys, SSH
handshake secret) from the high-level architecture page, fix accuracy
issues surfaced during deep audit, and expand uncommon acronyms on first
mention.
- Drop unsupported "cost-based routing" claim from Privacy Router row.
- Replace "brokers requests across the platform" with auth-boundary
description.
- Add "inference" to Policy Engine constraint list per AGENTS.md.
- Expand Deny rule to include SSRF, blocked control-plane port, and L7
deny paths in addition to deny-by-default.
- Switch Allow/Deny labels from hyphen to colon; remove em dashes and a
double space.
- Expand LLM, SSRF, L7, TLS, CA, PEM, SSH, OCSF, and JSONL on first use.
Signed-off-by: Miyoung Choi <miyoungc@nvidia.com>
Made-with: Cursor
* docs: address audit feedback on refresh PR
- observability/logging: rewrite the allowed_ips paragraph after the
Denial Reasons table; the previous wording said authors "can use"
invalid entries while also stating they were rejected, which was
contradictory and conflated load-time validation with the runtime
per-CONNECT denial phrases the section documents.
- about/architecture: split compound sentences in the new Gateway
Lifecycle and Observability sections so each clause stands alone.
- inference/about: drop the streaming-tolerance sentence from the prose
paragraph since the dedicated Streaming reliability table row already
covers it.
Signed-off-by: Miyoung Choi <miyoungc@nvidia.com>
Made-with: Cursor
* docs(architecture): address review feedback on architecture page
- Drop the Privacy Router row from the Components table; it is not
yet a separately exposed customer-facing component.
- Update the page description and intro count to match the remaining
three components (gateway, sandbox, policy engine).
- Split the policy decision into the three modes that the engine
actually implements: Explicit Deny (deny rules and hardening rules,
takes precedence), Allow, and Implicit Deny (no rule matched).
Signed-off-by: Miyoung Choi <miyoungc@nvidia.com>
Made-with: Cursor
* docs(architecture): convert policy decisions to a table
Promote the three policy decisions (Explicit Deny, Allow, Implicit
Deny) to a top-level table with Decision, When it applies, and
Outcome columns instead of a nested bulleted list under list item 5.
Top-level tables render reliably across markdown renderers, where
nested-in-list tables do not.
Signed-off-by: Miyoung Choi <miyoungc@nvidia.com>
Made-with: Cursor
* docs: repharse a bit
---------
Signed-off-by: Miyoung Choi <miyoungc@nvidia.com>
* Added gateway and tutorial
* add new content and polish, add license terms
* small fix
* tiny fix
* polish bit more
* change warning to important
* fix support matrix
* small fix
* add a note about example script
* fix link
* fix link
* improve
* fix links
* unclutter sandbax index more
* improve titles
* tiny style fix
* style guide on inferences, release notes link to gh
* polish
* nit
* switch version to 0.0.3
* incorporate feedback
---------
Co-authored-by: Kirit93 <kthadaka@nvidia.com>
* initial doc filling
* improvements
* stage provided get started, and add clean tutorials
* pull in Kirit's content and polish
* improve observability
* moving pieces
* move TOC around
* drop support matrix from concepts
* fix links
* minor fixes and fix badges
* minor fixes
* incorporate missed content
* minor improvements
* clean up
* run dori style guide review
* clean up
* updates impacting docs
* incorporate feedback
* minor fix
* some edits
* enterprise structure
* update cards
* improve
* add some emojis
* improve landing page with animated getting started code
* fix the animated code
* small improvements
* refresh content based on PR 156 and 158
* README as the source of truth for quickstart
* update README
* run edits
* change to the new prod name, text only, code swipe later
* add Home
* incorporate dev feedback on README
* krit's edits
* edit and improve index pages
* fix build
* revert README
* revert quickstart to not pull from README