* fix(network): normalize Windows policy binary paths
Match Windows executable identities using a stable case-insensitive, separator-normalized representation across policy data, L4 input, and L7 relay evaluation. Preserve exact matching on other platforms and keep the original path for hashing and filesystem access.
NVBug 6782969
Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>
* fix(network): harden Windows binary matching
Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>
* fix(ci): scope Windows relay test imports
* fix(network): harden Windows binary path matching
---------
Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>
* refactor(inference): remove managed inference routes
Closes#3172
Remove the inference route control plane, inference.local data path, built-in router crate, and SDK surface. Move inference workloads to explicitly imported provider profiles and native endpoints, with migration cleanup and updated tests and documentation.
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
* fix(policy): preserve alternate upstream isolation
Restore the provider policy activation guard so legacy OpenAI and Anthropic providers configured for alternate base URLs do not grant egress to the built-in public vendor endpoints.
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
---------
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
Add a reminder to the bug report template's Logs field and a new row in
the security best-practices Common Mistakes table advising reporters to
redact credentials, API keys, and tokens from stack traces before pasting.
Signed-off-by: Evan Lezar <elezar@nvidia.com>
* feat(providers): support SPIFFE-backed token grants
Add provider profile token_grant metadata and expand endpoint-specific
dynamic credentials so sandbox supervisors can request SPIFFE JWT-SVIDs,
exchange them with an OAuth-style token endpoint, cache returned access
tokens, and inject bearer tokens into matching HTTP requests.
Wire Kubernetes and Helm deployments to mount the provider SPIFFE Workload
API socket into sandbox pods for token grant exchange.
Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
Signed-off-by: Gordon Sim <gsim@redhat.com>
* test(examples): add SPIFFE token grant demo
Add a reusable alpha/beta demo that deploys a SPIFFE-verifying token issuer
and protected services, imports a token-grant provider profile, creates a
sandbox, and verifies endpoint-specific bearer tokens.
The script leaves Kubernetes workloads in place, deletes sandboxes through
openshell unless KEEP_SANDBOX=1, and prints protected service logs as proof
of life.
Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
* fix(providers): harden SPIFFE token grants
* fix(providers): harden dynamic token grants
Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
* fix(providers): harden token grant handling
---------
Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
Signed-off-by: Gordon Sim <gsim@redhat.com>
* feat(providers): add Google Vertex AI provider
Adds Vertex AI provider profiles, routing, credential refresh plumbing, CLI support, docs, and regression coverage. Keeps the related NETLINK_ROUTE seccomp allowance needed by Vertex client tooling that calls getifaddrs.
* docs: add Vertex AI sandbox usage for Claude Code and OpenCode
Cover the full end-to-end setup for running Claude Code and OpenCode
inside an OpenShell sandbox via inference.local with a Vertex AI backend:
- google-vertex-ai.mdx: add 'Use from a Sandbox' section with tabbed
examples for Claude Code (--bare flag, no /v1 suffix) and OpenCode
(/v1 suffix required). Add providers_v2_enabled prerequisite and
--no-verify note for global region. Document policy proposals table
covering metadata.google.internal (always blocked), downloads.claude.ai,
and storage.googleapis.com.
- inference-routing.mdx: expand 'Use the Local Endpoint' section with
tabbed examples for Claude Code, OpenCode, Python OpenAI SDK, and
Python Anthropic SDK. Add notes explaining the /v1 path suffix
difference between clients.
- supported-agents.mdx: update Claude Code and OpenCode rows to mention
inference.local support and correct base URL requirements.
* fix: address vertex review findings
* test(sandbox): retry on spurious Ok in fork-exec ambiguity test
On arm64 under heavy CI load, the /proc fd scan in
find_socket_inode_owners can transiently miss the parent process's
socket fd entry, returning only the child as an owner. This causes
resolve_process_identity to return Ok (single owner, no ambiguity
check fires) instead of the expected ambiguous-ownership Err.
Extend the retry loop to also handle unexpected Ok results, mirroring
the existing retry for transient Err results. 10 retries at 50ms gives
a 500ms settling window, which is sufficient for procfs to stabilize
on loaded arm64 runners.
* fix: address vertex review regressions
* docs(router): clarify stream_response semantics for Vertex rawPredict routing
Document the three call sites of prepare_backend_request and their
stream_response values in a caller table:
- send_backend_request: false → :rawPredict (unary endpoint)
- send_backend_request_streaming: true → :streamRawPredict
- verify_backend_endpoint: explicitly false to probe the unary endpoint
Cross-reference the table from build_provider_url and
is_vertex_anthropic_rawpredict_route so the stream_response=true guard
in the suffix upgrade branch is understood in full context.
Also note that is_vertex_anthropic_rawpredict_route is a structural
predicate (model_in_path + anthropic_messages + :rawPredict suffix),
not a named-provider check, so any future provider with the same route
shape inherits the transforms automatically.
Sandbox processes running on Deno couldn't TLS-verify connections to
services signed by OpenShell's per-sandbox ephemeral CA because Deno
doesn't consult `SSL_CERT_FILE` / `NODE_EXTRA_CA_CERTS`. Deno reads its
own `DENO_CERT` env var (additive PEM bundle, same shape as
`NODE_EXTRA_CA_CERTS`) for that.
Add `DENO_CERT` to `child_env::tls_env_vars` pointing at the standalone
OpenShell CA cert (not the combined system+sandbox bundle, since DENO_CERT
is additive). The three callers in `process.rs` / `ssh.rs` iterate the
array, so the size bump from 5 to 6 entries flows through without
caller changes.
Updates the security best-practices doc to enumerate `DENO_CERT`
alongside the other trust-store env vars.
Tracks: brevdev/dex-ui#5
Signed-off-by: Alec Fong <alecf@nvidia.com>
Migrate all sandbox and VM driver network policy enforcement from
iptables to nftables. nftables provides atomic ruleset loading, a
cleaner rule syntax, and is the standard netfilter interface in modern
kernels.
Sandbox bypass enforcement (openshell-sandbox):
- Replace iptables chain of individual rule insertions with a single
atomic nftables ruleset load via nft -f
- New nft_ruleset module with pure functions for ruleset generation
and unit tests
- Combine log and reject rules in one inet family table (handles both
IPv4 and IPv6 in a single ruleset)
- Fall back to reject-only ruleset when kernel lacks nft_log support
- Enable net.netfilter.nf_log_all_netns so log rules work from
non-init network namespaces
- Use temp file for nft ruleset loading instead of stdin for
compatibility with minimal VM guest environments
VM TAP networking (openshell-driver-vm):
- Replace iptables NAT/forwarding rules with nftables equivalents
- New nft_ruleset module for TAP network rule generation with unit
tests
- Atomic table-per-TAP-device lifecycle (create/destroy)
- Host-side rules provide NAT infrastructure and defense-in-depth
isolation (input chain restricts VM to gateway port only, forward
chain blocks unsolicited inbound); primary security enforcement
happens inside the VM guest via the sandbox supervisor's own rules
VM init script:
- Load nft kernel modules at sandbox init
- Enable nf_log_all_netns sysctl for bypass detection logging
OCSF / docs:
- Update firewall rule engine references from iptables to nftables
- Document host firewall interaction model and two-layer enforcement
architecture in VM driver README and compute drivers reference
Closes#1335
Signed-off-by: Russell Bryant <rbryant@redhat.com>
* feat(gateway): add TOML configuration file (RFC 0003)
Introduces an opt-in --config / OPENSHELL_GATEWAY_CONFIG flag that loads a
TOML file with gateway-wide settings and per-driver tables. Source
precedence is CLI > env > file > built-in default, implemented via clap's
ValueSource so existing flags and env vars keep their priority.
Driver crates (kubernetes, docker, podman, vm) now derive Deserialize on
their config structs. SupervisorSideloadMethod gains Deserialize with
kebab-case rename. A per-driver inheritance allowlist on the loader side
overlays [openshell.gateway] shared defaults (default_image,
supervisor_image, image_pull_policy, guest_tls_*, ssh_handshake_skew_secs,
client_tls_secret_name, host_gateway_ip, enable_user_namespaces) onto
each [openshell.drivers.<name>] table before deserialization.
The Helm chart renders a new gateway-config ConfigMap and mounts it at
/etc/openshell/gateway.toml. The migrated OPENSHELL_* env entries are
dropped from the StatefulSet — only the Secret-backed
OPENSHELL_SSH_HANDSHAKE_SECRET remains. database_url stays on --db-url.
Adds examples/gateway/gateway.example.toml and updates architecture/gateway.md
with the source precedence and inheritance rules.
* docs(gateway): drop ssh_handshake_skew_secs and ssh_handshake_secret from examples
Both fields are scheduled for removal. Remove the example values and the
env-only note so the gateway.toml example and the architecture doc stop
recommending settings that will not exist much longer.
* docs(rfc): correct OPENSHELL_CONFIG to OPENSHELL_GATEWAY_CONFIG in RFC 0003
* docs(gateway): add per-driver TOML example configurations
Adds focused single-driver examples next to the comprehensive
gateway.example.toml: kubernetes, docker, podman, and microvm. Each one
demonstrates the realistic settings for that driver plus how shared
[openshell.gateway] defaults inherit into the driver table.
A new unit test (`checked_in_examples_parse`) loads every example through
the config_file loader so schema drift fails CI rather than silently
shipping a broken example.
* refactor(gateway): drop image_pull_policy from shared inheritance
Kubernetes and Podman use mutually-incompatible vocabularies for the same
TOML key:
- Kubernetes: `Always | IfNotPresent | Never` (free-form string passed
verbatim to the K8s API).
- Podman: `always | missing | never | newer` (strict lowercase enum
deserialised into `ImagePullPolicy`).
No value means the same thing in both drivers. Sharing the key at
`[openshell.gateway]` scope and inheriting it into every active driver's
table meant any value safe for one driver was either wrong or silently
dropped for the other (`IfNotPresent` → `ImagePullPolicy::Missing` after
`.unwrap_or_default()`). Operators run one driver per gateway, so the
"shared default" never pays for itself.
Make `image_pull_policy` driver-local:
- Remove the field from `GatewayFileSection` and from
`inheritable_keys()` for both Kubernetes and Podman.
- Drop the file→`RunArgs` merge for the gateway-scope key.
- Stop unconditionally clobbering the driver value with
`config.sandbox_image_pull_policy` in the runtime wiring — only apply
the CLI/env override when it was set (and, for Podman, only when it
parses into the lowercase enum).
- Move the key under `[openshell.drivers.kubernetes]` and
`[openshell.drivers.podman]` in every example, the RFC, the
architecture doc, and the Helm-rendered gateway ConfigMap.
The supervisor pull policy follows the same shape: it is K8s-only and
moves into `[openshell.drivers.kubernetes]` alongside `image_pull_policy`
in the Helm template.
* fix(gateway): address review feedback on TOML configuration
Resolves the P1 and P2 issues raised in PR #1317:
- Helm gateway ConfigMap moves `grpc_endpoint` under
`[openshell.drivers.kubernetes]` so the default install no longer fails
the gateway's `deny_unknown_fields` schema check.
- `kubernetes_config_from_file` and `podman_config_from_file` only let
the gateway-wide CLI/env `grpc_endpoint` overwrite the driver-table
value when it was actually supplied, preserving file-only configs.
- Kubernetes driver default `image_pull_policy` is now empty (was Podman
vocabulary "missing"), so default deployments let the Kubernetes API
apply its own policy instead of being rejected.
- New `disable_tls` gateway field plumbs `.Values.server.disableTls`
through the TOML ConfigMap instead of relying on env vars dropped from
the StatefulSet.
- StatefulSet pod template now carries a `checksum/gateway-config`
annotation so `helm upgrade` rolls pods when the ConfigMap changes.
- Auxiliary listener resolution preserves the full `SocketAddr` from
`health_bind_address` / `metrics_bind_address`, so a loopback-pinned
health port is not silently relocated onto the public bind address.
- `ssh_session_ttl_secs` from the file is now applied to `Config` (it
was previously accepted by the loader but never read).
New regression coverage: cli-level merge tests for the new fields plus
helm-unittest assertions for the ConfigMap shape, checksum annotation,
and `disable_tls` rendering.
* docs(gateway): consolidate gateway TOML examples into docs reference
Replaces the per-driver example files under examples/gateway/ with a
single published reference page at docs/reference/gateway-config.mdx
covering source precedence, layout, the full example, and the four
per-driver examples (Kubernetes, Docker, Podman, microVM). Drew flagged
during PR #1317 review that the examples belong with the user-facing
docs rather than in a sibling examples/ directory.
The cross-references in architecture/gateway.md and RFC 0003 are updated
to point at the new docs page; the round-trip test in config_file.rs is
removed (schema coverage stays on the inline parses_full_example test
and per-field merge tests — doc-snippet drift belongs in a separate
docs-lint, not in a cross-tree Rust unit test).
* refactor(core): move DEFAULT_K8S_NAMESPACE into K8s driver
The constant is Kubernetes-specific (used only by KubernetesComputeConfig's
Default impl) and does not belong in openshell-core. Relocate it to the
driver crate that owns the K8s vocabulary; openshell-core retains only
truly cross-cutting defaults.
* refactor(core): move Podman bridge default into Podman driver
DEFAULT_NETWORK_NAME is Podman vocabulary, consumed only by the Podman
driver. Also drops the unused DEFAULT_IMAGE_PULL_POLICY constant.
* docs(auth): scrub remaining SSH handshake secret references
Sweeps the trailing mentions left after the rebase: the gateway
config-file module doc, the Helm gateway-config ConfigMap header,
the gateway-config.mdx env-only note, and the RPM systemd unit
comment for init-gateway-env.sh.
* docs(gateway): clarify OPENSHELL_GRPC_ENDPOINT applies to all drivers
The previous comment implied the callback endpoint was Kubernetes-only,
but the value is propagated to every compute driver (Kubernetes, Docker,
Podman, VM) and must be reachable from wherever the sandbox runs.
* refactor(gateway): move driver options into config (#1394)
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* fix(e2e): regenerate gateway config via TOML for docker + podman harnesses
The gateway CLI flags moved into TOML config tables in 560550d2 (#1394),
which made every existing e2e/with-{docker,podman}-gateway.sh invocation
fail with "unexpected argument '--sandbox-namespace'" (and a long tail
of similar driver-specific options) before the gateway could even bind.
Replace the obsolete CLI flags with a synthesized
`[openshell.drivers.<driver>]` table written to `${STATE_DIR}/gateway.toml`
and passed via the new `--config` flag. Only the gateway-wide flags
that survived 560550d2 (bind-address, port, drivers, db-url, tls-*,
disable-tls, log-level, health-port) stay on the command line.
Both scripts get a small `toml_string` helper to properly TOML-quote
the values (the previous `%q` printf format produced bash-escape, not
TOML-escape). The Docker harness also corrects two field names that
diverged from the driver schema: `docker_network_name` →
`network_name`, and the supervisor binary/image plumbing now reads
through to `supervisor_bin` / `supervisor_image` in the same table.
The Podman harness drops `--ssh-gateway-port` (deleted in 560550d2 —
gRPC + SSH are multiplexed on the same port now) and substitutes
`network_name` + `gateway_port` for the obsolete `--sandbox-namespace`
(which the Podman driver never had as a typed field).
* fix(core): swap bind-only 0.0.0.0 SSH gateway host for cluster URL host
CLI's resolve_ssh_gateway treated 0.0.0.0 as a loopback "keep as-is"
when the cluster URL was also loopback, so the SSH proxy connected to
0.0.0.0:port. The unspecified address is never a valid connect target
and is not present in any TLS cert SAN, which produced BadCertificate
TLS handshake failures during `openshell sandbox create -- ...` in
docker/podman e2e (e.g. bypass_detection).
Resolution: when the server returns 0.0.0.0 or :: as the gateway host
and both endpoints are loopback, fall back to the cluster URL's host
(which the CLI is already using to reach the gateway, so it must
resolve and match the cert).
* fix(e2e): repair podman harness on macOS
Podman 5.x with the applehv/libkrun provider no longer creates the legacy
~/.local/share/containers/podman/machine/podman.sock symlink, and
`podman system service` is a Linux-only subcommand — the macOS client
delegates the API service to the VM. Both assumptions in the harness
were stale, so the script tried to start a temporary service that podman
rejected with "unknown flag: --time".
- Discover the macOS socket via `podman machine inspect` instead of the
hardcoded path.
- On Darwin, fail fast with a "start podman machine" message rather than
attempting the Linux-only `podman system service` fallback.
- Write socket_path into [openshell.drivers.podman] so the in-process
driver picks up the discovered socket; the driver reads TOML only
after the config refactor (560550d2), so OPENSHELL_PODMAN_SOCKET alone
was no longer enough.
* fix(server): use clone_from for TLS client CA assignment
clippy 1.95.0 rejects assigning the result of `Clone::clone()` to an
existing variable under `-D warnings` (`assigning_clones`). Switch to
`clone_from(&...)` to satisfy the lint and avoid the redundant
allocation.
---------
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
Add opt-in support for Kubernetes user namespace isolation on sandbox
pods. When enabled, container UID 0 maps to an unprivileged host UID
and capabilities become namespaced, providing defense-in-depth for the
supervisor process.
Configuration is two-layered: a cluster-wide default via
OPENSHELL_ENABLE_USER_NAMESPACES (default false) and a per-sandbox
override via the new `user_namespaces` field on SandboxTemplate.
When user namespaces are active, the pod security context is extended
with SETUID, SETGID, and DAC_READ_SEARCH capabilities to match the
bounding-set requirements inside a user namespace.
Introduces SandboxPodParams struct to replace long argument lists on
sandbox_to_k8s_spec and sandbox_template_to_k8s.
Validated end-to-end on OCP 4.22 (K8s 1.35.3, CRI-O 1.35, RHEL
CoreOS, kernel 5.14) with full SSH tunnel and non-identity UID mapping.
Closes#879
Resolve policy hostnames against the sandbox's /etc/hosts before DNS so Kubernetes hostAliases are visible to the proxy while keeping the existing allowed_ips SSRF enforcement semantics intact.
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
* docs: refresh user-facing docs for recent sandbox and inference changes
- architecture: document system CA loading for upstream TLS, `tls: skip`
as the opt-out, gateway state persistence across restarts, and OCSF
structured logging surface.
- inference: document per-provider header allowlist, Authorization
stripping, 120s streaming idle tolerance, and extended-thinking
timeout guidance.
- manage-sandboxes: add "Execute a Command in a Sandbox" section for
`openshell sandbox exec` with flag reference.
- security best practices: expand seccomp denylist (unconditional and
conditional blocks), document two-phase Landlock probe, High-severity
`landlock-unavailable` finding, and inference keep-alive closure.
- observability logging: document port in HTTP log URLs, `[reason:...]`
denial suffixes, proxy 403/502 JSON error bodies, and Landlock
CONFIG:ENABLED/CONFIG:OTHER events.
Signed-off-by: Miyoung Choi <miyoungc@nvidia.com>
* docs(architecture): tighten high-level architecture page
Trim implementation detail (CA bundle paths, deprecated TLS keys, SSH
handshake secret) from the high-level architecture page, fix accuracy
issues surfaced during deep audit, and expand uncommon acronyms on first
mention.
- Drop unsupported "cost-based routing" claim from Privacy Router row.
- Replace "brokers requests across the platform" with auth-boundary
description.
- Add "inference" to Policy Engine constraint list per AGENTS.md.
- Expand Deny rule to include SSRF, blocked control-plane port, and L7
deny paths in addition to deny-by-default.
- Switch Allow/Deny labels from hyphen to colon; remove em dashes and a
double space.
- Expand LLM, SSRF, L7, TLS, CA, PEM, SSH, OCSF, and JSONL on first use.
Signed-off-by: Miyoung Choi <miyoungc@nvidia.com>
Made-with: Cursor
* docs: address audit feedback on refresh PR
- observability/logging: rewrite the allowed_ips paragraph after the
Denial Reasons table; the previous wording said authors "can use"
invalid entries while also stating they were rejected, which was
contradictory and conflated load-time validation with the runtime
per-CONNECT denial phrases the section documents.
- about/architecture: split compound sentences in the new Gateway
Lifecycle and Observability sections so each clause stands alone.
- inference/about: drop the streaming-tolerance sentence from the prose
paragraph since the dedicated Streaming reliability table row already
covers it.
Signed-off-by: Miyoung Choi <miyoungc@nvidia.com>
Made-with: Cursor
* docs(architecture): address review feedback on architecture page
- Drop the Privacy Router row from the Components table; it is not
yet a separately exposed customer-facing component.
- Update the page description and intro count to match the remaining
three components (gateway, sandbox, policy engine).
- Split the policy decision into the three modes that the engine
actually implements: Explicit Deny (deny rules and hardening rules,
takes precedence), Allow, and Implicit Deny (no rule matched).
Signed-off-by: Miyoung Choi <miyoungc@nvidia.com>
Made-with: Cursor
* docs(architecture): convert policy decisions to a table
Promote the three policy decisions (Explicit Deny, Allow, Implicit
Deny) to a top-level table with Decision, When it applies, and
Outcome columns instead of a nested bulleted list under list item 5.
Top-level tables render reliably across markdown renderers, where
nested-in-list tables do not.
Signed-off-by: Miyoung Choi <miyoungc@nvidia.com>
Made-with: Cursor
* docs: repharse a bit
---------
Signed-off-by: Miyoung Choi <miyoungc@nvidia.com>
Policies with allowed_ips entries targeting loopback, link-local, or
unspecified ranges now fail at connection time instead of being silently
blocked at runtime. The shorthand log format for DENIED events includes
a [reason:...] suffix so operators can distinguish 'allowlist miss' from
'structurally un-allowable'. The mechanistic mapper skips proposals for
always-blocked destinations, preventing the infinite TUI notification
loop. The gateway validates proposed rules on approval as defense-in-depth.
- Extract shared IP helpers (is_always_blocked_ip, is_always_blocked_net,
is_internal_ip) to openshell_core::net
- Reject always-blocked entries in parse_allowed_ips with hard error
- Skip implicit allowed_ips synthesis for always-blocked literal IP hosts
- Add status_detail to HttpActivityBuilder for denial reason propagation
- Enrich NET and HTTP shorthand with [reason:...] for DENIED events
- Add engine: tag to HTTP shorthand (consistency with NET shorthand)
- Filter always-blocked proposals in mechanistic mapper generate_proposals
- Add validate_rule_not_always_blocked server-side defense-in-depth
- Update architecture docs, published docs, and E2E test assertions