* refactor(inference): remove managed inference routes
Closes#3172
Remove the inference route control plane, inference.local data path, built-in router crate, and SDK surface. Move inference workloads to explicitly imported provider profiles and native endpoints, with migration cleanup and updated tests and documentation.
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
* fix(policy): preserve alternate upstream isolation
Restore the provider policy activation guard so legacy OpenAI and Anthropic providers configured for alternate base URLs do not grant egress to the built-in public vendor endpoints.
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
---------
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
* chore(deps): replace ring with AWS-LC
Signed-off-by: Simon Scatton <sscatton@nvidia.com>
* fix(lint): address warnings after dependency upgrades
Signed-off-by: Simon Scatton <sscatton@nvidia.com>
* fix(tls): limit provider initialization to reqwest clients
Signed-off-by: Simon Scatton <sscatton@nvidia.com>
---------
Signed-off-by: Simon Scatton <sscatton@nvidia.com>
The Maturin-based wheel packaging was a historical remnant from when the local gateway launch path and OpenShell CLI were coupled in one binary. The gateway and CLI now ship as standalone artifacts, so the Python distribution should contain only the SDK.
Build a single platform-independent setuptools wheel, verify that it cannot contain native code or an openshell entry point, and simplify the release jobs and documentation for SDK-only PyPI installs.
Signed-off-by: Simon Scatton <sscatton@nvidia.com>
The supervisor binary runs inside sandbox images whose libc and glibc
version are unknown at build time, so it must be statically linked. Add
SUPERVISOR_LIBC to select between the default musl variant and a new
glibc-static variant that builds the GNU target with +crt-static.
glibc-static has no cross-compile path: zig cc accepts -static for
*-linux-gnu targets and emits a dynamically linked binary anyway. The
staging script therefore refuses a cross-arch request for that variant
rather than silently degrading linkage, and requires a native
per-architecture build.
Add verify-static-binary.sh, run after every supervisor build in both the
staging script and CI so linkage cannot regress unnoticed for either
variant. It inspects via readelf (or greadelf/llvm-readelf) and fails closed
rather than trusting the tool's exit status: every inspection must produce no
diagnostics, the input must be an executable ELF (ET_EXEC, or ET_DYN with
DF_1_PIE) whose PT_LOAD segments all lie within the file, whose dynamic table
agrees with PT_DYNAMIC, and which carries no PT_INTERP and no DT_NEEDED. That
rejects a dynamically linked, truncated, corrupt, non-ELF, or shared-object
input that naive parsing would misread as static. Hosts without any inspector
(e.g. macOS, which ships no binutils) skip with a warning; Linux, including
CI, requires one and fails closed.
No image or release workflow builds the glibc-static variant, so add a
dedicated supervisor-static-validate workflow that builds it on both
architectures and runs the verifier. rust-native-build.yml uses self-hosted
runners, which reject pull_request-triggered jobs, so it validates in the merge
queue and on pushes to main that touch the build inputs, plus a nightly
schedule, so the GNU + crt-static build branch cannot regress unnoticed.
The default is unchanged, so image, release, and CI behavior is identical.
Selecting glibc-static statically links LGPL glibc into a redistributed
binary, which is why it is opt-in.
Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
Co-authored-by: Mrunal Patel <mrunalp@gmail.com>
The OpenClaw community sandbox was removed from NVIDIA/OpenShell-Community
in PR #73 (May 16, 2026). The Docker Compose tutorial and docker-compose.yml
comment still referenced the stale --from openclaw command and GHCR image.
Replace the broken OpenClaw tab with a redirect to the NemoClaw Quickstart,
which is the supported path per docs/about/supported-agents.mdx. Remove the
stale pre-pull command for the removed image.
Fixes#2404
Signed-off-by: Matias Schimuneck <schimuneck.matias@gmail.com>
* feat(kubernetes): add sidecar supervisor topology
Add the Kubernetes sidecar supervisor topology, its Helm/Skaffold configuration, topology documentation, and sidecar e2e matrix coverage. Skip root-only sandbox identity rewriting when process enforcement is network-only so the low-permission sidecar process container can start successfully.
Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
* fix(supervisor): avoid similar process id names
Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
* fix(supervisor): avoid similar process id names
Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
* fix(sandbox): avoid similar proxy id names
Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
* docs(kubernetes): clarify sidecar topology limits
Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
* fix(kubernetes): keep sidecar process leaf capless
Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
* fix(kubernetes): refresh sidecar provider env snapshots
Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
* test(supervisor): align hot-swap identity regression
Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
* fix(kubernetes): stage sidecar mtls files before proxy chown
Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
* fix(kubernetes): simplify sidecar supervisor topology
Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
* chore(helm): reuse sidecar skaffold values
Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
* fix(supervisor): avoid similar iptables helper names
Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
* fix(e2e): harden kube gateway wrapper setup
Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
* fix(supervisor): avoid nft batch rollback on OCP
Run nftables setup as individual commands so optional conntrack and log expressions can fail without rolling back required table, chain, and reject rules.
Signed-off-by: Seth Jennings <sjenning@redhat.com>
Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
* fix(kubernetes): preserve process identity in sidecar topology
Render sidecar pods with a shared process namespace, keep binary-aware network policy enabled, and move Kubernetes sidecar settings under the nested sidecar config table.
Also apply unprivileged Landlock/seccomp setup in NetworkOnly supervisor mode so sidecar topology keeps sandbox child hardening without privileged process setup.
Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
* refactor(kubernetes): replace sidecar snapshots with control socket
Coordinate sidecar policy and provider bootstrap over a local Unix socket so the process leaf no longer reads policy/provider snapshot files.
Report entrypoint startup through the control channel and keep gateway credentials confined to the network sidecar.
Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
* feat(kubernetes): support relaxed sidecar network identity
Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
* fix(sandbox): satisfy sidecar clippy lint
Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
* refactor(kubernetes): standardize topology naming
Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
* fix(sandbox): satisfy linux clippy timeout import
Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
* fix(kubernetes): support kata sidecar on ipv4 pods
Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
* fix(kubernetes): satisfy linux clippy for sidecar fallback
Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
* chore(kubernetes): remove stale supervisor topology references
Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
* fix(kubernetes): enable sidecar binary policy inspection
Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
* fix(kubernetes): harden sidecar control boundary
Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
* fix(kubernetes): couple sidecar supervisor lifecycles
Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
---------
Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
Signed-off-by: Seth Jennings <sjenning@redhat.com>
Co-authored-by: Seth Jennings <sjenning@redhat.com>
The existing docs omitted or misstated several requirements when running
the gateway as a container with the Docker compute driver:
- OPENSHELL_GRPC_ENDPOINT is required; the Docker driver uses only the
scheme (http/https) — host and port are substituted automatically with
host.openshell.internal and the gateway's own bind port
- Supervisor binary must be extracted to a host path before starting the
gateway; bind-mount sources are resolved by the host Docker daemon so
the path must be identical inside and outside the gateway container
- Docker socket access requires adding the docker group (UID 1000 default)
- Port binding should remain 127.0.0.1; Docker driver adds a bridge
listener automatically
- add --server-san host.openshell.internal to generate-certs for mTLS
- Complete the mTLS docker run with all Docker driver requirements
- Add deploy/docker/gateway.toml — TOML config for the Docker driver
- Add deploy/docker/docker-compose.yml referencing the TOML
- Add docs/get-started/tutorials/docker-compose.mdx tutorial page
- Remote gateway registration instructions (--remote flag)
Address reviewer feedback:
- Move Docker Compose tutorials card to the bottom of the list
- Replace inline YAML snippet in Docker Compose section with a reference
to deploy/docker/ to avoid drift
- Clarify OPENSHELL_DB_URL is safe in compose.yml (plain SQLite path,
no credentials); the TOML block targets credential-bearing DSNs
- Note that ./ in source: resolves relative to the compose file directory
- Clarify that only the scheme from OPENSHELL_GRPC_ENDPOINT matters
- Add note that the tilde volume mount resolves to the same absolute
path on both host and container
Refresh the CI image tool pins so Go-built tools are rebuilt with patched Go releases and move the sandbox Python runtime to 3.14.5.
Rebase the gateway runtime to a pinned distroless Debian 13 image with glibc 2.41-12+deb13u3 while preserving the existing UID/GID 1000 runtime identity for upgrade compatibility. Update rustls-webpki to 0.103.13 and clarify Linux k3d guidance now that k3d is not installed through mise on Linux.
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
Adds a helm:test mise task that installs the helm-unittest plugin if
not present and runs chart unit tests under deploy/helm/openshell.
Installs the plugin into Dockerfile.ci so CI runs do not need to
download it each time.
Closes#1281
Signed-off-by: Mesut Oezdil <versusfinem@gmail.com>
* wip
* refactor(docker): use native rust builds for split gateway/supervisor images
Drop the in-Docker BUILD_FROM_SOURCE path so both images consume only
prebuilt binaries staged natively via tasks/scripts/stage-prebuilt-binaries.sh.
This mirrors what CI does and reuses the host's cargo target cache and
sccache across rebuilds.
- Dockerfile.gateway: nvcr.io/nvidia/distroless/cc:v4.0.4 base (the 4.0.0
tag does not exist on nvcr.io; the registry uses a v prefix). GNU-linked
binary copied to /usr/local/bin.
- Dockerfile.supervisor: scratch base, static musl binary. Static linkage
lets the image stay scratch while still being executable as a Kubernetes
init container.
- skaffold.yaml: each artifact invokes tasks/scripts/docker-build-image.sh,
which stages the binary natively (cargo / cargo-zigbuild) and then builds
the image. Drops the cross-build.sh dependency from the supervisor build.
- seccomp.rs: add a local SYS_kexec_file_load constant for musl/aarch64.
libc 0.2.185 omits the symbol from its musl/aarch64 bindings, so the
supervisor's seccomp filter previously failed to compile for that target.
- architecture/build.md: describe the native-first pipeline and per-image
runtime choices.
Local validation: gateway image 101MB (was 194MB), supervisor image 21.7MB.
helm:skaffold:run deploys cleanly; the static musl supervisor binary runs
correctly in a non-glibc agent container.
* refactor(docker): tighten binary perms via --chown + 0550
Replace `COPY --chmod=755` with `COPY --chown=<user> --chmod=0550` in
the gateway and supervisor Dockerfiles. The binary is no longer
world-readable or world-executable; ownership is pinned to the runtime
user.
- Gateway uses `--chown=nvs:nvs` + `USER nvs:nvs`, matching the only
non-root user defined in `nvcr.io/nvidia/distroless/cc` (UID 1000) and
the Helm chart's `securityContext.runAsUser: 1000`, which overrides
the Dockerfile USER at runtime.
- Supervisor uses numeric `--chown=65534:65534` because the scratch base
has no `/etc/passwd` for name resolution. The supervisor image is
only consumed by the init-container copy-self path; the destination
pod's runAsUser governs execute access.
Validated by deploying to a local k3d cluster via `helm:skaffold:run`
and confirming the gateway StatefulSet reaches 1/1 Running.
* ci(gpu): repoint GPU probe image lookup at Dockerfile.gateway
The previous awk parsed `FROM <image> AS gateway` from the now-deleted
`Dockerfile.images`. The new `Dockerfile.gateway` uses an ARG with a
default (`ARG GATEWAY_BASE_IMAGE=nvcr.io/nvidia/distroless/cc:v4.0.4`)
and `FROM ${GATEWAY_BASE_IMAGE} AS gateway`, so the old script returns
nothing.
Parse the ARG default value directly so the GPU prerequisites check
keeps using the gateway base image as a `nvidia-smi` probe target.
* ci(gpu): pin GPU probe to nvcr.io/nvidia/base/ubuntu:noble
The previous probe parsed the gateway base image out of the Dockerfile,
relying on the fact that the gateway ran on `nvcr.io/nvidia/base/ubuntu`
and that NVIDIA Container Toolkit CDI injection would populate
`nvidia-smi` and the supporting libs at runtime. The new gateway base
(`nvcr.io/nvidia/distroless/cc`) lacks `ldconfig`, a populated
`/usr/bin`, and the broader filesystem layout CDI injection assumes,
so it cannot serve as a GPU probe.
Pin the probe image explicitly to the NVIDIA-managed Ubuntu base. The
probe is independent of the gateway runtime and survives future base
swaps.
* fix(e2e-gpu): pass GPU probe image via env, drop Dockerfile.images parse
The Rust e2e test `gpu_request_for_each_discovered_device_matches_plain_container`
was still parsing the deleted `Dockerfile.images` to derive its GPU probe
image, panicking with `No such file or directory` after this PR's
Dockerfile split.
Move the probe image to a single source of truth in the workflow
(`OPENSHELL_E2E_GPU_PROBE_IMAGE` env at the job level) and require the
e2e test to read it from there. No silent codebase default — the test
panics with a pointer at the workflow if the env is missing, so it
fails loudly rather than drifting from CI.
The prereq probe step in the workflow now consumes the same env, so
the probe image is declared exactly once.
* fix(docker): keep supervisor binary root-owned for rootless Podman
The Podman driver mounts the supervisor image read-only into the sandbox
container at /opt/openshell/bin and runs that container as UID 0, but
deliberately drops DAC_OVERRIDE for hardening (container.rs:419). With
--chown=65534:65534 --chmod=0550 the binary was r-xr-x--- owned by UID
65534, so the container's UID 0 fell into "other" with no read or exec
access and the supervisor crashed on start (ContainerExited code 1).
Docker and Kubernetes both retain DAC_OVERRIDE, so root could still
exec the file — which is why this regression only surfaced in the
Podman e2e job.
Drop --chown from the supervisor COPY so the binary stays root-owned.
Keep --chmod=0550: the security win was dropping world-execute, not
changing the owner. The chown bought nothing here because the container
is always UID 0 regardless of driver, but it actively broke the only
driver that drops DAC_OVERRIDE.
---------
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
* fix(sandbox): add copy-self subcommand for scratch-image init container
The supervisor image is built FROM scratch and contains only the
openshell-sandbox binary. The Kubernetes side-load init container
previously invoked `sh -c "cp ..."`, which failed with
`executable file 'sh' not found in $PATH` because the scratch image
has no shell.
Add a `copy-self <DEST>` subcommand to the supervisor binary that
resolves its own executable, copies it to the destination path, and
sets mode 0755. Update the Kubernetes driver init container to invoke
this subcommand directly. Mirrors argoexec's emissary self-copy pattern.
* fix(docker): use ubuntu base for supervisor image
`FROM scratch` cannot exec the dynamically linked supervisor binary
(no `/lib/ld-linux-*.so.*`), so the Kubernetes side-load init container
hit `exec /openshell-sandbox: no such file or directory` even after
the `copy-self` subcommand replaced the broken `sh -c` command.
Switch the supervisor stage to `nvcr.io/nvidia/base/ubuntu:noble-20251013`
— the same base the gateway image already uses — so glibc and the
dynamic loader are present. Docker/Podman drivers continue to work
unchanged (they extract or volume-mount the binary into agent containers
that already have glibc); the Kubernetes init container now succeeds
end-to-end.
Verified on a k3d cluster: init container `openshell-supervisor-install`
exits 0 on first try, binary lands at `/opt/openshell/bin/openshell-sandbox`
(mode 0755), and the agent container reaches Ready.
* docs(docker): refresh supervisor target description
The file header still described the supervisor image as
`FROM scratch, binary only`. Update it to match the current
Ubuntu-base supervisor stage.
The openshell-providers crate uses `include_str!("../../../providers/*.yaml")`
to embed provider profile YAML at compile time. The rust-builder stage in
Dockerfile.images only copied `crates/` and `proto/`, so source builds
(BUILD_FROM_SOURCE=1, used by Skaffold dev) failed with:
error: couldn't read `crates/openshell-providers/src/../../../providers/anthropic.yaml`
Copy `providers/` into the build context alongside the other source trees
so `include_str!` resolves.
* feat(auth): add OIDC/Keycloak authentication with RBAC
Add OAuth2/OIDC authentication to the gateway server with role-based
access control, CLI login flows, and full deployment plumbing.
Server: JWT validation against configurable OIDC issuer (oidc.rs),
JWKS key caching with TTL and rotation handling, method classification
(unauthenticated/sandbox-secret/dual-auth/bearer), identity extraction
with provider-agnostic Identity type, and RBAC enforcement via
AuthzPolicy with configurable admin/user roles and auth-only mode.
CLI: browser-based Authorization Code + PKCE flow, Client Credentials
flow for CI/automation, token storage with refresh, gateway add/login/
logout commands, OIDC bearer token injection over mTLS transport,
discovery endpoint for auto-configuration.
Security: sandbox-secret scope restriction on UpdateConfig (policy
sync only), anti-spoofing header stripping, dual-auth fallthrough
from sandbox-secret to Bearer token.
Deployment: OIDC config wired through DeployOptions, Docker env vars,
Helm values/templates, HelmChart manifest, cluster-entrypoint.sh, and
bootstrap scripts. Keycloak dev server script with pre-configured
realm (test users, roles, PKCE client, CI client).
Tested with Keycloak. The roles claim path and role names are
configurable to support other OIDC providers.
* feat(auth): add OAuth2 scope-based fine-grained permissions
Add opt-in scope enforcement on top of existing OIDC role-based access
control. When --oidc-scopes-claim is set, the server extracts scopes
from the JWT and checks them per-method against an exhaustive scope map.
Scopes: sandbox:read, sandbox:write, provider:read, provider:write,
config:read, config:write, inference:read, inference:write, and
openshell:all (wildcard). Methods not in the scope map require
openshell:all. Scopes layer on top of roles and cannot escalate
privilege. Auth-only mode (empty role names) still enforces scopes
when enabled.
Server: scopes_claim in OidcConfig, scope extraction from JWT
(space-delimited and JSON array formats), standard OIDC scope
filtering, scope check in AuthzPolicy after role check.
CLI: --oidc-scopes on gateway add/start stored in metadata and
consumed by gateway login, --oidc-scopes-claim on gateway start
forwarded to server, scopes parameter in browser and client
credentials OAuth2 flows with openid deduplication.
Deployment: oidc_scopes_claim wired through DeployOptions, docker.rs,
Helm, bootstrap scripts, and cluster entrypoint.
Keycloak: realm config updated with built-in OIDC scopes and 9
OpenShell client scopes as optional on openshell-cli and openshell:all
as default on openshell-ci.
* fix(auth): address branch review findings
Add GetInferenceBundle to sandbox-secret methods so sandbox inference
route refresh works under OIDC. Make GetSandboxConfig dual-auth so CLI
users can read sandbox settings with Bearer tokens.
Preserve OIDC gateway metadata on restart — a bare gateway start
without --oidc-* flags no longer erases the stored OIDC registration.
Document CI client ID requirement (openshell-ci vs openshell-cli) in
the testing guide. Add security note about auth-only mode blast radius
for GitHub Actions.
* fix(auth): complete review findings for OIDC auth boundary
Move OpenShell/GetSandboxConfig from sandbox-secret-only to dual-auth
so CLI users can read sandbox settings with Bearer tokens while sandbox
supervisors continue using the shared secret.
Add sandbox secret interceptor to the inference bundle fetch path so
GetInferenceBundle works under OIDC-enabled gateways. Extract shared
interceptor constructor to avoid duplication.
Add GetSandboxConfig to the config:read scope map so scope enforcement
applies consistently when scopes are enabled.
Refactor OIDC metadata preservation into apply_oidc_gateway_metadata()
with explicit resume semantics — only preserve existing OIDC metadata
on real resume paths, not on fresh deployments.
Update architecture docs and testing guide to reflect the corrected
method classifications and add new test coverage for interceptor
injection, scope requirements, metadata preservation, and dual-auth
classification.
* refactor(auth): use oauth2 crate for CLI OIDC flows
Replace hand-written PKCE generation, authorization URL construction,
token exchange, client credentials, and token refresh with the oauth2
crate's typed API.
Eliminates sha2, hex, and getrandom dependencies from the CLI. The
custom urlencoded() helper and manual form POST logic are replaced by
BasicClient methods with proper type-state safety.
Discovery and the callback server remain custom since the oauth2 crate
does not provide OIDC discovery or a localhost redirect listener.
* refactor(auth): move server auth modules into auth/ directory
Group oidc.rs, authz.rs, identity.rs, and the auth HTTP endpoints
under src/auth/ module directory. No behavioral changes.
auth/mod.rs — module root, re-exports HTTP router
auth/oidc.rs — JWT validation, JWKS caching, method classification
auth/authz.rs — role and scope authorization policy
auth/identity.rs — provider-agnostic Identity type
auth/http.rs — /auth/connect and /auth/oidc-config endpoints
* fix(auth): use RequestBody auth type for client credentials flow
The oauth2 crate defaults to BasicAuth (HTTP Basic header) but Keycloak
and most OIDC providers expect client_secret_post (credentials in the
request body). Set AuthType::RequestBody explicitly to match the
pre-refactor behavior.
Also re-export Identity, IdentityProvider, and JwksCache from the auth
module so ServerState's public API remains nameable by external consumers.
* fix(auth): forward OPENSHELL_OIDC_SCOPES through cluster bootstrap
Pass --oidc-scopes to gateway start so the metadata includes requested
scopes after cluster bootstrap. Without this, users had to manually
edit metadata.json to set scopes for gateway login.
Usage: OPENSHELL_OIDC_SCOPES="openshell:all" mise run cluster
* test(auth): add OIDC e2e tests for RBAC, scopes, and client credentials
Add 10 end-to-end tests covering OIDC authentication against a live
K3s cluster with Keycloak:
RBAC (5 tests): admin can create providers, user cannot, user can list
sandboxes, unauthenticated requests rejected, health probe works
without auth.
Scopes (4 tests): sandbox-scoped token can list sandboxes but not
providers, openshell:all grants full access, no-scopes token denied.
Client credentials (1 test): CI token via client_credentials grant.
Tests are opt-in via OPENSHELL_E2E_OIDC=1 and OPENSHELL_E2E_OIDC_SCOPES=1
env vars. They derive the Keycloak URL from gateway metadata to match
the server's configured issuer.
Run with:
OPENSHELL_E2E_OIDC=1 OPENSHELL_E2E_OIDC_SCOPES=1 \
PYTHONPATH=python uv run pytest e2e/python/oidc/ -v
* fix(docs): fix markdown lint errors in OIDC architecture docs
Add blank lines before lists and fenced code blocks to satisfy
markdownlint MD031 and MD032 rules.
* ci(docker): use prebuilt Rust binaries by default
Flip Docker image builds to consume staged native Rust artifacts, remove in-Docker Rust build stages, and publish per-arch images with a manifest merge.
Add local staging support for prebuilt gateway and sandbox binaries so development image builds continue to work without CI artifacts.
Signed-off-by: Jonas Toelke <jtoelke@nvidia.com>
* ci(docker): address prebuilt build review feedback
* ci(rust): allow existing vfio complexity
* ci(rust): pin toolchain to 1.95
---------
Signed-off-by: Jonas Toelke <jtoelke@nvidia.com>
* feat(docker): add BINARY_SOURCE selector for prebuilt Rust binaries
Signed-off-by: Jonas Toelke <jtoelke@nvidia.com>
* fix(docker): preserve exec bit on prebuilt binary COPY
Adds --chmod=755 to the COPY instructions in the scratch-based
prebuilt binary stages. Without this, binaries produced by PR 4a and
shuttled through actions/upload-artifact + download-artifact lose
their executable bit during the roundtrip, and the resulting image's
ENTRYPOINT fails at runtime.
Signed-off-by: Jonas Toelke <jtoelke@nvidia.com>
---------
Signed-off-by: Jonas Toelke <jtoelke@nvidia.com>
* feat(podman): add Podman compute driver for rootless sandbox management
Adds openshell-driver-podman, a new compute driver that manages OpenShell
sandboxes as rootless Podman containers via the Podman REST API over a
Unix socket. Enables local workstation sandboxes without Kubernetes.
Driver features:
- Bridge networking with ephemeral host-port mapping for rootless SSH reachability
- Named volumes for workspace storage, Podman native health checks, GPU via CDI
- Supervisor binary sideloaded via image volume mount (BYOC-compatible)
- SSH handshake secret injected via Podman secrets API (not plaintext env)
- Typed ContainerSpec structs, input validation, and path-traversal guards
- Cgroups v2 required; fails fast on v1 hosts
- Bounded event stream buffer; watch stream reconnection handled by server watch_loop
- Graceful shutdown and standalone driver binary with gRPC bridge
Rootless-specific fixes:
- Skip drop_privileges when user namespace lacks SETUID/SETGID/DAC_READ_SEARCH caps
- Add /run/netns tmpfs mount for ip netns in rootless containers
- Use secret_env map (not secrets array) for env-var injection in libpod API
- Resolve SSH endpoint to 127.0.0.1:<host_port> instead of unreachable bridge IP
Server/sandbox hardening:
- Split loopback and link-local SSRF gates; Podman/VM drivers allow loopback
- Close SSRF bypass in SSH tunnel Host path by resolving DNS before connecting
- Prevent OPENSHELL_* env var override by user-supplied spec environment maps
- Disable SQLite pool idle_timeout/max_lifetime for in-memory databases
- Emit deleted_event on 404-during-inspect instead of regressing sandbox phase
- Key delete cleanup by stable sandbox_id to survive container label drift
CLI fixes:
- Restore --name as a named flag on sandbox create (not positional)
- Fix exec command arg parsing to not consume sandboxed-command flags
- Propagate SSH verbosity via OPENSHELL_SSH_LOG_LEVEL
Build tooling:
- Add tasks/scripts/container-engine.sh: auto-detects Podman or Docker, exposes
unified ce_* helpers; all build/cluster/VM scripts updated to use it
- Add docker:build:supervisor mise task for standalone supervisor image
- Add openshell-driver-podman to Dockerfile.images pre-fetch/build stages
- Add e2e/rust/e2e-podman.sh and e2e:podman mise task for full lifecycle testing
Signed-off-by: Adam Miller <admiller@redhat.com>
* fix(driver-podman): derive grpc endpoint from server bind port
When a user starts the gateway on a non-default port (e.g. --port 8081),
sandbox containers were receiving OPENSHELL_ENDPOINT pointing at the
default port 8080. The driver's auto-detection fallback read
OPENSHELL_BIND_ADDRESS from the environment, which was stale or unset,
and fell back to DEFAULT_SERVER_PORT.
Add gateway_port to PodmanComputeConfig and thread config.bind_address.port()
from the server into the driver so the fallback uses the actual listening
port. Remove the OPENSHELL_BIND_ADDRESS env var read and the
extract_port_from_bind_address helper which are no longer needed.
Add --gateway-port / OPENSHELL_GATEWAY_PORT to the standalone driver
binary for parity when the driver is run outside the embedded server path.
Signed-off-by: Adam Miller <admiller@redhat.com>
* fix(driver-podman): address PR feedback on env test safety and cluster DNS docs
Replace hand-rolled unsafe TempEnvVar RAII guard with temp_env::with_vars
and a static ENV_LOCK mutex, fixing a data race in parallel test execution.
The prior safety comment incorrectly claimed Cargo runs tests single-threaded.
Update debug-openshell-cluster skill to accurately document the DNS proxy
strategy (setup_dns_proxy + public DNS fallback) and clarify the separation
between cluster DNS and sandbox agent DNS enforcement.
Signed-off-by: Adam Miller <admiller@redhat.com>
* fix(e2e): resolve CI failures in auth timeout, test harness, and formatting
- Short-circuit browser_auth_flow when OPENSHELL_NO_BROWSER=1 instead
of waiting the full 120s AUTH_TIMEOUT for a callback that never arrives
- Add timeout to SandboxGuard::create() and create_with_upload() to
prevent indefinite hangs (matches create_keep() which already had one)
- Add missing '--' separator in no_proxy test before command args
- Add #![cfg(feature = "e2e")] gate to sandbox_lifecycle.rs
- Run cargo fmt on openshell-driver-podman
- Refine cluster DNS docs for Podman in debug-openshell-cluster skill
Signed-off-by: Adam Miller <admiller@redhat.com>
* refactor(server): remove allows_loopback_endpoints from ComputeRuntime
SSRF protection is now handled at the network and proxy layers
(openshell-core net.rs, openshell-sandbox proxy.rs) rather than
requiring per-driver flags on ComputeRuntime. Update architecture
docs to reflect supervisor relay SSH transport and add rootless
networking deep-dive.
Signed-off-by: Adam Miller <admiller@redhat.com>
---------
Signed-off-by: Adam Miller <admiller@redhat.com>
Move /health, /healthz, and /readyz to a dedicated plaintext HTTP port
(default 8081) so Kubernetes probes work without mTLS client certificates.
- Add health_bind_address to Config with --health-port CLI arg
- Spawn standalone axum::serve for health_router on the health port
- Remove health routes from the main multiplexed HTTP router
- Update Helm statefulset probes from tcpSocket to httpGet on health port
- Fix cluster-healthcheck.sh to open/close TCP without sending data,
avoiding InvalidContentType TLS errors in the gateway log