32 Commits
Author SHA1 Message Date
Mark CampbellandKris Hicks 679b190677 feat(testing): support independent gateway and supervisor image overrides (#3341)
* feat(testing): normalize configurable test images

Signed-off-by: Bobbins228 <mcampbel@redhat.com>
Signed-off-by: Kris Hicks <khicks@nvidia.com>

* feat(helm): add global image overrides

Signed-off-by: Bobbins228 <mcampbel@redhat.com>

* feat(helm): support image registry overrides

Signed-off-by: Bobbins228 <mcampbel@redhat.com>
Signed-off-by: Kris Hicks <khicks@nvidia.com>

* refactor(helm): simplify image configuration

Signed-off-by: Bobbins228 <mcampbel@redhat.com>

* fix(e2e): avoid reloading reused kind sandbox image

Signed-off-by: Bobbins228 <mcampbel@redhat.com>

* fix(helm): default sandbox image to nvcr.io/nvidia/base/ubuntu:24.04

Signed-off-by: Kris Hicks <khicks@nvidia.com>

---------

Signed-off-by: Bobbins228 <mcampbel@redhat.com>
Signed-off-by: Kris Hicks <khicks@nvidia.com>
Co-authored-by: Kris Hicks <khicks@nvidia.com>
2026-09-23 23:35:30 +00:00
50230616d5 refactor(runtime): retire Community image dependencies (#3386)
* feat(sandbox): default to official Alpine sandbox image

default_sandbox_image() now returns docker.io/library/alpine:3.22, a generic
version-qualified official image, so a fresh install no longer depends on the
community sandbox image catalog. All compute drivers (docker, podman,
kubernetes, vm) inherit this fallback.

Part of #3116.

Signed-off-by: Akram
Signed-off-by: Akram <akram.benaissi@gmail.com>

* feat(deploy): default deployment configs to the official Alpine sandbox image

Update the shared gateway default_image, Helm chart values, the standalone
Kubernetes manifest, and the dev gateway task scripts to use
docker.io/library/alpine:3.22 instead of the community base image, consistent
with default_sandbox_image(). GPU e2e image-build base is left unchanged (CUDA
needs a glibc base).

Part of #3116.

Signed-off-by: Akram
Signed-off-by: Akram <akram.benaissi@gmail.com>

* feat(driver): default to numeric non-root identity for USER-less images

With the default sandbox image now Alpine, images that declare no OCI USER
must start instead of being rejected. When the image declares no USER and
the policy requests none, the Podman and Docker drivers now supply a numeric
non-root identity (DEFAULT_SANDBOX_UID/GID = 1000) instead of rejecting,
matching the numeric-identity behavior of the Kubernetes and VM drivers. The
supervisor's resolved-identity path runs the sandbox as a synthesized
non-root account without the account existing in the image. Images that
declare a USER keep the OCI resolution path unchanged.

Part of #3116.

Signed-off-by: Akram <akram.benaissi@gmail.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(conformance): use Alpine workload image

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* refactor(policy): drop community image /app path from default policy

The restrictive default policy granted read-only access to /app, a directory
that only existed in the community base image. A generic Alpine default has no
/app, so remove it. Landlock best-effort already ignores absent paths; this
just stops advertising a community-specific layout in the default.

Part of #3116.

Signed-off-by: Akram
Signed-off-by: Akram <akram.benaissi@gmail.com>

* docs(config): document Alpine default images

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(podman): report early sandbox termination

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(podman): initialize rootless workspace ownership

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(sandbox): qualify NVIDIA Ubuntu default

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): initialize rootful default workspace

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(sftp): add native sandbox adapter

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sftp): gate runtime helper support to Linux

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sftp): support standard OpenSSH file operations

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sftp): harden rename and special file handling

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(runtime): remove community image dependencies

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(e2e): build provider readiness tool fixture

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(e2e): use a dedicated Noble fixture for Docker tests

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Akram
Signed-off-by: Akram <akram.benaissi@gmail.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Evan Lezar <elezar@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
2026-09-22 14:43:51 +02:00
0a770d9173 feat(kubernetes): support HA gateway rebalancing (#1868)
* feat(kubernetes): support HA gateway rebalancing

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* perf(server): cache peer connections, tokens, and owner lookups

Every forwarded relay rebuilt its setup from scratch: an owner lookup, a
blocking read of the peer token, a TLS connect to the owning replica, and
a TokenReview plus Pod GET on the receiving side. Sandbox service routing
does this per HTTP request, so the apiserver calls scaled with traffic.

Cache all of it on ServerState:

- peer channels pooled per endpoint, so relays multiplex over one
  connection instead of redialing
- peer tokens keyed by SHA-256, expiring at min(ttl, token exp) so a hit
  cannot accept an expired token
- owner records for 3s against a 45s ownership TTL, still freshness
  checked before use

Entries are evicted when a relay fails. Also raise HTTP/2
max_concurrent_streams to 1024, since pooling funnels every relay between
two replicas onto one connection and hyper's default of 200 sits below
the 256 pending-relay budget.

Signed-off-by: divesh <dgude@nvidia.com>

* perf(server): pool upstream connections for sandbox services

Each HTTP request to a sandbox service opened its own supervisor relay,
paying a new TCP connection and HTTP/1 handshake every time. Worse, it
counted against the 32 in-flight relay cap, so a service handling more
than 32 concurrent requests failed outright.

Pool idle upstreams per endpoint and port, up to 8 each for 15s. Reuse is
safe because the pool only returns a connection hyper reports as ready,
and HTTP/1 cannot start a request until the previous body has drained.
Upgrades are never pooled since they take the connection over, and a
failed send evicts that endpoint. Pruning is bounded per key, with the
full sweep limited to once per 30s.

Signed-off-by: divesh <dgude@nvidia.com>

* fix(server): address HA gateway review findings (#3449)

- Let a gateway own supervisor sessions without a peer endpoint. Requiring
  one whenever the store is PostgreSQL broke every single-instance
  PostgreSQL deployment, because no sandbox supervisor could connect.
  A cross-replica request to an owner that advertises no endpoint now fails
  immediately naming the cause, instead of retrying until the wait timeout.
- Close a supervisor session on heartbeat only when another replica owns it,
  or after renewals fail for the ownership TTL. A database error no longer
  drops every session heartbeating during an outage.
- Clamp owner record ages at zero so a skewed or corrupt stored timestamp
  cannot produce a negative age.
- Bound the cross-object advisory lock with a lock timeout, so a stuck holder
  fails instead of blocking every mutation in the fleet.
- Refuse to start when a peer endpoint is configured on a multi-replica
  backend but peer authentication is unavailable, and warn when a
  multi-replica backend has no peer endpoint at all.
- Reject a plaintext peer endpoint when the gateway serves TLS.
- Skip the sandbox watch poller on single-replica backends, where the local
  update bus already sees every write.
- Rate-limit the peer owner cache sweep so an insert no longer scans the
  whole map under the lock.
- Retry GET and HEAD on a pooled upstream the sandbox closed, instead of
  returning 502, and drop an emptied endpoint from the pool right away.
- Document the gateway peer environment variables and the post-rollout
  ownership skew operators should expect.

Signed-off-by: divesh <dgude@nvidia.com>

* fix(server): harden HA supervisor ownership

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Signed-off-by: divesh <dgude@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: divesh <dgude@nvidia.com>
Co-authored-by: Divesh Chowdary <47188680+FrostGod@users.noreply.github.com>
2026-09-21 21:22:12 +00:00
Roshni Malani 586c385bdf chore(k8s): use upstream agent-sandbox manifest in CI/e2e (#1657)
Drop the vendored deploy/kube/manifests/agent-sandbox.yaml. CI and e2e
scripts now apply manifest.yaml directly from
github.com/kubernetes-sigs/agent-sandbox releases, pinned via
AGENT_SANDBOX_VERSION env (default v0.4.6, overridable) so internal runs
stay reproducible.

Public docs and the helm chart README continue to reference
/releases/latest/download/ — no OpenShell/agent-sandbox support matrix
exists yet (see #1649).

Adds an air-gap note in docs/kubernetes/setup.mdx enumerating the
manifest and image operators need to mirror to an internal registry.

Signed-off-by: Roshni Malani <rmalani@nvidia.com>
2026-06-04 11:35:06 -07:00
Taylor MutchandDrew Newberry b61a98dbad feat(gateway): add TOML configuration file (RFC 0003) (#1317)
* feat(gateway): add TOML configuration file (RFC 0003)

Introduces an opt-in --config / OPENSHELL_GATEWAY_CONFIG flag that loads a
TOML file with gateway-wide settings and per-driver tables. Source
precedence is CLI > env > file > built-in default, implemented via clap's
ValueSource so existing flags and env vars keep their priority.

Driver crates (kubernetes, docker, podman, vm) now derive Deserialize on
their config structs. SupervisorSideloadMethod gains Deserialize with
kebab-case rename. A per-driver inheritance allowlist on the loader side
overlays [openshell.gateway] shared defaults (default_image,
supervisor_image, image_pull_policy, guest_tls_*, ssh_handshake_skew_secs,
client_tls_secret_name, host_gateway_ip, enable_user_namespaces) onto
each [openshell.drivers.<name>] table before deserialization.

The Helm chart renders a new gateway-config ConfigMap and mounts it at
/etc/openshell/gateway.toml. The migrated OPENSHELL_* env entries are
dropped from the StatefulSet — only the Secret-backed
OPENSHELL_SSH_HANDSHAKE_SECRET remains. database_url stays on --db-url.

Adds examples/gateway/gateway.example.toml and updates architecture/gateway.md
with the source precedence and inheritance rules.

* docs(gateway): drop ssh_handshake_skew_secs and ssh_handshake_secret from examples

Both fields are scheduled for removal. Remove the example values and the
env-only note so the gateway.toml example and the architecture doc stop
recommending settings that will not exist much longer.

* docs(rfc): correct OPENSHELL_CONFIG to OPENSHELL_GATEWAY_CONFIG in RFC 0003

* docs(gateway): add per-driver TOML example configurations

Adds focused single-driver examples next to the comprehensive
gateway.example.toml: kubernetes, docker, podman, and microvm. Each one
demonstrates the realistic settings for that driver plus how shared
[openshell.gateway] defaults inherit into the driver table.

A new unit test (`checked_in_examples_parse`) loads every example through
the config_file loader so schema drift fails CI rather than silently
shipping a broken example.

* refactor(gateway): drop image_pull_policy from shared inheritance

Kubernetes and Podman use mutually-incompatible vocabularies for the same
TOML key:

  - Kubernetes: `Always | IfNotPresent | Never` (free-form string passed
    verbatim to the K8s API).
  - Podman: `always | missing | never | newer` (strict lowercase enum
    deserialised into `ImagePullPolicy`).

No value means the same thing in both drivers. Sharing the key at
`[openshell.gateway]` scope and inheriting it into every active driver's
table meant any value safe for one driver was either wrong or silently
dropped for the other (`IfNotPresent` → `ImagePullPolicy::Missing` after
`.unwrap_or_default()`). Operators run one driver per gateway, so the
"shared default" never pays for itself.

Make `image_pull_policy` driver-local:

  - Remove the field from `GatewayFileSection` and from
    `inheritable_keys()` for both Kubernetes and Podman.
  - Drop the file→`RunArgs` merge for the gateway-scope key.
  - Stop unconditionally clobbering the driver value with
    `config.sandbox_image_pull_policy` in the runtime wiring — only apply
    the CLI/env override when it was set (and, for Podman, only when it
    parses into the lowercase enum).
  - Move the key under `[openshell.drivers.kubernetes]` and
    `[openshell.drivers.podman]` in every example, the RFC, the
    architecture doc, and the Helm-rendered gateway ConfigMap.

The supervisor pull policy follows the same shape: it is K8s-only and
moves into `[openshell.drivers.kubernetes]` alongside `image_pull_policy`
in the Helm template.

* fix(gateway): address review feedback on TOML configuration

Resolves the P1 and P2 issues raised in PR #1317:

- Helm gateway ConfigMap moves `grpc_endpoint` under
  `[openshell.drivers.kubernetes]` so the default install no longer fails
  the gateway's `deny_unknown_fields` schema check.
- `kubernetes_config_from_file` and `podman_config_from_file` only let
  the gateway-wide CLI/env `grpc_endpoint` overwrite the driver-table
  value when it was actually supplied, preserving file-only configs.
- Kubernetes driver default `image_pull_policy` is now empty (was Podman
  vocabulary "missing"), so default deployments let the Kubernetes API
  apply its own policy instead of being rejected.
- New `disable_tls` gateway field plumbs `.Values.server.disableTls`
  through the TOML ConfigMap instead of relying on env vars dropped from
  the StatefulSet.
- StatefulSet pod template now carries a `checksum/gateway-config`
  annotation so `helm upgrade` rolls pods when the ConfigMap changes.
- Auxiliary listener resolution preserves the full `SocketAddr` from
  `health_bind_address` / `metrics_bind_address`, so a loopback-pinned
  health port is not silently relocated onto the public bind address.
- `ssh_session_ttl_secs` from the file is now applied to `Config` (it
  was previously accepted by the loader but never read).

New regression coverage: cli-level merge tests for the new fields plus
helm-unittest assertions for the ConfigMap shape, checksum annotation,
and `disable_tls` rendering.

* docs(gateway): consolidate gateway TOML examples into docs reference

Replaces the per-driver example files under examples/gateway/ with a
single published reference page at docs/reference/gateway-config.mdx
covering source precedence, layout, the full example, and the four
per-driver examples (Kubernetes, Docker, Podman, microVM). Drew flagged
during PR #1317 review that the examples belong with the user-facing
docs rather than in a sibling examples/ directory.

The cross-references in architecture/gateway.md and RFC 0003 are updated
to point at the new docs page; the round-trip test in config_file.rs is
removed (schema coverage stays on the inline parses_full_example test
and per-field merge tests — doc-snippet drift belongs in a separate
docs-lint, not in a cross-tree Rust unit test).

* refactor(core): move DEFAULT_K8S_NAMESPACE into K8s driver

The constant is Kubernetes-specific (used only by KubernetesComputeConfig's
Default impl) and does not belong in openshell-core. Relocate it to the
driver crate that owns the K8s vocabulary; openshell-core retains only
truly cross-cutting defaults.

* refactor(core): move Podman bridge default into Podman driver

DEFAULT_NETWORK_NAME is Podman vocabulary, consumed only by the Podman
driver. Also drops the unused DEFAULT_IMAGE_PULL_POLICY constant.

* docs(auth): scrub remaining SSH handshake secret references

Sweeps the trailing mentions left after the rebase: the gateway
config-file module doc, the Helm gateway-config ConfigMap header,
the gateway-config.mdx env-only note, and the RPM systemd unit
comment for init-gateway-env.sh.

* docs(gateway): clarify OPENSHELL_GRPC_ENDPOINT applies to all drivers

The previous comment implied the callback endpoint was Kubernetes-only,
but the value is propagated to every compute driver (Kubernetes, Docker,
Podman, VM) and must be reachable from wherever the sandbox runs.

* refactor(gateway): move driver options into config (#1394)

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(e2e): regenerate gateway config via TOML for docker + podman harnesses

The gateway CLI flags moved into TOML config tables in 560550d2 (#1394),
which made every existing e2e/with-{docker,podman}-gateway.sh invocation
fail with "unexpected argument '--sandbox-namespace'" (and a long tail
of similar driver-specific options) before the gateway could even bind.

Replace the obsolete CLI flags with a synthesized
`[openshell.drivers.<driver>]` table written to `${STATE_DIR}/gateway.toml`
and passed via the new `--config` flag. Only the gateway-wide flags
that survived 560550d2 (bind-address, port, drivers, db-url, tls-*,
disable-tls, log-level, health-port) stay on the command line.

Both scripts get a small `toml_string` helper to properly TOML-quote
the values (the previous `%q` printf format produced bash-escape, not
TOML-escape). The Docker harness also corrects two field names that
diverged from the driver schema: `docker_network_name` →
`network_name`, and the supervisor binary/image plumbing now reads
through to `supervisor_bin` / `supervisor_image` in the same table.

The Podman harness drops `--ssh-gateway-port` (deleted in 560550d2 —
gRPC + SSH are multiplexed on the same port now) and substitutes
`network_name` + `gateway_port` for the obsolete `--sandbox-namespace`
(which the Podman driver never had as a typed field).

* fix(core): swap bind-only 0.0.0.0 SSH gateway host for cluster URL host

CLI's resolve_ssh_gateway treated 0.0.0.0 as a loopback "keep as-is"
when the cluster URL was also loopback, so the SSH proxy connected to
0.0.0.0:port. The unspecified address is never a valid connect target
and is not present in any TLS cert SAN, which produced BadCertificate
TLS handshake failures during `openshell sandbox create -- ...` in
docker/podman e2e (e.g. bypass_detection).

Resolution: when the server returns 0.0.0.0 or :: as the gateway host
and both endpoints are loopback, fall back to the cluster URL's host
(which the CLI is already using to reach the gateway, so it must
resolve and match the cert).

* fix(e2e): repair podman harness on macOS

Podman 5.x with the applehv/libkrun provider no longer creates the legacy
~/.local/share/containers/podman/machine/podman.sock symlink, and
`podman system service` is a Linux-only subcommand — the macOS client
delegates the API service to the VM. Both assumptions in the harness
were stale, so the script tried to start a temporary service that podman
rejected with "unknown flag: --time".

- Discover the macOS socket via `podman machine inspect` instead of the
  hardcoded path.
- On Darwin, fail fast with a "start podman machine" message rather than
  attempting the Linux-only `podman system service` fallback.
- Write socket_path into [openshell.drivers.podman] so the in-process
  driver picks up the discovered socket; the driver reads TOML only
  after the config refactor (560550d2), so OPENSHELL_PODMAN_SOCKET alone
  was no longer enough.

* fix(server): use clone_from for TLS client CA assignment

clippy 1.95.0 rejects assigning the result of `Clone::clone()` to an
existing variable under `-D warnings` (`assigning_clones`). Switch to
`clone_from(&...)` to satisfy the lint and avoid the redundant
allocation.

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
2026-05-15 12:43:48 -07:00
Seth Jennings c94cddbfb8 feat(server): separate HTTPS from mTLS authentication (#1351)
Make --tls-client-ca optional and make client certificates always
optional when a CA is configured. This decouples HTTPS encryption
from mTLS authentication, allowing mTLS and OIDC bearer tokens to
coexist as parallel authentication mechanisms.

When --tls-client-ca is provided, client certificates are validated
against the CA when presented but never required. Clients may connect
with or without a certificate — authentication is handled at the
application layer (e.g. OIDC).

Two TLS modes are now supported:
- HTTPS with optional mTLS (--tls-client-ca provided)
- HTTPS-only (--tls-client-ca omitted)

The --disable-gateway-auth flag is preserved for backward
compatibility but is now a no-op. The allow_unauthenticated field
has been removed from TlsConfig. The Helm chart conditionally
includes the client-ca volume and env var based on whether
clientCaSecretName is configured.
2026-05-15 09:43:30 -07:00
Taylor Mutch 316c788eac fix(helm): derive grpcEndpoint from chart context (#1241)
* fix(helm): derive grpcEndpoint from chart context

The chart hardcoded server.grpcEndpoint to
https://openshell.openshell.svc.cluster.local:8080, which only matched the
in-cluster Service DNS for the standard release name and namespace. A new
helper now builds <scheme>://<fullname>.<namespace>.svc.cluster.local:<port>
from chart context, picking the scheme from server.disableTls. An explicit
server.grpcEndpoint override is passed through verbatim.

* chore(scripts): validate k3d cluster name length early

helm-k3s-local.sh derives the cluster name from the current branch suffix.
Long branch names produced names exceeding k3d's 32-char cap and failed
deep inside k3d cluster create with a confusing validation error. cmd_create
now bails out before invoking docker/k3d with a copy-pasteable
HELM_K3S_CLUSTER_NAME override hint. Status, start, stop, delete, and help
remain unaffected so an over-long derived name does not block diagnostics.
2026-05-08 14:10:19 -07:00
Seth Jennings bb4dbd7c65 fix(kube): add RBAC rule for sandbox finalizer updates (#1203) 2026-05-06 10:12:38 -07:00
Taylor Mutch 5116cc27b7 feat(helm): add kubernetes local-dev environment (#1158) 2026-05-05 13:42:21 -07:00
John T. Myers 043bde279a feat(providers): add profile-backed policy composition (#1037)
Foundation for providers v2. Add provider profiles and provider profile composition with user policies.
2026-05-04 18:34:33 -07:00
jtoelke2 32857eb650 fix(helm): grant node read access for GPU capacity checks (#1106)
* fix(helm): grant node read access for GPU capacity checks

Signed-off-by: Jonas Toelke <jtoelke@nvidia.com>

* ci: add WSL GPU failure diagnostics

Signed-off-by: Jonas Toelke <jtoelke@nvidia.com>

* fix(gpu): bump device plugin for WSL CDI

* fix(sandbox): allow WSL GPU paths in Landlock

Allow /dev/dxg and /usr/lib/wsl as GPU baseline paths so WSL CDI GPU sandboxes can initialize NVML. Native Linux skips these entries when the paths do not exist.

* ci: remove temporary WSL GPU diagnostics

---------

Signed-off-by: Jonas Toelke <jtoelke@nvidia.com>
2026-05-01 17:49:49 -05:00
Mrunal Patel 084505425b feat(auth): add OIDC/Keycloak authentication with RBAC and scope-based permissions (#935)
* feat(auth): add OIDC/Keycloak authentication with RBAC

Add OAuth2/OIDC authentication to the gateway server with role-based
access control, CLI login flows, and full deployment plumbing.

Server: JWT validation against configurable OIDC issuer (oidc.rs),
JWKS key caching with TTL and rotation handling, method classification
(unauthenticated/sandbox-secret/dual-auth/bearer), identity extraction
with provider-agnostic Identity type, and RBAC enforcement via
AuthzPolicy with configurable admin/user roles and auth-only mode.

CLI: browser-based Authorization Code + PKCE flow, Client Credentials
flow for CI/automation, token storage with refresh, gateway add/login/
logout commands, OIDC bearer token injection over mTLS transport,
discovery endpoint for auto-configuration.

Security: sandbox-secret scope restriction on UpdateConfig (policy
sync only), anti-spoofing header stripping, dual-auth fallthrough
from sandbox-secret to Bearer token.

Deployment: OIDC config wired through DeployOptions, Docker env vars,
Helm values/templates, HelmChart manifest, cluster-entrypoint.sh, and
bootstrap scripts. Keycloak dev server script with pre-configured
realm (test users, roles, PKCE client, CI client).

Tested with Keycloak. The roles claim path and role names are
configurable to support other OIDC providers.

* feat(auth): add OAuth2 scope-based fine-grained permissions

Add opt-in scope enforcement on top of existing OIDC role-based access
control. When --oidc-scopes-claim is set, the server extracts scopes
from the JWT and checks them per-method against an exhaustive scope map.

Scopes: sandbox:read, sandbox:write, provider:read, provider:write,
config:read, config:write, inference:read, inference:write, and
openshell:all (wildcard). Methods not in the scope map require
openshell:all. Scopes layer on top of roles and cannot escalate
privilege. Auth-only mode (empty role names) still enforces scopes
when enabled.

Server: scopes_claim in OidcConfig, scope extraction from JWT
(space-delimited and JSON array formats), standard OIDC scope
filtering, scope check in AuthzPolicy after role check.

CLI: --oidc-scopes on gateway add/start stored in metadata and
consumed by gateway login, --oidc-scopes-claim on gateway start
forwarded to server, scopes parameter in browser and client
credentials OAuth2 flows with openid deduplication.

Deployment: oidc_scopes_claim wired through DeployOptions, docker.rs,
Helm, bootstrap scripts, and cluster entrypoint.

Keycloak: realm config updated with built-in OIDC scopes and 9
OpenShell client scopes as optional on openshell-cli and openshell:all
as default on openshell-ci.

* fix(auth): address branch review findings

Add GetInferenceBundle to sandbox-secret methods so sandbox inference
route refresh works under OIDC. Make GetSandboxConfig dual-auth so CLI
users can read sandbox settings with Bearer tokens.

Preserve OIDC gateway metadata on restart — a bare gateway start
without --oidc-* flags no longer erases the stored OIDC registration.

Document CI client ID requirement (openshell-ci vs openshell-cli) in
the testing guide. Add security note about auth-only mode blast radius
for GitHub Actions.

* fix(auth): complete review findings for OIDC auth boundary

Move OpenShell/GetSandboxConfig from sandbox-secret-only to dual-auth
so CLI users can read sandbox settings with Bearer tokens while sandbox
supervisors continue using the shared secret.

Add sandbox secret interceptor to the inference bundle fetch path so
GetInferenceBundle works under OIDC-enabled gateways. Extract shared
interceptor constructor to avoid duplication.

Add GetSandboxConfig to the config:read scope map so scope enforcement
applies consistently when scopes are enabled.

Refactor OIDC metadata preservation into apply_oidc_gateway_metadata()
with explicit resume semantics — only preserve existing OIDC metadata
on real resume paths, not on fresh deployments.

Update architecture docs and testing guide to reflect the corrected
method classifications and add new test coverage for interceptor
injection, scope requirements, metadata preservation, and dual-auth
classification.

* refactor(auth): use oauth2 crate for CLI OIDC flows

Replace hand-written PKCE generation, authorization URL construction,
token exchange, client credentials, and token refresh with the oauth2
crate's typed API.

Eliminates sha2, hex, and getrandom dependencies from the CLI. The
custom urlencoded() helper and manual form POST logic are replaced by
BasicClient methods with proper type-state safety.

Discovery and the callback server remain custom since the oauth2 crate
does not provide OIDC discovery or a localhost redirect listener.

* refactor(auth): move server auth modules into auth/ directory

Group oidc.rs, authz.rs, identity.rs, and the auth HTTP endpoints
under src/auth/ module directory. No behavioral changes.

  auth/mod.rs      — module root, re-exports HTTP router
  auth/oidc.rs     — JWT validation, JWKS caching, method classification
  auth/authz.rs    — role and scope authorization policy
  auth/identity.rs — provider-agnostic Identity type
  auth/http.rs     — /auth/connect and /auth/oidc-config endpoints

* fix(auth): use RequestBody auth type for client credentials flow

The oauth2 crate defaults to BasicAuth (HTTP Basic header) but Keycloak
and most OIDC providers expect client_secret_post (credentials in the
request body). Set AuthType::RequestBody explicitly to match the
pre-refactor behavior.

Also re-export Identity, IdentityProvider, and JwksCache from the auth
module so ServerState's public API remains nameable by external consumers.

* fix(auth): forward OPENSHELL_OIDC_SCOPES through cluster bootstrap

Pass --oidc-scopes to gateway start so the metadata includes requested
scopes after cluster bootstrap. Without this, users had to manually
edit metadata.json to set scopes for gateway login.

Usage: OPENSHELL_OIDC_SCOPES="openshell:all" mise run cluster

* test(auth): add OIDC e2e tests for RBAC, scopes, and client credentials

Add 10 end-to-end tests covering OIDC authentication against a live
K3s cluster with Keycloak:

RBAC (5 tests): admin can create providers, user cannot, user can list
sandboxes, unauthenticated requests rejected, health probe works
without auth.

Scopes (4 tests): sandbox-scoped token can list sandboxes but not
providers, openshell:all grants full access, no-scopes token denied.

Client credentials (1 test): CI token via client_credentials grant.

Tests are opt-in via OPENSHELL_E2E_OIDC=1 and OPENSHELL_E2E_OIDC_SCOPES=1
env vars. They derive the Keycloak URL from gateway metadata to match
the server's configured issuer.

Run with:

  OPENSHELL_E2E_OIDC=1 OPENSHELL_E2E_OIDC_SCOPES=1 \
  PYTHONPATH=python uv run pytest e2e/python/oidc/ -v

* fix(docs): fix markdown lint errors in OIDC architecture docs

Add blank lines before lists and fenced code blocks to satisfy
markdownlint MD031 and MD032 rules.
2026-04-30 10:37:23 -07:00
Drew Newberry ddb85b1704 feat(vm): add openshell-vm crate with libkrun microVM gateway (#611) 2026-04-08 22:00:01 -07:00
Drew Newberry e837849009 feat(bootstrap): resume gateway from existing state and persist SSH handshake secret (#488) 2026-04-02 09:50:10 -07:00
Evan Lezar 122bc74948 feat(sandbox): switch device plugin to CDI injection mode (#503)
* feat(sandbox): switch device plugin to CDI injection mode

Configure the NVIDIA device plugin to use deviceListStrategy=cdi-cri so
that GPU devices are injected via direct CDI device requests in the CRI.
Sandbox pods now only require the nvidia.com/gpu resource request —
runtimeClassName is no longer set on GPU pods.

Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-03-31 08:35:39 -07:00
Evan Lezar dac6cd953d feat(gpu): disable NFD/GFD and remove nodeAffinity from device plugin chart (#497)
Disables GPU Feature Discovery and Node Feature Discovery DaemonSets and
overrides the device plugin's default nodeAffinity to null so it schedules
unconditionally on the single-node gateway without requiring NFD/GFD labels.

Setting affinity to an empty map ({}) does not override the chart defaults
because Helm deep-merges user values with chart defaults. Using null explicitly
removes the key, causing the chart template to skip the affinity block entirely.
2026-03-20 10:45:30 -07:00
Drew NewberryandPiotr Mlocek 83af7a245b feat(sandbox): inject host gateway hostAliases into sandbox pods (#306)
* feat(sandbox): inject host gateway hostAliases into sandbox pods

Sandbox pods running in the k3s cluster cannot resolve host.docker.internal
by default, preventing them from reaching services on the Docker host.

Detect the host gateway IP (default route) in the cluster entrypoint,
thread it through the Helm chart to the gateway server, and inject
hostAliases entries (host.docker.internal, host.openshell.internal)
into every sandbox pod spec. The injection is conditional -- when the
IP is empty (non-Docker deployments), no hostAliases are added.

---------

Co-authored-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-03-14 16:54:03 -07:00
Drew Newberry 89d21d7852 refactor(sandbox): sandboxes are managed as separate community images (#267) 2026-03-12 22:06:52 -07:00
Drew NewberryandPiotr Mlocek d94d4e1166 feat(cluster): add NVIDIA GPU passthrough support for gateway start (#234)
Add --gpu flag to 'openshell gateway start' that enables NVIDIA GPU
access inside the cluster container. When enabled, the Docker container
receives GPU device requests via bollard, k3s auto-detects the NVIDIA
runtime, and the NVIDIA k8s-device-plugin is deployed via HelmChart CR
so Kubernetes workloads can request nvidia.com/gpu resources.

- Add GPU device request to Docker container creation (bollard DeviceRequest)
- Set GPU_ENABLED=true env var to gate GPU manifest deployment
- Multi-stage Dockerfile build to bundle NVIDIA Container Toolkit binaries
- Conditional GPU manifest copy in cluster entrypoint
- HelmChart CR for nvidia-device-plugin v0.18.2 with GFD and NFD enabled
- Thread gpu bool through DeployOptions, CLI flag, and run dispatch

---------

Co-authored-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-03-11 16:53:24 -07:00
Drew Newberry f97270f988 refactor(docker): rename server image to gateway (#246)
* refactor(docker): rename server image to gateway

Rename Dockerfile.server to Dockerfile.gateway and update all image
references from openshell/server to openshell/gateway across Helm
charts, Kubernetes manifests, mise tasks, build/deploy scripts, CI
workflows, and documentation.

The underlying Rust binary (navigator-server) is unchanged -- this
rename only affects the Docker image name and Dockerfile.

* fix: catch remaining server->gateway references in docs and comments
2026-03-11 16:01:26 -07:00
Drew Newberry d6c6e97679 chore: remove navigator references from codebase (#208) 2026-03-10 14:46:11 -07:00
Drew Newberry 984d1a6e5c chore: rename project from NemoClaw to OpenShell (#198) 2026-03-10 11:49:09 -07:00
Drew Newberry 1cf54ca055 feat(bootstrap): switch container registry from CloudFront CDN to GHCR with token auth (#167) 2026-03-09 23:46:23 -07:00
Drew Newberry a2de1f24e5 feat: add Cloudflare tunnel auth support (#178) 2026-03-09 18:39:42 -07:00
John T. MyersandJohn Myers e9f10719f4 fix(security): harden sandbox SSH with mandatory HMAC secret, NetworkPolicy, and nonce replay detection (#127)
* fix(security): harden sandbox SSH with mandatory HMAC secret, NetworkPolicy, and nonce replay detection

Closes #25

- Make NEMOCLAW_SSH_HANDSHAKE_SECRET mandatory: server and sandbox both
  refuse to start if the secret is empty/unset. Cluster deployments
  auto-generate it via openssl rand in the entrypoint script.
- Add Kubernetes NetworkPolicy restricting sandbox port 2222 ingress to
  the gateway pod only, preventing lateral movement from other cluster
  workloads.
- Add NSSH1 nonce replay detection with a TTL-bounded cache, rejecting
  replayed handshakes within the timestamp validity window.
- Add unit tests for verify_preface (valid, replay, expired, bad HMAC,
  malformed) and env injection.

* fix(deploy): pass sshHandshakeSecret in fast deploy helm upgrade

---------

Co-authored-by: John Myers <johntmyers@users.noreply.github.com>
2026-03-05 12:21:18 -08:00
Drew Newberry d920d39dd4 chore: more contributing improvements (#103) 2026-03-04 13:20:24 -08:00
Drew Newberry 9099bc3972 chore: rename Navigator to NemoClaw across user facing contracts (#73) 2026-03-03 11:41:19 -08:00
Alexander Watson 1d7909cb38 chore: add open-source compliance files and SPDX headers (#71)
Add Apache 2.0 licensing, SPDX copyright headers on all source files,
DCO enforcement, third-party notices, and CI enforcement.

- LICENSE: Apache License 2.0 full text
- DCO: Developer Certificate of Origin 1.1
- SPDX headers on all 176 source files (.rs, .py, .proto, .rego, .sh,
  .toml, .yaml, Dockerfiles)
- scripts/update_license_headers.py: header management with --check mode
- scripts/generate_third_party_notices.py: dependency license aggregation
- THIRD-PARTY-NOTICES: generated listing of all Rust and Python deps
- build/license.toml: mise tasks for license:check and license:update
- CI: license-headers job in checks.yml, DCO check workflow
- CONTRIBUTING.md: DCO sign-off requirement and license header docs
- Cargo.toml: license changed to Apache-2.0, repository URL updated
- pyproject.toml: license field added

Closes #58
2026-03-03 09:30:56 -08:00
Drew Newberry 2d85338940 feat(platform): cleanup api surface area and mtls flows (!39)
Closes #48, #52

## Summary
- Replace the envoy-gateway-based TLS setup with inline PKI generation during cluster bootstrap, generating CA, server, and client certificates directly in the `navigator-bootstrap` crate
- Remove all envoy gateway Helm templates (`gateway.yaml`, `gatewayclass.yaml`, `grpcroute.yaml`, PKI job, traffic policies) and the `Dockerfile.pki-job`
- Add native mTLS support to the navigator server with `tokio-rustls`, mounting client TLS certs as volumes into sandbox pods
- Update cluster entrypoint, healthcheck, and deploy scripts to work with the new direct-TLS architecture
- Add TLS security e2e test and fix formatting/clippy warnings

## Test Plan
- All unit tests pass (`cargo test --workspace`)
- Clippy clean (`cargo clippy --workspace --all-targets`)
- Format clean (`cargo fmt --all -- --check`)
- Python tests pass (`uv run pytest python/`)
- Full `mise run pre-commit` passes
2026-02-24 11:33:33 -08:00
Drew Newberry 24b9654ea8 feat(cluster): add remote SSH deployment 2026-02-10 08:25:43 -08:00
Drew Newberry 00f432de8c feat: add mtls support to plaform 2026-02-05 17:24:48 -08:00
Drew Newberry 207ebe4a43 chore: cleanup docker/kube/helm infra 2026-02-04 23:00:17 -08:00