generate_pki minted a CA with no key usage and server and client leaves
with no Authority Key Identifier. RFC 5280 requires both, and verifiers
that enforce it reject the chain: OpenSSL X509_STRICT fails with
"Missing Authority Key Identifier", and Python 3.13 turned that flag on
by default in ssl.create_default_context(). rustls and BoringSSL do not
enforce it, so gRPC clients kept working while an HTTPS client built on
Python 3.13 (for example a platform proxying to an exposed sandbox
service) could not complete a handshake with a pkiInitJob-provisioned
gateway at all. cert-manager PKI was unaffected.
Set keyCertSign and cRLSign on the CA and use_authority_key_identifier
on both leaves, matching what the sandbox L7 CA already does. Add a
test that parses the bundle and asserts the extensions, including that
each leaf AKI matches the CA SKI.
Verified: openssl verify -x509_strict accepts both leaves, and a strict
Python 3.13 client completes an mTLS handshake against a server using
the new bundle where the previous bundle reproduces the failure.
Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com>
Establish the Phase 1 reference implementation proposed by RFC 0012
while retaining the existing Cargo and Mise workflows during evaluation.
- pin Bazel 9.1.1 and configure Bzlmod, rules_rs, LLVM, protobuf, and
Rust 1.95 toolchains
- import third-party crates from Cargo metadata and propagate the
workspace version into Bazel targets
- add library, binary, proc-macro, unit-test, and integration-test
targets across the supported Rust workspace crates and drivers
- generate protobuf Rust sources and descriptor sets under Bazel while
preserving Cargo-compatible generated-code imports
- annotate aws-lc-sys and zstd-sys native dependencies, build Z3 4.15.2
from source, and generate z3-sys bindings
- make CLI and procfs test fixtures available as explicit Bazel inputs
without relying on fixed host binary paths
- define optimized release targets for Linux x86_64 and aarch64 CLI,
sandbox, and gateway binaries, plus macOS aarch64 artifacts
- add Bazel, buildifier, and lcov to the Nix development environment
RFC: 0012 (rfc12 branch)
Refs: #2491
Signed-off-by: Simon Scatton <sscatton@nvidia.com>
* fix(supervisor-process): use slice patterns when rewriting passwd/group lines
rewrite_passwd_at and rewrite_group_at indexed into a Vec of fields
with fields[N] after checking fields.len(). Use slice patterns so a
malformed line falls through to the no-op branch instead of risking a
panic if the guard is ever changed. Add regression tests for
malformed sandbox entries.
Signed-off-by: Andrew White <andrewh@cdw.com>
* fix(bootstrap): avoid byte-index slice after @ in SSH destination
extract_host_from_ssh_destination sliced dest[at_pos + 1..] after
finding '@'. '@' is ASCII so this is currently safe, but it is a
latent panic surface if the split logic changes. Use get() instead
and add a multi-byte hostname test.
Signed-off-by: Andrew White <andrewh@cdw.com>
* fix(bootstrap): avoid byte-index slice in .dockerignore glob matcher
glob_match sliced path[idx + 1..] after matching '/'. '/' is ASCII so
this is currently safe, but it is a latent UTF-8 panic surface. Use
get() instead.
Signed-off-by: Andrew White <andrewh@cdw.com>
* fix(router): avoid literal byte offset when stripping /v1 prefix
build_backend_url sliced &path[3..] after verifying the /v1 prefix.
Use strip_prefix() so the code stays correct if the prefix length
ever changes.
Signed-off-by: Andrew White <andrewh@cdw.com>
* fix(supervisor-network): avoid byte-index slice in inference path matching
The Bedrock-style /*/ path matcher sliced rest[slash_at + 1..] after
finding '/'. '/' is ASCII so this is currently safe, but it is a
latent UTF-8 panic surface. Use get() instead.
Signed-off-by: Andrew White <andrewh@cdw.com>
---------
Signed-off-by: Andrew White <andrewh@cdw.com>
* feat(workspace): implement workspace model (Phase 1 of RFC 0011)
Implements workspace and membership model providing hard isolation
boundaries for multi-player OpenShell deployments.
Workspace CRUD with Kubernetes-style Terminating phase for graceful
deletion. All resources scoped by workspace via ObjectMeta. Membership
RPCs for workspace access control. Persistence migration shifts name
uniqueness to (object_type, workspace, name). Provider profiles support
platform and workspace scoping. Service routing uses workspace-prefixed
DNS labels. Inference routes renamed and workspace-scoped with
DeleteInferenceRoute RPC. Python SDK with WorkspaceClient, two-method
list pattern (workspace-scoped and for_all_workspaces), and workspace
parameter on all methods. CLI workspace flags, TUI workspace cycling.
K8s driver filters unmanaged CRs and uses delete preconditions. Podman
driver uses immutable container IDs. Label serialization fixed across
all put_if call sites.
Signed-off-by: Derek Carr <decarr@redhat.com>
* fix(cli): delegate sandbox upload command to existing upload function
The standalone `sandbox upload` command reimplemented upload logic
inline with two bugs: it used `Path::exists()` which follows symlinks
(rejecting dangling symlinks), and it ran git-aware filtering on
symlink sources. The `run::sandbox_upload()` function already handles
both cases correctly via `sandbox_upload_plan()`. Replace the inline
logic with a call to the existing function.
Signed-off-by: Derek Carr <decarr@redhat.com>
* fix(e2e): shorten sandbox names and fix test compatibility
Shorten the sandbox name in initial_sparse_policy_is_acknowledged_as_loaded
from 'e2e-2159-sparse-enrich' (22 chars) to 'e2e-sparse-enrich' (17 chars)
to comply with MAX_ROUTABLE_NAME_LEN (19 chars).
Also capture stderr in create_keep_with_args so future sandbox creation
failures include the actual CLI error instead of reporting empty output.
Signed-off-by: Derek Carr <decarr@redhat.com>
* test(workspace): add test coverage for workspace CRUD and persistence isolation
Add unit tests for workspace create happy path, get round-trip, get
not-found, get empty-name rejection, already-exists error, and
resolve_workspace not-found. Add persistence test proving cross-workspace
name uniqueness (same name in different workspaces produces separate
records). Add workspace name max-length boundary tests. Fix e2e harness
to include stderr in name-parse-failure error path. Align Python e2e
test_workspace_crud with try/finally pattern. Document provider profile
catalog workspace scoping gap in RFC 0011.
Signed-off-by: Derek Carr <decarr@redhat.com>
* fix(examples): update examples for workspace model compatibility
Shorten sandbox names in demo scripts to fit the 19-character
MAX_ROUTABLE_NAME_LEN limit: policy-demo prefix to pd-, multi-agent
notepad derives a short SANDBOX_TAG from the run ID, governance
interceptor uses gs-PID-RANDOM. Update vscode-remote-sandbox.md SSH
host aliases from openshell-{name} to openshell-{name}.{workspace}
format.
Signed-off-by: Derek Carr <decarr@redhat.com>
* feat(sdk): add workspace-scoped client and workspace CRUD
Add WorkspaceScopedClient modeled after kube::Api::namespaced — captures
workspace once and injects it into every sandbox request. Add workspace
CRUD methods (create, get, list, delete) and list_sandboxes_all_workspaces
on OpenShellClient. Extend SandboxRef with workspace field and add
WorkspaceRef type. Include mock tests for all new operations.
Signed-off-by: Derek Carr <decarr@redhat.com>
* fix(lint): resolve clippy warnings in workspace test assertions
Signed-off-by: Derek Carr <decarr@redhat.com>
* fix(docs): convert indented code blocks to fenced in RFC 0011
Signed-off-by: Derek Carr <decarr@redhat.com>
* fix(lint): resolve clippy warnings and apply cargo fmt across workspace
Auto-format with cargo fmt and fix clippy warnings exposed by the
reformat: unnecessary qualifications, map_unwrap_or, identical match
arms, unused variable prefix, dead code annotations, and let-unit-value
in e2e harness.
Signed-off-by: Derek Carr <decarr@redhat.com>
* fix(workspace): address workspace scoping issues from review
- Add workspace field to settings JSON output (CLI)
- Skip Podman containers missing workspace label instead of defaulting
to empty string, matching K8s driver behavior
- Add resource_version to list_by_scope SELECT in both SQLite and
Postgres backends, with regression test
- Gate PolicyLocalContext proposal/lookup routes on workspace readiness,
returning 503 when workspace is not yet discovered
- Block sandbox and provider creation in TUI all-workspaces mode
- Clear workspace vectors in TUI reset_sandbox_state
Signed-off-by: Derek Carr <decarr@redhat.com>
* fix(workspace): make provider profile catalog workspace-aware
Thread workspace through snapshot_catalog so the
EffectiveProviderProfileCatalog enforces workspace boundaries on both
read and write paths. UserProviderProfileSource now loads platform-scoped
profiles (workspace "") plus the target workspace's profiles, preventing
cross-workspace duplicate profile ID collisions that previously caused
global catalog failures.
Update RFC 0011 to reflect catalog scoping is implemented in Phase 1
rather than deferred to future work.
Signed-off-by: Derek Carr <decarr@redhat.com>
* fix(persistence): include workspace column in atomic policy revision INSERT
put_policy_revision_atomic omitted the workspace column from the INSERT
into the objects table in both SQLite and Postgres backends, causing
atomically-written policy revisions to lose their workspace association.
Add workspace field to AtomicPolicyRevisionWrite and thread it through
both backend INSERT statements, matching the non-atomic put_policy_revision
path which already included it.
Signed-off-by: Derek Carr <decarr@redhat.com>
* fix(proxy): skip ancestor walk when socket owner is the entrypoint
collect_ancestor_identities walked the entire process tree above the
entrypoint when the connecting process was the entrypoint itself,
SHA256-hashing every ancestor binary (IDE, shell, container runtime).
On dev machines with large binaries in the ancestor chain this exceeded
the 30-second test timeout. When start_pid == stop_pid there are no
intermediate ancestors to verify, so return an empty list immediately.
Signed-off-by: Derek Carr <decarr@redhat.com>
* fix(workspace): make provider profile catalog scope-aware
Allow the same profile ID at platform and workspace scopes by
introducing layered catalog entries where workspace profiles shadow
platform profiles. Add source and scope fields to the ProviderProfile
proto and CLI output. Migrate List/Get handlers to the catalog,
fixing divergence with runtime profile resolution.
Signed-off-by: Derek Carr <decarr@redhat.com>
* fix(e2e): align podman e2e labels with centralized driver constants
The podman driver moved its container labels to the centralized
openshell.ai/ prefix, but the e2e test harness and cleanup script
still referenced the old openshell.sandbox-* keys, causing the
local_driver_token_restart test to fail on container lookup.
Signed-off-by: Derek Carr <decarr@redhat.com>
* fix(e2e): align python profile isolation test with scope-aware catalog
Platform profiles are now visible in workspace listings as fallbacks
per the layered catalog design. Update the assertion to match.
Signed-off-by: Derek Carr <decarr@redhat.com>
* fix(workspace): honor profile_workspace in runtime profile resolution
Runtime profile lookups now consult provider.profile_workspace via
get_type_profile_for_scope. Providers created with --global-profile
(profile_workspace="") resolve to the platform profile even when a
workspace profile shadows the same ID. All 6 runtime call sites
updated; type-only call sites remain scope-agnostic.
Signed-off-by: Derek Carr <decarr@redhat.com>
---------
Signed-off-by: Derek Carr <decarr@redhat.com>
* feat(bootstrap): add system gateway registry for installer defaults
Adds a read-only installer-seeded gateway registry that the CLI consults after per-user gateway config. The registry uses the same layout as per-user config with `active_gateway` at the root and `gateways/<name>/metadata.json` beneath it. By default the system config root is `/etc/openshell`, while `OPENSHELL_SYSTEM_GATEWAY_DIR` remains available as an override for packages that need a different location. User-managed gateways continue to shadow installer entries on name collision.
Originally-authored-by: Mark Shuttleworth <mark@ubuntu.com>
Signed-off-by: Alex Lewontin <alex.lewontin@canonical.com>
* feat(cli): show gateway config source in list and term
Expose whether a gateway registration comes from user or system config in `openshell gateway list`, the TUI gateway pane, and list JSON output. The CLI also refuses to remove system-managed registrations and the smoke tests cover the new list output.
Signed-off-by: Alex Lewontin <alex.lewontin@canonical.com>
* fix(bootstrap): preserve user shadowing on invalid metadata
Signed-off-by: Alex Lewontin <alex.lewontin@canonical.com>
* fix(cli): keep system fallback out of rollback state
* docs(gateway): describe system config fallback layout
* test(bootstrap): cover system gateway last_sandbox persistence
* test(gateway): cover system-only removal rejection
* fix(bootstrap): validate gateway names before path joins
---------
Signed-off-by: Alex Lewontin <alex.lewontin@canonical.com>
* feat(telemetry): add build-time option to compile out telemetry
Gate anonymous telemetry emission behind a default-on `telemetry` Cargo
feature in openshell-core. The data model (enums, validation, emit_*/enabled*
signatures) stays always-compiled, while the endpoint, HTTP client, queue, and
emission code are feature-gated. With the feature off, enabled() returns false
and emit_* are no-ops, so dependent crates compile unchanged and no telemetry
endpoint, HTTP client, or emission code is included in the binary.
chrono and reqwest become optional dependencies of openshell-core, dropped from
its dependency graph when telemetry is disabled.
Thread the switch through the workspace: every crate depends on openshell-core
with default-features = false, and the default-on `telemetry` passthrough lives
on the binary crates that emit or collect telemetry (openshell-server,
openshell-sandbox, openshell-driver-vm). In-process drivers inherit it via
resolver v2 feature unification.
Build a telemetry-free binary with, e.g.:
cargo build --release -p openshell-server --no-default-features
The runtime OPENSHELL_TELEMETRY_ENABLED switch is unchanged for default builds.
Signed-off-by: Russell Bryant <russell.bryant@gmail.com>
* ci(telemetry): guard that telemetry can be compiled out
Add tasks/scripts/verify-telemetry-compiled-out.sh, which inspects a built
binary for telemetry markers (the telemetry endpoint host and client ID) that
exist only when emission code is compiled in. The rust:verify:telemetry-off
mise task builds the gateway with default features (positive control: markers
must be present, so the absent checks can never be silently vacuous) and with
--no-default-features (markers must be absent), and checks the
--no-default-features sandbox binary as well.
Wire the task into the Rust branch-checks job so a regression that reintroduces
telemetry code into a --no-default-features build fails CI.
Signed-off-by: Russell Bryant <russell.bryant@gmail.com>
---------
Signed-off-by: Russell Bryant <russell.bryant@gmail.com>
Docker BuildKit exposes TARGET* for the requested image platform. It exposes
BUILD* for the daemon running build stages. The bootstrap path did not pass an
explicit platform or seed those args, so Dockerfile behavior could depend on daemon defaults.
Set the build request platform from TARGETPLATFORM when supplied. Use the daemon
platform as the fallback. Query the daemon for BUILD*. Seed implicit args without
overwriting caller values. Do not inject empty variant args.
Tests
- mise run pre-commit
Signed-off-by: Shiju <shiju@nvidia.com>
* feat(rpm): replace init-pki.sh with openshell-gateway generate-certs
Cuts the RPM gateway over to the unified Rust certgen path. The systemd
user unit's first ExecStartPre now invokes:
/usr/bin/openshell-gateway generate-certs --output-dir %S/openshell/tls
producing the same six-PEM layout init-pki.sh built (ca.{crt,key},
server/tls.{crt,key}, client/tls.{crt,key}) and the same CLI mTLS copy
under $XDG_CONFIG_HOME/openshell/gateways/openshell/mtls/. None of the
OPENSHELL_TLS_* / OPENSHELL_PODMAN_TLS_* paths in the unit change.
Adds host.containers.internal to the gateway's built-in SAN list so
podman containers reaching their host validate cleanly with no
per-deployment --server-san flag. Docker (host.docker.internal) and
Kubernetes (cluster.local DNS) were already covered.
Drops 197 lines of openssl shell, the install/file lines for the script
itself, and updates the docs (man page, RPM CONFIGURATION.md, env-file
generator comment) to point at the new entrypoint. The %S state dir,
unit security hardening, and consumer paths are untouched.
* docs(certgen): remove stale init-pki.sh references in comments
---------
Co-authored-by: Taylor Mutch <taylormutch@gmail.com>
Make --tls-client-ca optional and make client certificates always
optional when a CA is configured. This decouples HTTPS encryption
from mTLS authentication, allowing mTLS and OIDC bearer tokens to
coexist as parallel authentication mechanisms.
When --tls-client-ca is provided, client certificates are validated
against the CA when presented but never required. Clients may connect
with or without a certificate — authentication is handled at the
application layer (e.g. OIDC).
Two TLS modes are now supported:
- HTTPS with optional mTLS (--tls-client-ca provided)
- HTTPS-only (--tls-client-ca omitted)
The --disable-gateway-auth flag is preserved for backward
compatibility but is now a no-op. The allow_unauthenticated field
has been removed from TlsConfig. The Helm chart conditionally
includes the client-ca volume and env var based on whether
clientCaSecretName is configured.
* feat(auth): add OIDC/Keycloak authentication with RBAC
Add OAuth2/OIDC authentication to the gateway server with role-based
access control, CLI login flows, and full deployment plumbing.
Server: JWT validation against configurable OIDC issuer (oidc.rs),
JWKS key caching with TTL and rotation handling, method classification
(unauthenticated/sandbox-secret/dual-auth/bearer), identity extraction
with provider-agnostic Identity type, and RBAC enforcement via
AuthzPolicy with configurable admin/user roles and auth-only mode.
CLI: browser-based Authorization Code + PKCE flow, Client Credentials
flow for CI/automation, token storage with refresh, gateway add/login/
logout commands, OIDC bearer token injection over mTLS transport,
discovery endpoint for auto-configuration.
Security: sandbox-secret scope restriction on UpdateConfig (policy
sync only), anti-spoofing header stripping, dual-auth fallthrough
from sandbox-secret to Bearer token.
Deployment: OIDC config wired through DeployOptions, Docker env vars,
Helm values/templates, HelmChart manifest, cluster-entrypoint.sh, and
bootstrap scripts. Keycloak dev server script with pre-configured
realm (test users, roles, PKCE client, CI client).
Tested with Keycloak. The roles claim path and role names are
configurable to support other OIDC providers.
* feat(auth): add OAuth2 scope-based fine-grained permissions
Add opt-in scope enforcement on top of existing OIDC role-based access
control. When --oidc-scopes-claim is set, the server extracts scopes
from the JWT and checks them per-method against an exhaustive scope map.
Scopes: sandbox:read, sandbox:write, provider:read, provider:write,
config:read, config:write, inference:read, inference:write, and
openshell:all (wildcard). Methods not in the scope map require
openshell:all. Scopes layer on top of roles and cannot escalate
privilege. Auth-only mode (empty role names) still enforces scopes
when enabled.
Server: scopes_claim in OidcConfig, scope extraction from JWT
(space-delimited and JSON array formats), standard OIDC scope
filtering, scope check in AuthzPolicy after role check.
CLI: --oidc-scopes on gateway add/start stored in metadata and
consumed by gateway login, --oidc-scopes-claim on gateway start
forwarded to server, scopes parameter in browser and client
credentials OAuth2 flows with openid deduplication.
Deployment: oidc_scopes_claim wired through DeployOptions, docker.rs,
Helm, bootstrap scripts, and cluster entrypoint.
Keycloak: realm config updated with built-in OIDC scopes and 9
OpenShell client scopes as optional on openshell-cli and openshell:all
as default on openshell-ci.
* fix(auth): address branch review findings
Add GetInferenceBundle to sandbox-secret methods so sandbox inference
route refresh works under OIDC. Make GetSandboxConfig dual-auth so CLI
users can read sandbox settings with Bearer tokens.
Preserve OIDC gateway metadata on restart — a bare gateway start
without --oidc-* flags no longer erases the stored OIDC registration.
Document CI client ID requirement (openshell-ci vs openshell-cli) in
the testing guide. Add security note about auth-only mode blast radius
for GitHub Actions.
* fix(auth): complete review findings for OIDC auth boundary
Move OpenShell/GetSandboxConfig from sandbox-secret-only to dual-auth
so CLI users can read sandbox settings with Bearer tokens while sandbox
supervisors continue using the shared secret.
Add sandbox secret interceptor to the inference bundle fetch path so
GetInferenceBundle works under OIDC-enabled gateways. Extract shared
interceptor constructor to avoid duplication.
Add GetSandboxConfig to the config:read scope map so scope enforcement
applies consistently when scopes are enabled.
Refactor OIDC metadata preservation into apply_oidc_gateway_metadata()
with explicit resume semantics — only preserve existing OIDC metadata
on real resume paths, not on fresh deployments.
Update architecture docs and testing guide to reflect the corrected
method classifications and add new test coverage for interceptor
injection, scope requirements, metadata preservation, and dual-auth
classification.
* refactor(auth): use oauth2 crate for CLI OIDC flows
Replace hand-written PKCE generation, authorization URL construction,
token exchange, client credentials, and token refresh with the oauth2
crate's typed API.
Eliminates sha2, hex, and getrandom dependencies from the CLI. The
custom urlencoded() helper and manual form POST logic are replaced by
BasicClient methods with proper type-state safety.
Discovery and the callback server remain custom since the oauth2 crate
does not provide OIDC discovery or a localhost redirect listener.
* refactor(auth): move server auth modules into auth/ directory
Group oidc.rs, authz.rs, identity.rs, and the auth HTTP endpoints
under src/auth/ module directory. No behavioral changes.
auth/mod.rs — module root, re-exports HTTP router
auth/oidc.rs — JWT validation, JWKS caching, method classification
auth/authz.rs — role and scope authorization policy
auth/identity.rs — provider-agnostic Identity type
auth/http.rs — /auth/connect and /auth/oidc-config endpoints
* fix(auth): use RequestBody auth type for client credentials flow
The oauth2 crate defaults to BasicAuth (HTTP Basic header) but Keycloak
and most OIDC providers expect client_secret_post (credentials in the
request body). Set AuthType::RequestBody explicitly to match the
pre-refactor behavior.
Also re-export Identity, IdentityProvider, and JwksCache from the auth
module so ServerState's public API remains nameable by external consumers.
* fix(auth): forward OPENSHELL_OIDC_SCOPES through cluster bootstrap
Pass --oidc-scopes to gateway start so the metadata includes requested
scopes after cluster bootstrap. Without this, users had to manually
edit metadata.json to set scopes for gateway login.
Usage: OPENSHELL_OIDC_SCOPES="openshell:all" mise run cluster
* test(auth): add OIDC e2e tests for RBAC, scopes, and client credentials
Add 10 end-to-end tests covering OIDC authentication against a live
K3s cluster with Keycloak:
RBAC (5 tests): admin can create providers, user cannot, user can list
sandboxes, unauthenticated requests rejected, health probe works
without auth.
Scopes (4 tests): sandbox-scoped token can list sandboxes but not
providers, openshell:all grants full access, no-scopes token denied.
Client credentials (1 test): CI token via client_credentials grant.
Tests are opt-in via OPENSHELL_E2E_OIDC=1 and OPENSHELL_E2E_OIDC_SCOPES=1
env vars. They derive the Keycloak URL from gateway metadata to match
the server's configured issuer.
Run with:
OPENSHELL_E2E_OIDC=1 OPENSHELL_E2E_OIDC_SCOPES=1 \
PYTHONPATH=python uv run pytest e2e/python/oidc/ -v
* fix(docs): fix markdown lint errors in OIDC architecture docs
Add blank lines before lists and fenced code blocks to satisfy
markdownlint MD031 and MD032 rules.
* ci(docker): use prebuilt Rust binaries by default
Flip Docker image builds to consume staged native Rust artifacts, remove in-Docker Rust build stages, and publish per-arch images with a manifest merge.
Add local staging support for prebuilt gateway and sandbox binaries so development image builds continue to work without CI artifacts.
Signed-off-by: Jonas Toelke <jtoelke@nvidia.com>
* ci(docker): address prebuilt build review feedback
* ci(rust): allow existing vfio complexity
* ci(rust): pin toolchain to 1.95
---------
Signed-off-by: Jonas Toelke <jtoelke@nvidia.com>
the transfer of very large sandbox images to containerd can
timeout depending on the size of the image and the speed of
the local host.
Made-with: Cursor
* fix(sandbox): add GPU device nodes and nvidia-persistenced to landlock baseline
Landlock READ_FILE/WRITE_FILE restricts open(2) on character device files
even when DAC permissions would otherwise allow it. GPU sandboxes need
/dev/nvidiactl, /dev/nvidia-uvm, /dev/nvidia-uvm-tools, /dev/nvidia-modeset,
and per-GPU /dev/nvidiaX nodes in the policy to allow NVML initialization.
Additionally, CDI bind-mounts /run/nvidia-persistenced/socket into the
container. NVML tries to connect to this socket at init time; if the
directory is not in the landlock policy, it receives EACCES (not
ECONNREFUSED), which causes NVML to abort with NVML_ERROR_INSUFFICIENT_PERMISSIONS
even though nvidia-persistenced is optional.
Both classes of paths are auto-added to the baseline when /dev/nvidiactl is
present. Per-GPU device nodes are enumerated at runtime to handle multi-GPU
configurations.
Replace manual Header::set_path() + append() with
builder.append_path_with_name() which emits GNU LongName extensions
for paths exceeding the 100-byte POSIX tar name field limit.
Fixes#705
Co-authored-by: John Myers <johntmyers@users.noreply.github.com>
The image push pipeline buffered the entire Docker image tar 3x in
memory (export, tar wrap, Bytes copy), causing OOM kills for images
over ~1-2 GB. Replace the in-memory pipeline with a temp-file +
streaming upload: export to a NamedTempFile, then stream the outer
tar (header, 8 MiB file chunks, footer) directly into
upload_to_container via body_try_stream. Peak memory drops from ~3x
image size to ~8 MiB constant.
Also adds incremental export progress reporting every 100 MiB.
* feat(bootstrap): switch GPU injection to CDI where supported
Use an explicit CDI device request (driver="cdi", device_ids=["nvidia.com/gpu=all"])
when the Docker daemon reports CDI spec directories via GET /info (SystemInfo.CDISpecDirs).
This makes device injection declarative and decouples spec generation from consumption.
When the daemon reports no CDI spec directories, fall back to the legacy NVIDIA device
request (driver="nvidia", count=-1) which relies on the NVIDIA Container Runtime hook.
Failure modes for both paths are equivalent: a missing or stale NVIDIA Container Toolkit
installation will cause container start to fail.
CDI spec generation is out of scope for this change; specs are expected to be
pre-generated out-of-band, for example by the NVIDIA Container Toolkit.
---------
Signed-off-by: Evan Lezar <elezar@nvidia.com>
Co-authored-by: Piotr Mlocek <pmlocek@nvidia.com>
* fix(cli): clear stale last-used sandbox on deletion
When a sandbox is deleted, the locally stored last-used record now gets
cleared if it matches the deleted sandbox name. This prevents subsequent
commands from falling back to a sandbox that no longer exists, which
previously caused confusing gRPC errors.
Adds clear_last_sandbox_if_matches() to openshell-bootstrap and calls
it from sandbox_delete() after each successful deletion.
Closes#172
Signed-off-by: Serge Panev <spanev@nvidia.com>
* style: fix rustfmt import ordering and line wrapping
Signed-off-by: Serge Panev <spanev@nvidia.com>
* fix(cli): pass gateway to sandbox_delete in finalize_sandbox_create_session
Signed-off-by: John Myers <johntmyers@users.noreply.github.com>
---------
Signed-off-by: Serge Panev <spanev@nvidia.com>
Signed-off-by: John Myers <johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <johntmyers@users.noreply.github.com>
Error messages and recovery hints were showing 'openshell gateway destroy <name>'
as a positional argument, but the CLI defines name as a --name flag. This caused
'unexpected argument' errors when users copied the suggested commands.
Docker Desktop 29.x defaults to private cgroupns which prevents k3s
kubelet from accessing cgroup v2 controllers (cpu, cpuset, memory,
pids, hugetlb). This causes ContainerManager to fail during startup.
Explicitly set cgroupns_mode to host, which is backwards compatible
with all Docker versions and matches what k3s-in-Docker tooling
(k3d) requires.
Closes#273
Verify inference endpoints synchronously on the server during set/update, expose a --no-verify escape hatch in the CLI and Python helper, and return actionable failures when validation does not pass.
PR #281 removed the shared openshell-cluster Docker network in favor of
the default bridge. This restores custom bridge networking but makes each
gateway use its own isolated network named openshell-cluster-{name},
matching the existing container/volume naming convention.
Changes:
- Add network_name() to constants.rs for per-gateway network naming
- Add ensure_network() with retry/backoff and force_remove_network()
parameterized by network name instead of a global constant
- Attach containers to their per-gateway network via network_mode
- Disconnect and remove the network during gateway destroy
- Wire ensure_network() into the deploy flow before ensure_volume()
- Update architecture docs to reflect per-gateway network isolation