* feat(api): expose structured gateway errors across SDKs
Refs #3051. Add standard validation, conflict, and retry details; preserve raw transport status in Rust, Go, TypeScript, and Python; document status and recovery guidance.
This is the structured-error foundation only. Mutation result shapes, allow_missing, durable request deduplication, and exec retry semantics remain follow-up work.
Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
* fix(python): preserve wrapped RPC cleanup handling
Inspect the original gRPC call when handling missing sandboxes during deletion waits and managed cleanup. Add intercepted cleanup regressions and clarify the error-wrapper migration contract.
Addresses the cleanup review on #3313; part of #3051.
Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
---------
Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
* refactor(inference): remove managed inference routes
Closes#3172
Remove the inference route control plane, inference.local data path, built-in router crate, and SDK surface. Move inference workloads to explicitly imported provider profiles and native endpoints, with migration cleanup and updated tests and documentation.
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
* fix(policy): preserve alternate upstream isolation
Restore the provider policy activation guard so legacy OpenAI and Anthropic providers configured for alternate base URLs do not grant egress to the built-in public vendor endpoints.
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
---------
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
* feat(server): add sandbox workload templates
Signed-off-by: Gordon Sim <gsim@redhat.com>
* feat(go-sdk): add sandbox workload template support
Signed-off-by: Gordon Sim <gsim@redhat.com>
* feat(rust-sdk): add sandbox workload template support
Signed-off-by: Gordon Sim <gsim@redhat.com>
* feat(python-sdk): add sandbox workload template support
Signed-off-by: Gordon Sim <gsim@redhat.com>
* feat(typescript-sdk): add sandbox workload template support
Signed-off-by: Gordon Sim <gsim@redhat.com>
* docs(agents): document sandbox workload templates
Signed-off-by: Gordon Sim <gsim@redhat.com>
* fix(cli): support default GPU requests in sandbox templates
Signed-off-by: Gordon Sim <gsim@redhat.com>
* fix(server): cap sandbox templates per workspace
Signed-off-by: Gordon Sim <gsim@redhat.com>
* docs(architecture): document sandbox workload template boundaries
Signed-off-by: Gordon Sim <gsim@redhat.com>
* feat(cli+sdk): expose sandbox workload template provenance
Signed-off-by: Gordon Sim <gsim@redhat.com>
* fix(cli): include sandbox template annotations in output
Signed-off-by: Gordon Sim <gsim@redhat.com>
* feat(sandbox): add label selectors to template listing
Signed-off-by: Gordon Sim <gsim@redhat.com>
* test(e2e): cover sandbox template failure paths
Signed-off-by: Gordon Sim <gsim@redhat.com>
* fix(server): preserve command and ttl when creating sandbox from template
Signed-off-by: Gordon Sim <gsim@redhat.com>
* test(server): add field coverage test for template merge
Signed-off-by: Gordon Sim <gsim@redhat.com>
* fix(go-sdk): add pagination support to fake client
Signed-off-by: Gordon Sim <gsim@redhat.com>
* fix(cli): warn on env vars that looks like secrets
Signed-off-by: Gordon Sim <gsim@redhat.com>
* fix(sdk-ts): propagate sandbox workspace through lifecycle calls
Signed-off-by: Gordon Sim <gsim@redhat.com>
* fix(docs): update workspace management docs
Signed-off-by: Gordon Sim <gsim@redhat.com>
* fix(ts-sdk): support command and tty when creating from template
Signed-off-by: Gordon Sim <gsim@redhat.com>
* fix(go-sdk): support command and tty when creating from template
Signed-off-by: Gordon Sim <gsim@redhat.com>
* test(python-sdk): verify command and tty handling when creating from template
Signed-off-by: Gordon Sim <gsim@redhat.com>
* fix(go-sdk): update docs and ClientInterface
Signed-off-by: Gordon Sim <gsim@redhat.com>
* fix(server): validate sandbox create specs before I/O
Signed-off-by: Gordon Sim <gsim@redhat.com>
* fix(go-sdk): guard empty DNS-1123 label validation
Signed-off-by: Gordon Sim <gsim@redhat.com>
* fix(python-sdk): allow empty template builder mappings
Signed-off-by: Gordon Sim <gsim@redhat.com>
* fix(cli): align template GPU JSON default output
Signed-off-by: Gordon Sim <gsim@redhat.com>
---------
Signed-off-by: Gordon Sim <gsim@redhat.com>
The Maturin-based wheel packaging was a historical remnant from when the local gateway launch path and OpenShell CLI were coupled in one binary. The gateway and CLI now ship as standalone artifacts, so the Python distribution should contain only the SDK.
Build a single platform-independent setuptools wheel, verify that it cannot contain native code or an openshell entry point, and simplify the release jobs and documentation for SDK-only PyPI installs.
Signed-off-by: Simon Scatton <sscatton@nvidia.com>
* feat(workspace): implement workspace model (Phase 1 of RFC 0011)
Implements workspace and membership model providing hard isolation
boundaries for multi-player OpenShell deployments.
Workspace CRUD with Kubernetes-style Terminating phase for graceful
deletion. All resources scoped by workspace via ObjectMeta. Membership
RPCs for workspace access control. Persistence migration shifts name
uniqueness to (object_type, workspace, name). Provider profiles support
platform and workspace scoping. Service routing uses workspace-prefixed
DNS labels. Inference routes renamed and workspace-scoped with
DeleteInferenceRoute RPC. Python SDK with WorkspaceClient, two-method
list pattern (workspace-scoped and for_all_workspaces), and workspace
parameter on all methods. CLI workspace flags, TUI workspace cycling.
K8s driver filters unmanaged CRs and uses delete preconditions. Podman
driver uses immutable container IDs. Label serialization fixed across
all put_if call sites.
Signed-off-by: Derek Carr <decarr@redhat.com>
* fix(cli): delegate sandbox upload command to existing upload function
The standalone `sandbox upload` command reimplemented upload logic
inline with two bugs: it used `Path::exists()` which follows symlinks
(rejecting dangling symlinks), and it ran git-aware filtering on
symlink sources. The `run::sandbox_upload()` function already handles
both cases correctly via `sandbox_upload_plan()`. Replace the inline
logic with a call to the existing function.
Signed-off-by: Derek Carr <decarr@redhat.com>
* fix(e2e): shorten sandbox names and fix test compatibility
Shorten the sandbox name in initial_sparse_policy_is_acknowledged_as_loaded
from 'e2e-2159-sparse-enrich' (22 chars) to 'e2e-sparse-enrich' (17 chars)
to comply with MAX_ROUTABLE_NAME_LEN (19 chars).
Also capture stderr in create_keep_with_args so future sandbox creation
failures include the actual CLI error instead of reporting empty output.
Signed-off-by: Derek Carr <decarr@redhat.com>
* test(workspace): add test coverage for workspace CRUD and persistence isolation
Add unit tests for workspace create happy path, get round-trip, get
not-found, get empty-name rejection, already-exists error, and
resolve_workspace not-found. Add persistence test proving cross-workspace
name uniqueness (same name in different workspaces produces separate
records). Add workspace name max-length boundary tests. Fix e2e harness
to include stderr in name-parse-failure error path. Align Python e2e
test_workspace_crud with try/finally pattern. Document provider profile
catalog workspace scoping gap in RFC 0011.
Signed-off-by: Derek Carr <decarr@redhat.com>
* fix(examples): update examples for workspace model compatibility
Shorten sandbox names in demo scripts to fit the 19-character
MAX_ROUTABLE_NAME_LEN limit: policy-demo prefix to pd-, multi-agent
notepad derives a short SANDBOX_TAG from the run ID, governance
interceptor uses gs-PID-RANDOM. Update vscode-remote-sandbox.md SSH
host aliases from openshell-{name} to openshell-{name}.{workspace}
format.
Signed-off-by: Derek Carr <decarr@redhat.com>
* feat(sdk): add workspace-scoped client and workspace CRUD
Add WorkspaceScopedClient modeled after kube::Api::namespaced — captures
workspace once and injects it into every sandbox request. Add workspace
CRUD methods (create, get, list, delete) and list_sandboxes_all_workspaces
on OpenShellClient. Extend SandboxRef with workspace field and add
WorkspaceRef type. Include mock tests for all new operations.
Signed-off-by: Derek Carr <decarr@redhat.com>
* fix(lint): resolve clippy warnings in workspace test assertions
Signed-off-by: Derek Carr <decarr@redhat.com>
* fix(docs): convert indented code blocks to fenced in RFC 0011
Signed-off-by: Derek Carr <decarr@redhat.com>
* fix(lint): resolve clippy warnings and apply cargo fmt across workspace
Auto-format with cargo fmt and fix clippy warnings exposed by the
reformat: unnecessary qualifications, map_unwrap_or, identical match
arms, unused variable prefix, dead code annotations, and let-unit-value
in e2e harness.
Signed-off-by: Derek Carr <decarr@redhat.com>
* fix(workspace): address workspace scoping issues from review
- Add workspace field to settings JSON output (CLI)
- Skip Podman containers missing workspace label instead of defaulting
to empty string, matching K8s driver behavior
- Add resource_version to list_by_scope SELECT in both SQLite and
Postgres backends, with regression test
- Gate PolicyLocalContext proposal/lookup routes on workspace readiness,
returning 503 when workspace is not yet discovered
- Block sandbox and provider creation in TUI all-workspaces mode
- Clear workspace vectors in TUI reset_sandbox_state
Signed-off-by: Derek Carr <decarr@redhat.com>
* fix(workspace): make provider profile catalog workspace-aware
Thread workspace through snapshot_catalog so the
EffectiveProviderProfileCatalog enforces workspace boundaries on both
read and write paths. UserProviderProfileSource now loads platform-scoped
profiles (workspace "") plus the target workspace's profiles, preventing
cross-workspace duplicate profile ID collisions that previously caused
global catalog failures.
Update RFC 0011 to reflect catalog scoping is implemented in Phase 1
rather than deferred to future work.
Signed-off-by: Derek Carr <decarr@redhat.com>
* fix(persistence): include workspace column in atomic policy revision INSERT
put_policy_revision_atomic omitted the workspace column from the INSERT
into the objects table in both SQLite and Postgres backends, causing
atomically-written policy revisions to lose their workspace association.
Add workspace field to AtomicPolicyRevisionWrite and thread it through
both backend INSERT statements, matching the non-atomic put_policy_revision
path which already included it.
Signed-off-by: Derek Carr <decarr@redhat.com>
* fix(proxy): skip ancestor walk when socket owner is the entrypoint
collect_ancestor_identities walked the entire process tree above the
entrypoint when the connecting process was the entrypoint itself,
SHA256-hashing every ancestor binary (IDE, shell, container runtime).
On dev machines with large binaries in the ancestor chain this exceeded
the 30-second test timeout. When start_pid == stop_pid there are no
intermediate ancestors to verify, so return an empty list immediately.
Signed-off-by: Derek Carr <decarr@redhat.com>
* fix(workspace): make provider profile catalog scope-aware
Allow the same profile ID at platform and workspace scopes by
introducing layered catalog entries where workspace profiles shadow
platform profiles. Add source and scope fields to the ProviderProfile
proto and CLI output. Migrate List/Get handlers to the catalog,
fixing divergence with runtime profile resolution.
Signed-off-by: Derek Carr <decarr@redhat.com>
* fix(e2e): align podman e2e labels with centralized driver constants
The podman driver moved its container labels to the centralized
openshell.ai/ prefix, but the e2e test harness and cleanup script
still referenced the old openshell.sandbox-* keys, causing the
local_driver_token_restart test to fail on container lookup.
Signed-off-by: Derek Carr <decarr@redhat.com>
* fix(e2e): align python profile isolation test with scope-aware catalog
Platform profiles are now visible in workspace listings as fallbacks
per the layered catalog design. Update the assertion to match.
Signed-off-by: Derek Carr <decarr@redhat.com>
* fix(workspace): honor profile_workspace in runtime profile resolution
Runtime profile lookups now consult provider.profile_workspace via
get_type_profile_for_scope. Providers created with --global-profile
(profile_workspace="") resolve to the platform profile even when a
workspace profile shadows the same ID. All 6 runtime call sites
updated; type-only call sites remain scope-agnostic.
Signed-off-by: Derek Carr <decarr@redhat.com>
---------
Signed-off-by: Derek Carr <decarr@redhat.com>
* fix(sandbox): acknowledge initial policy revision
The supervisor loaded and enforced a sandbox-scoped policy but never told
the gateway which revision it loaded. The policy poll loop seeded itself
with the initial revision's hash on its first poll, so `policy_changed`
was never true for that revision and `ReportPolicyStatus(LOADED)` — which
only ran in the hot-reload branch — was never called. The revision stayed
`Pending` and `current_policy_version` stayed 0 even though the sandbox was
`Ready` and the policy was effective. This was most visible with sparse
policies that get baseline-enriched into a new revision during startup.
After the OPA engine is constructed, report the exact sandbox revision the
supervisor loaded as LOADED, and seed the poll loop from that revision so
it is not re-reported. Report FAILED with the original construction error
if engine construction or conversion fails. Only sandbox-sourced revisions
(version > 0) whose canonical content matches the loaded policy are
acknowledged; global and local-file policies are untouched. Delivery uses
the shared bounded retry, is non-fatal on transient failure, and a pending
initial acknowledgement is delivered before any newer revision so policy
history is never reordered.
Signed-off-by: Kyle Zheng <kyzheng@nvidia.com>
* feat(python): expose sandbox labels and selectors
The gateway protobuf and CLI already support request-level sandbox labels
(`CreateSandboxRequest.name`/`labels`) and selector-based listing
(`ListSandboxesRequest.label_selector`), but the public Python SDK dropped
them, so Python-created sandboxes could not be found via
`openshell sandbox list --selector ...`.
Add optional, source-compatible `name`/`labels` to `SandboxClient.create`,
`create_session`, and the high-level `Sandbox`, and `label_selector` to
`list`/`list_ids`. `SandboxRef` now carries the gateway labels as an
immutable mapping (default empty, so `SandboxRef(id, name, status)` still
works). Caller-provided label mappings are copied. Attaching the high-level
`Sandbox` to an existing sandbox rejects `name`/`labels` since creation
metadata cannot change on attach. Template labels remain a separate concept.
No protobuf changes are required.
Signed-off-by: Kyle Zheng <kyzheng@nvidia.com>
* fix(python): keep SandboxRef hashable and copy high-level labels
Excluding the new immutable `labels` field from SandboxRef equality/hash
(`compare=False`) preserves the original (id, name, status) identity and keeps
the frozen dataclass hashable — a MappingProxyType field would otherwise make
`hash(SandboxRef(...))` raise. Also defensively copy caller-provided labels in
the high-level `Sandbox` so later caller mutation cannot change what is sent.
Signed-off-by: Kyle Zheng <kyzheng@nvidia.com>
* fix(sandbox): bound initial-policy-ack retries
The poll loop retried a pending initial acknowledgement before processing any
newer revision, but retried unconditionally forever. A permanently
undeliverable ack (e.g. the revision was superseded before it could be
reported) would then stall all later policy hot-reloads and provider-env
refreshes. Cap the retries; after the bound, give up and resume normal polling
so the loop cannot livelock on a stuck acknowledgement.
Signed-off-by: Kyle Zheng <kyzheng@nvidia.com>
* test(sandbox): add sparse-policy revision-2 acknowledgement e2e
Regression for #2159: create a sandbox with the network-only policy-advisor
fixture, which the supervisor enriches with baseline filesystem paths during
startup (creating revision 2, superseding revision 1). Assert the effective
policy reaches revision 2 and no revision remains Pending once the supervisor
acknowledges the load. Adds SandboxGuard::create_keep_with_args to create a
kept sandbox with an initial --policy.
Signed-off-by: Kyle Zheng <kyzheng@nvidia.com>
* fix(ci): correct sandbox checks
Signed-off-by: Kyle Zheng <kyzheng@nvidia.com>
* test(cli): serialize mTLS environment access
Signed-off-by: Kyle Zheng <kyzheng@nvidia.com>
* fix(sandbox): address policy review feedback
Signed-off-by: Kyle Zheng <kyzheng@nvidia.com>
* fix(sandbox): preserve exact policy acknowledgements
Signed-off-by: Kyle Zheng <kyzheng@nvidia.com>
* fix(sandbox): preserve local policy overrides
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
---------
Signed-off-by: Kyle Zheng <kyzheng@nvidia.com>
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <9696606+johntmyers@users.noreply.github.com>
* fix(python): add encoding=utf-8 to all file reads and writes in sandbox.py
read_text(), fdopen() without encoding use the system locale, which can
differ on non-UTF-8 systems. All config files (metadata.json,
oidc_token.json, active_gateway) are UTF-8 text.
* fix(python): catch UnicodeDecodeError in _read_oidc_token_bundle, add encoding tests
Corrupt or non-UTF-8 oidc_token.json now returns None consistently.
Add regression tests for non-ASCII UTF-8 paths in metadata, oidc token,
and active_gateway files.
* fix(python): use ensure_ascii=False in encoding tests, add from_active_cluster UTF-8 bytes test
* fix(python): assert non-ASCII cluster name and endpoint in encoding test
* fix(python): wrap long write_bytes line for ruff format
* fix(python): apply ruff format to sandbox_test.py
* fix(gateway): align package TLS bootstrap path
Closes#1593
Default package-managed gateway services to a stable local TLS directory and use that same value for certificate generation and runtime startup.
Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
* test(packaging): validate package asset paths exist
Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
* ci(e2e): pin mise in kubernetes job
Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
---------
Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
* feat(python-sdk): support OIDC Bearer auth on SandboxClient
PR #1596 hardened the gateway side of the OIDC story; the Python SDK
was the remaining gap — it only supported plaintext or mTLS, with no
Bearer metadata anywhere. Deployments with OIDC enabled (the
recommended posture since PR #935 / PR #1404) were unreachable from
the SDK.
Adds:
- `bearer_token: str | Callable[[], str] | None` kwarg on
`SandboxClient`. Static strings or zero-arg callables (the latter
is invoked per RPC, so callers can drop in a refresh loop or
token-file watcher without reconstructing the client). Composes
with `tls` for OIDC-over-mTLS deployments.
- `_BearerAuthInterceptor` implementing all four
`grpc.{Unary,Stream}{Unary,Stream}ClientInterceptor` types.
Appends `authorization: Bearer <token>` to outgoing metadata.
Implemented as an interceptor (not call credentials) so it works
on both plaintext (`disableTls=true` dev) and TLS channels without
`grpc.composite_channel_credentials`.
- `TlsConfig` ergonomics: all three fields (`ca_path`, `cert_path`,
`key_path`) are now optional with `cert_path` / `key_path`
required-together-or-not-at-all (enforced in `__post_init__`). This
unlocks three transport profiles from one dataclass:
* full mTLS (all three)
* CA-only trust (`ca_path` only)
* system roots (`TlsConfig()` — for OIDC gateways behind a
public CA)
- `from_active_cluster` mirrors `crates/openshell-tui/src/lib.rs`
`build_oidc_channel`:
* For any `https://` gateway, always build a secure channel.
Pick the strongest TLS profile available in `mtls/` (full
mTLS → CA-only → system roots). No more `insecure_channel`
fallback for HTTPS.
* Gate OIDC bearer attachment on
`metadata.json["auth_mode"] == "oidc"`. Matches
`crates/openshell-cli/src/main.rs:132` and the TUI; a stale
`oidc_token.json` next to a non-OIDC gateway no longer causes
the SDK to attach a bearer.
- `_OidcRefresher` — thread-safe, in-process native OAuth2 refresh
modeled on `google.oauth2.credentials.Credentials` and
`botocore.tokens.SSOTokenProvider`. Lazily checks expiry on every
RPC; when stale, re-reads disk first (the CLI may have rotated
the bundle), and only then exchanges the refresh_token against
the IdP's token endpoint discovered via OIDC discovery
(`/.well-known/openid-configuration`, cached after first call).
Concurrent RPCs share a single refresh via `threading.Lock` (no
IdP stampede). Honors refresh-token rotation. Surfaces IdP
failures as `SandboxError` with the RFC 6749 error body included
for diagnostics.
Mirrors the Rust CLI's HTTP-policy posture from
`crates/openshell-cli/src/oidc_auth.rs`:
* `follow_redirects=False` so a 3xx during discovery can't
steer us to an attacker-controlled token endpoint.
* Discovery `issuer` is validated against the configured
issuer; a discovery document claiming a different issuer is
rejected, preventing the SDK from POSTing the refresh_token
to a malicious endpoint.
* `insecure: bool` flag plumbed through to httpx's
`verify=` so self-signed-cert deployments work the same way
they do in the Rust CLI.
Built on `httpx` (chosen over `urllib` specifically for
follow_redirects + verify control as kwargs). The OAuth2
refresh-token grant itself (RFC 6749 §6) is one form-encoded
POST — handled inline rather than via a dedicated OAuth library;
tried `authlib`'s `OAuth2Client` first but it auto-injects an
Authorization header on every request, which breaks the
unauthenticated discovery GET.
- `_make_cluster_bearer_provider(..., auto_refresh=True,
write_back=True, insecure=False)` factory. Defaults to the
refresher path with write-back enabled; `auto_refresh=False`
falls back to the read-only fail-closed behavior for callers that
don't want the SDK to make outbound HTTP calls to the IdP.
`write_back=True` is the default (changed from the first round of
review): IdPs with refresh-token rotation (Keycloak with
rotation, Entra in strict mode) invalidate the old refresh_token
on each refresh, so an in-memory-only refresh would leave the
on-disk bundle pointing at an invalidated value — any second
process starting from disk would `invalid_grant`. With write-back
enabled by default, the SDK keeps the shared cache consistent
with the IdP.
- `from_active_cluster` exposes `auto_refresh`, `write_back`, and
`insecure` kwargs (defaults: True / True / False). The
high-level `Sandbox` context manager surfaces the same three
kwargs and forwards them through, so callers using the wrapper
have parity with `SandboxClient` for OIDC-protected gateways.
- `SandboxClient.close()` chains to a `_bearer_close` hook so the
`_OidcRefresher`'s underlying `httpx.Client` is released
deterministically instead of leaking sockets/FDs until GC runs
`__del__`. Idempotent.
- `_OidcRefresher._write_to_disk` uses `tempfile.mkstemp` (PID +
random suffix) instead of a fixed `.oidc_token.json.tmp` path,
so two writers racing on the same gateway directory don't
trample each other's tmp content. Success path atomically
replaces; failure path unlinks the orphan.
OAuth2 refresh policy and write-back semantics deliberately mirror
what the major Python SDKs do — see
github.com/googleapis/google-auth-library-python (`Credentials`)
and github.com/boto/botocore (`SSOTokenProvider`):
| Library | Native refresh | Writes back |
|-------------------------------|----------------|-------------|
| google-auth Credentials | yes | no |
| botocore SSOTokenProvider | yes | yes |
| openshell SandboxClient (here)| yes (opt-out) | yes (opt-out)|
OpenShell sits between the two; chose write-back-by-default because
the rotation invariant matters more for our deployments than the
"CLI is the only writer" assumption that fits google-auth.
Adds `httpx>=0.27` as a runtime dependency. No new OAuth2 library —
the refresh grant is a single POST.
Tested:
- 42 sandbox_test.py tests pass (5 pre-existing + 37 new across
the bearer interceptor, fail-closed provider, refresher
behavior, TlsConfig validation, from_active_cluster auth ladder,
security-review regressions, Sandbox-wrapper kwarg forwarding,
and lifecycle / concurrency probes).
`mise run test:python` → 47 passed total across the python
suite.
- `mise run python:lint` (ruff) clean.
- End-to-end against a Keycloak-protected gateway on OpenShift:
* unauthenticated `Health` bypass works
* admin + `openshell:all` reaches user-callable methods
* reader (`sandbox:read`) denied on `CreateSandbox` by scope
* admin + `openshell:all` denied on PR #1596 sandbox-only
methods at the router (the new gate is honored from the SDK)
* full provider CRUD lifecycle via the SDK
* callable token provider rotates per RPC as expected
- Regression-probed against three pre-review security findings:
* **Discovery issuer validation** — a discovery document
claiming a different `issuer` than the configured one is
rejected with a clear `SandboxError` before any refresh POST
can reach the attacker-controlled endpoint.
* **Redirect during discovery** — `follow_redirects=False` on
the underlying httpx client means a 3xx during discovery
surfaces as a SandboxError rather than silently chasing the
redirect.
* **Cross-process rotation** — a two-process simulation shows
process B starting from disk and successfully refreshing
with the rotated refresh_token, because process A's
write-back updated the shared cache.
- Refresher unit tests cover: cached-fresh fast path, disk-rotated
re-read before refresh, OAuth2 exchange against the discovered
token endpoint, refresh-token rotation, atomic write-back at
0600 mode (default), default-on write_back proven by test,
concurrent N-thread coordination (one refresh shared across 8
threads), IdP failure surfaced with the error body, the
client_credentials / no-refresh_token error path, issuer-
mismatch rejection, redirect-during-discovery rejection,
insecure flag plumbing.
- Lifecycle / concurrency regression tests added: `close()`
invokes the `_bearer_close` hook (idempotent), the refresher's
`httpx.Client` is marked closed after `SandboxClient.close()`,
and 16 concurrent writers don't leave orphan tmp files behind
while producing a valid final bundle. The `Sandbox` wrapper has
direct forwarding tests proving `auto_refresh`, `write_back`,
and `insecure` reach `from_active_cluster` (both explicit
values and defaults).
- End-to-end against a real OpenShift + Keycloak cluster from
inside a pod: real OIDC discovery against
`keycloak.keycloak.svc.cluster.local:8080`, refresh-token grant
POST, atomic write-back of the rotated bundle at 0600, and a
follow-up RPC reusing the freshly-rotated in-memory token —
full round-trip in ~170ms.
Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
* fix(python-sdk): adopt newer on-disk OIDC bundle before refreshing
_OidcRefresher.current_access_token() only adopted the on-disk
oidc_token.json when its access token was still fresh; otherwise it
refreshed using the in-memory bundle. With refresh-token rotation
enabled (Keycloak with rotation, Entra strict mode), this let a process
keep using an invalidated refresh_token:
1. Process A holds a stale in-memory bundle with refresh_token=r1.
2. Process B refreshes first and writes a rotated (r2) but now
near-expiry bundle to disk.
3. Process A re-reads disk, sees the access token is not fresh, ignores
the disk bundle, and POSTs the stale r1 — which the IdP has already
invalidated, yielding invalid_grant.
Fix: when the cached bundle is stale, adopt the on-disk bundle if it was
refreshed more recently than ours, even when its access token is also
stale. "More recently" is decided by expires_at — a refresh mints a new
access token with a forward expiry alongside the rotated refresh_token,
so the later expiry carries the newest refresh_token. Comparing by
expiry (rather than unconditionally preferring disk) preserves the
write_back=False case, where the in-memory bundle has already rotated
past the on-disk copy and must not be clobbered. When the adopted
bundle's issuer differs, the cached token endpoint is reset so the
refresh re-discovers against the new issuer.
Adds regression tests for the cross-process rotation race and the
issuer-change re-discovery path.
* fix(python-sdk): recover from invalid_grant on lost rotation race
The expiry-based disk re-read narrows but does not fully close the
cross-process refresh-token rotation race: two processes sharing a
gateway directory can both enter their refresh window, both POST their
copy of the refresh_token, and with rotation enabled the IdP invalidates
the loser's token (invalid_grant). Neither google-auth nor botocore
close this window without an OS file lock; a Python-only flock would not
coordinate with the Rust CLI/TUI that also write oidc_token.json, so
locking is not worth its cost here.
Recover instead of prevent: distinguish an OAuth2 invalid_grant (the
refresh_token was rejected) from transport/5xx failures via a private
_InvalidGrantError, and on invalid_grant re-read oidc_token.json once. If
a peer wrote a different refresh_token (it won the race), adopt and retry
with it — returning early if it is already fresh — so the loser succeeds
transparently instead of forcing a re-authenticate. If disk offers no new
token, the rejection is genuine and surfaces the re-authenticate hint as
before. The retry is single-shot; a second invalid_grant propagates.
Adds tests for the peer-rotation recovery and the genuine-rejection
(no-retry) paths.
---------
Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
* refactor(proto): move phase and current_policy_version into SandboxStatus
Move phase and current_policy_version from SandboxSpec into
SandboxStatus to correctly model mutable runtime state. Update all
callers in the gateway server, TUI, and Python SDK to read and write
these fields through SandboxStatus accessors.
Signed-off-by: Derek Carr <decarr@redhat.com>
* fix(server): preserve sandbox status on statusless driver updates
When a driver update arrives without a status payload (e.g. before
Kubernetes populates the status subresource), preserve the stored
phase, conditions, and current policy version instead of resetting
them. Adds a regression test covering the edge case.
Signed-off-by: Derek Carr <decarr@redhat.com>
---------
Signed-off-by: Derek Carr <decarr@redhat.com>
* ci: pin azure/setup-helm and helm/kind-action to commit SHAs
* chore(python): add py.typed marker for PEP 561 compliance
* ci: use full semver in pinned action version comments
Signed-off-by: mesutoezdil <mesudozdil@gmail.com>
---------
Signed-off-by: mesutoezdil <mesudozdil@gmail.com>
* feat(gateway): add TOML configuration file (RFC 0003)
Introduces an opt-in --config / OPENSHELL_GATEWAY_CONFIG flag that loads a
TOML file with gateway-wide settings and per-driver tables. Source
precedence is CLI > env > file > built-in default, implemented via clap's
ValueSource so existing flags and env vars keep their priority.
Driver crates (kubernetes, docker, podman, vm) now derive Deserialize on
their config structs. SupervisorSideloadMethod gains Deserialize with
kebab-case rename. A per-driver inheritance allowlist on the loader side
overlays [openshell.gateway] shared defaults (default_image,
supervisor_image, image_pull_policy, guest_tls_*, ssh_handshake_skew_secs,
client_tls_secret_name, host_gateway_ip, enable_user_namespaces) onto
each [openshell.drivers.<name>] table before deserialization.
The Helm chart renders a new gateway-config ConfigMap and mounts it at
/etc/openshell/gateway.toml. The migrated OPENSHELL_* env entries are
dropped from the StatefulSet — only the Secret-backed
OPENSHELL_SSH_HANDSHAKE_SECRET remains. database_url stays on --db-url.
Adds examples/gateway/gateway.example.toml and updates architecture/gateway.md
with the source precedence and inheritance rules.
* docs(gateway): drop ssh_handshake_skew_secs and ssh_handshake_secret from examples
Both fields are scheduled for removal. Remove the example values and the
env-only note so the gateway.toml example and the architecture doc stop
recommending settings that will not exist much longer.
* docs(rfc): correct OPENSHELL_CONFIG to OPENSHELL_GATEWAY_CONFIG in RFC 0003
* docs(gateway): add per-driver TOML example configurations
Adds focused single-driver examples next to the comprehensive
gateway.example.toml: kubernetes, docker, podman, and microvm. Each one
demonstrates the realistic settings for that driver plus how shared
[openshell.gateway] defaults inherit into the driver table.
A new unit test (`checked_in_examples_parse`) loads every example through
the config_file loader so schema drift fails CI rather than silently
shipping a broken example.
* refactor(gateway): drop image_pull_policy from shared inheritance
Kubernetes and Podman use mutually-incompatible vocabularies for the same
TOML key:
- Kubernetes: `Always | IfNotPresent | Never` (free-form string passed
verbatim to the K8s API).
- Podman: `always | missing | never | newer` (strict lowercase enum
deserialised into `ImagePullPolicy`).
No value means the same thing in both drivers. Sharing the key at
`[openshell.gateway]` scope and inheriting it into every active driver's
table meant any value safe for one driver was either wrong or silently
dropped for the other (`IfNotPresent` → `ImagePullPolicy::Missing` after
`.unwrap_or_default()`). Operators run one driver per gateway, so the
"shared default" never pays for itself.
Make `image_pull_policy` driver-local:
- Remove the field from `GatewayFileSection` and from
`inheritable_keys()` for both Kubernetes and Podman.
- Drop the file→`RunArgs` merge for the gateway-scope key.
- Stop unconditionally clobbering the driver value with
`config.sandbox_image_pull_policy` in the runtime wiring — only apply
the CLI/env override when it was set (and, for Podman, only when it
parses into the lowercase enum).
- Move the key under `[openshell.drivers.kubernetes]` and
`[openshell.drivers.podman]` in every example, the RFC, the
architecture doc, and the Helm-rendered gateway ConfigMap.
The supervisor pull policy follows the same shape: it is K8s-only and
moves into `[openshell.drivers.kubernetes]` alongside `image_pull_policy`
in the Helm template.
* fix(gateway): address review feedback on TOML configuration
Resolves the P1 and P2 issues raised in PR #1317:
- Helm gateway ConfigMap moves `grpc_endpoint` under
`[openshell.drivers.kubernetes]` so the default install no longer fails
the gateway's `deny_unknown_fields` schema check.
- `kubernetes_config_from_file` and `podman_config_from_file` only let
the gateway-wide CLI/env `grpc_endpoint` overwrite the driver-table
value when it was actually supplied, preserving file-only configs.
- Kubernetes driver default `image_pull_policy` is now empty (was Podman
vocabulary "missing"), so default deployments let the Kubernetes API
apply its own policy instead of being rejected.
- New `disable_tls` gateway field plumbs `.Values.server.disableTls`
through the TOML ConfigMap instead of relying on env vars dropped from
the StatefulSet.
- StatefulSet pod template now carries a `checksum/gateway-config`
annotation so `helm upgrade` rolls pods when the ConfigMap changes.
- Auxiliary listener resolution preserves the full `SocketAddr` from
`health_bind_address` / `metrics_bind_address`, so a loopback-pinned
health port is not silently relocated onto the public bind address.
- `ssh_session_ttl_secs` from the file is now applied to `Config` (it
was previously accepted by the loader but never read).
New regression coverage: cli-level merge tests for the new fields plus
helm-unittest assertions for the ConfigMap shape, checksum annotation,
and `disable_tls` rendering.
* docs(gateway): consolidate gateway TOML examples into docs reference
Replaces the per-driver example files under examples/gateway/ with a
single published reference page at docs/reference/gateway-config.mdx
covering source precedence, layout, the full example, and the four
per-driver examples (Kubernetes, Docker, Podman, microVM). Drew flagged
during PR #1317 review that the examples belong with the user-facing
docs rather than in a sibling examples/ directory.
The cross-references in architecture/gateway.md and RFC 0003 are updated
to point at the new docs page; the round-trip test in config_file.rs is
removed (schema coverage stays on the inline parses_full_example test
and per-field merge tests — doc-snippet drift belongs in a separate
docs-lint, not in a cross-tree Rust unit test).
* refactor(core): move DEFAULT_K8S_NAMESPACE into K8s driver
The constant is Kubernetes-specific (used only by KubernetesComputeConfig's
Default impl) and does not belong in openshell-core. Relocate it to the
driver crate that owns the K8s vocabulary; openshell-core retains only
truly cross-cutting defaults.
* refactor(core): move Podman bridge default into Podman driver
DEFAULT_NETWORK_NAME is Podman vocabulary, consumed only by the Podman
driver. Also drops the unused DEFAULT_IMAGE_PULL_POLICY constant.
* docs(auth): scrub remaining SSH handshake secret references
Sweeps the trailing mentions left after the rebase: the gateway
config-file module doc, the Helm gateway-config ConfigMap header,
the gateway-config.mdx env-only note, and the RPM systemd unit
comment for init-gateway-env.sh.
* docs(gateway): clarify OPENSHELL_GRPC_ENDPOINT applies to all drivers
The previous comment implied the callback endpoint was Kubernetes-only,
but the value is propagated to every compute driver (Kubernetes, Docker,
Podman, VM) and must be reachable from wherever the sandbox runs.
* refactor(gateway): move driver options into config (#1394)
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* fix(e2e): regenerate gateway config via TOML for docker + podman harnesses
The gateway CLI flags moved into TOML config tables in 560550d2 (#1394),
which made every existing e2e/with-{docker,podman}-gateway.sh invocation
fail with "unexpected argument '--sandbox-namespace'" (and a long tail
of similar driver-specific options) before the gateway could even bind.
Replace the obsolete CLI flags with a synthesized
`[openshell.drivers.<driver>]` table written to `${STATE_DIR}/gateway.toml`
and passed via the new `--config` flag. Only the gateway-wide flags
that survived 560550d2 (bind-address, port, drivers, db-url, tls-*,
disable-tls, log-level, health-port) stay on the command line.
Both scripts get a small `toml_string` helper to properly TOML-quote
the values (the previous `%q` printf format produced bash-escape, not
TOML-escape). The Docker harness also corrects two field names that
diverged from the driver schema: `docker_network_name` →
`network_name`, and the supervisor binary/image plumbing now reads
through to `supervisor_bin` / `supervisor_image` in the same table.
The Podman harness drops `--ssh-gateway-port` (deleted in 560550d2 —
gRPC + SSH are multiplexed on the same port now) and substitutes
`network_name` + `gateway_port` for the obsolete `--sandbox-namespace`
(which the Podman driver never had as a typed field).
* fix(core): swap bind-only 0.0.0.0 SSH gateway host for cluster URL host
CLI's resolve_ssh_gateway treated 0.0.0.0 as a loopback "keep as-is"
when the cluster URL was also loopback, so the SSH proxy connected to
0.0.0.0:port. The unspecified address is never a valid connect target
and is not present in any TLS cert SAN, which produced BadCertificate
TLS handshake failures during `openshell sandbox create -- ...` in
docker/podman e2e (e.g. bypass_detection).
Resolution: when the server returns 0.0.0.0 or :: as the gateway host
and both endpoints are loopback, fall back to the cluster URL's host
(which the CLI is already using to reach the gateway, so it must
resolve and match the cert).
* fix(e2e): repair podman harness on macOS
Podman 5.x with the applehv/libkrun provider no longer creates the legacy
~/.local/share/containers/podman/machine/podman.sock symlink, and
`podman system service` is a Linux-only subcommand — the macOS client
delegates the API service to the VM. Both assumptions in the harness
were stale, so the script tried to start a temporary service that podman
rejected with "unknown flag: --time".
- Discover the macOS socket via `podman machine inspect` instead of the
hardcoded path.
- On Darwin, fail fast with a "start podman machine" message rather than
attempting the Linux-only `podman system service` fallback.
- Write socket_path into [openshell.drivers.podman] so the in-process
driver picks up the discovered socket; the driver reads TOML only
after the config refactor (560550d2), so OPENSHELL_PODMAN_SOCKET alone
was no longer enough.
* fix(server): use clone_from for TLS client CA assignment
clippy 1.95.0 rejects assigning the result of `Clone::clone()` to an
existing variable under `-D warnings` (`assigning_clones`). Switch to
`clone_from(&...)` to satisfy the lint and avoid the redundant
allocation.
---------
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
Closes#273
Verify inference endpoints synchronously on the server during set/update, expose a --no-verify escape hatch in the CLI and Python helper, and return actionable failures when validation does not pass.