* feat(sandbox): default to official Alpine sandbox image
default_sandbox_image() now returns docker.io/library/alpine:3.22, a generic
version-qualified official image, so a fresh install no longer depends on the
community sandbox image catalog. All compute drivers (docker, podman,
kubernetes, vm) inherit this fallback.
Part of #3116.
Signed-off-by: Akram
Signed-off-by: Akram <akram.benaissi@gmail.com>
* feat(deploy): default deployment configs to the official Alpine sandbox image
Update the shared gateway default_image, Helm chart values, the standalone
Kubernetes manifest, and the dev gateway task scripts to use
docker.io/library/alpine:3.22 instead of the community base image, consistent
with default_sandbox_image(). GPU e2e image-build base is left unchanged (CUDA
needs a glibc base).
Part of #3116.
Signed-off-by: Akram
Signed-off-by: Akram <akram.benaissi@gmail.com>
* feat(driver): default to numeric non-root identity for USER-less images
With the default sandbox image now Alpine, images that declare no OCI USER
must start instead of being rejected. When the image declares no USER and
the policy requests none, the Podman and Docker drivers now supply a numeric
non-root identity (DEFAULT_SANDBOX_UID/GID = 1000) instead of rejecting,
matching the numeric-identity behavior of the Kubernetes and VM drivers. The
supervisor's resolved-identity path runs the sandbox as a synthesized
non-root account without the account existing in the image. Images that
declare a USER keep the OCI resolution path unchanged.
Part of #3116.
Signed-off-by: Akram <akram.benaissi@gmail.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>
* test(conformance): use Alpine workload image
Signed-off-by: Evan Lezar <elezar@nvidia.com>
* refactor(policy): drop community image /app path from default policy
The restrictive default policy granted read-only access to /app, a directory
that only existed in the community base image. A generic Alpine default has no
/app, so remove it. Landlock best-effort already ignores absent paths; this
just stops advertising a community-specific layout in the default.
Part of #3116.
Signed-off-by: Akram
Signed-off-by: Akram <akram.benaissi@gmail.com>
* docs(config): document Alpine default images
Signed-off-by: Evan Lezar <elezar@nvidia.com>
* fix(podman): report early sandbox termination
Signed-off-by: Evan Lezar <elezar@nvidia.com>
* fix(podman): initialize rootless workspace ownership
Signed-off-by: Evan Lezar <elezar@nvidia.com>
* fix(sandbox): qualify NVIDIA Ubuntu default
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* fix(podman): initialize rootful default workspace
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* feat(sftp): add native sandbox adapter
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* fix(sftp): gate runtime helper support to Linux
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* fix(sftp): support standard OpenSSH file operations
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* fix(sftp): harden rename and special file handling
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* refactor(runtime): remove community image dependencies
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* test(e2e): build provider readiness tool fixture
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* fix(e2e): use a dedicated Noble fixture for Docker tests
Signed-off-by: Evan Lezar <elezar@nvidia.com>
---------
Signed-off-by: Akram
Signed-off-by: Akram <akram.benaissi@gmail.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Evan Lezar <elezar@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
* refactor(inference): remove managed inference routes
Closes#3172
Remove the inference route control plane, inference.local data path, built-in router crate, and SDK surface. Move inference workloads to explicitly imported provider profiles and native endpoints, with migration cleanup and updated tests and documentation.
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
* fix(policy): preserve alternate upstream isolation
Restore the provider policy activation guard so legacy OpenAI and Anthropic providers configured for alternate base URLs do not grant egress to the built-in public vendor endpoints.
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
---------
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
* feat(workspace): implement workspace model (Phase 1 of RFC 0011)
Implements workspace and membership model providing hard isolation
boundaries for multi-player OpenShell deployments.
Workspace CRUD with Kubernetes-style Terminating phase for graceful
deletion. All resources scoped by workspace via ObjectMeta. Membership
RPCs for workspace access control. Persistence migration shifts name
uniqueness to (object_type, workspace, name). Provider profiles support
platform and workspace scoping. Service routing uses workspace-prefixed
DNS labels. Inference routes renamed and workspace-scoped with
DeleteInferenceRoute RPC. Python SDK with WorkspaceClient, two-method
list pattern (workspace-scoped and for_all_workspaces), and workspace
parameter on all methods. CLI workspace flags, TUI workspace cycling.
K8s driver filters unmanaged CRs and uses delete preconditions. Podman
driver uses immutable container IDs. Label serialization fixed across
all put_if call sites.
Signed-off-by: Derek Carr <decarr@redhat.com>
* fix(cli): delegate sandbox upload command to existing upload function
The standalone `sandbox upload` command reimplemented upload logic
inline with two bugs: it used `Path::exists()` which follows symlinks
(rejecting dangling symlinks), and it ran git-aware filtering on
symlink sources. The `run::sandbox_upload()` function already handles
both cases correctly via `sandbox_upload_plan()`. Replace the inline
logic with a call to the existing function.
Signed-off-by: Derek Carr <decarr@redhat.com>
* fix(e2e): shorten sandbox names and fix test compatibility
Shorten the sandbox name in initial_sparse_policy_is_acknowledged_as_loaded
from 'e2e-2159-sparse-enrich' (22 chars) to 'e2e-sparse-enrich' (17 chars)
to comply with MAX_ROUTABLE_NAME_LEN (19 chars).
Also capture stderr in create_keep_with_args so future sandbox creation
failures include the actual CLI error instead of reporting empty output.
Signed-off-by: Derek Carr <decarr@redhat.com>
* test(workspace): add test coverage for workspace CRUD and persistence isolation
Add unit tests for workspace create happy path, get round-trip, get
not-found, get empty-name rejection, already-exists error, and
resolve_workspace not-found. Add persistence test proving cross-workspace
name uniqueness (same name in different workspaces produces separate
records). Add workspace name max-length boundary tests. Fix e2e harness
to include stderr in name-parse-failure error path. Align Python e2e
test_workspace_crud with try/finally pattern. Document provider profile
catalog workspace scoping gap in RFC 0011.
Signed-off-by: Derek Carr <decarr@redhat.com>
* fix(examples): update examples for workspace model compatibility
Shorten sandbox names in demo scripts to fit the 19-character
MAX_ROUTABLE_NAME_LEN limit: policy-demo prefix to pd-, multi-agent
notepad derives a short SANDBOX_TAG from the run ID, governance
interceptor uses gs-PID-RANDOM. Update vscode-remote-sandbox.md SSH
host aliases from openshell-{name} to openshell-{name}.{workspace}
format.
Signed-off-by: Derek Carr <decarr@redhat.com>
* feat(sdk): add workspace-scoped client and workspace CRUD
Add WorkspaceScopedClient modeled after kube::Api::namespaced — captures
workspace once and injects it into every sandbox request. Add workspace
CRUD methods (create, get, list, delete) and list_sandboxes_all_workspaces
on OpenShellClient. Extend SandboxRef with workspace field and add
WorkspaceRef type. Include mock tests for all new operations.
Signed-off-by: Derek Carr <decarr@redhat.com>
* fix(lint): resolve clippy warnings in workspace test assertions
Signed-off-by: Derek Carr <decarr@redhat.com>
* fix(docs): convert indented code blocks to fenced in RFC 0011
Signed-off-by: Derek Carr <decarr@redhat.com>
* fix(lint): resolve clippy warnings and apply cargo fmt across workspace
Auto-format with cargo fmt and fix clippy warnings exposed by the
reformat: unnecessary qualifications, map_unwrap_or, identical match
arms, unused variable prefix, dead code annotations, and let-unit-value
in e2e harness.
Signed-off-by: Derek Carr <decarr@redhat.com>
* fix(workspace): address workspace scoping issues from review
- Add workspace field to settings JSON output (CLI)
- Skip Podman containers missing workspace label instead of defaulting
to empty string, matching K8s driver behavior
- Add resource_version to list_by_scope SELECT in both SQLite and
Postgres backends, with regression test
- Gate PolicyLocalContext proposal/lookup routes on workspace readiness,
returning 503 when workspace is not yet discovered
- Block sandbox and provider creation in TUI all-workspaces mode
- Clear workspace vectors in TUI reset_sandbox_state
Signed-off-by: Derek Carr <decarr@redhat.com>
* fix(workspace): make provider profile catalog workspace-aware
Thread workspace through snapshot_catalog so the
EffectiveProviderProfileCatalog enforces workspace boundaries on both
read and write paths. UserProviderProfileSource now loads platform-scoped
profiles (workspace "") plus the target workspace's profiles, preventing
cross-workspace duplicate profile ID collisions that previously caused
global catalog failures.
Update RFC 0011 to reflect catalog scoping is implemented in Phase 1
rather than deferred to future work.
Signed-off-by: Derek Carr <decarr@redhat.com>
* fix(persistence): include workspace column in atomic policy revision INSERT
put_policy_revision_atomic omitted the workspace column from the INSERT
into the objects table in both SQLite and Postgres backends, causing
atomically-written policy revisions to lose their workspace association.
Add workspace field to AtomicPolicyRevisionWrite and thread it through
both backend INSERT statements, matching the non-atomic put_policy_revision
path which already included it.
Signed-off-by: Derek Carr <decarr@redhat.com>
* fix(proxy): skip ancestor walk when socket owner is the entrypoint
collect_ancestor_identities walked the entire process tree above the
entrypoint when the connecting process was the entrypoint itself,
SHA256-hashing every ancestor binary (IDE, shell, container runtime).
On dev machines with large binaries in the ancestor chain this exceeded
the 30-second test timeout. When start_pid == stop_pid there are no
intermediate ancestors to verify, so return an empty list immediately.
Signed-off-by: Derek Carr <decarr@redhat.com>
* fix(workspace): make provider profile catalog scope-aware
Allow the same profile ID at platform and workspace scopes by
introducing layered catalog entries where workspace profiles shadow
platform profiles. Add source and scope fields to the ProviderProfile
proto and CLI output. Migrate List/Get handlers to the catalog,
fixing divergence with runtime profile resolution.
Signed-off-by: Derek Carr <decarr@redhat.com>
* fix(e2e): align podman e2e labels with centralized driver constants
The podman driver moved its container labels to the centralized
openshell.ai/ prefix, but the e2e test harness and cleanup script
still referenced the old openshell.sandbox-* keys, causing the
local_driver_token_restart test to fail on container lookup.
Signed-off-by: Derek Carr <decarr@redhat.com>
* fix(e2e): align python profile isolation test with scope-aware catalog
Platform profiles are now visible in workspace listings as fallbacks
per the layered catalog design. Update the assertion to match.
Signed-off-by: Derek Carr <decarr@redhat.com>
* fix(workspace): honor profile_workspace in runtime profile resolution
Runtime profile lookups now consult provider.profile_workspace via
get_type_profile_for_scope. Providers created with --global-profile
(profile_workspace="") resolve to the platform profile even when a
workspace profile shadows the same ID. All 6 runtime call sites
updated; type-only call sites remain scope-agnostic.
Signed-off-by: Derek Carr <decarr@redhat.com>
---------
Signed-off-by: Derek Carr <decarr@redhat.com>
* fix(providers): allow git clone/fetch via default GitHub provider
The github.com:443 git-transport endpoint used the read-only access
preset, which expands to GET/HEAD/OPTIONS only. Git smart HTTP requires
a POST to */git-upload-pack for clone and fetch, so the L7 proxy denied
those operations and `gh repo clone` / `git clone https://...` failed.
Replace the preset with explicit rules that permit the read-only methods
plus POST */git-upload-pack, so clone/fetch work while push
(git-receive-pack) stays blocked. Enabling push still requires an
explicit policy proposal.
Why allowing this POST is still read-only: in git's smart HTTP protocol
POST is an RPC transport, not a write. A clone/fetch does GET
*/info/refs (ref discovery) followed by POST */git-upload-pack, whose
body is only the client's want/have negotiation; the server responds
with a packfile and nothing on the server is modified (data flows
server -> client). The service names are from the server's perspective:
git-upload-pack = the server uploads a pack to the client (a read/
download), while git-receive-pack = the server receives a pack from the
client (the actual write/push). The new rule is scoped to
*/git-upload-pack only, so push (git-receive-pack) and arbitrary POSTs
to github.com remain denied.
Add a provider-profile regression test and a rego enforcement test
covering ref discovery, upload-pack (allowed), and receive-pack (denied).
Closes#1769
Signed-off-by: Russell Bryant <rbryant@redhat.com>
* test(providers): strengthen git-transport regression and add clone e2e
Pin the exact allowed rule set for the built-in github git-transport
endpoint in both the provider-profile and composed-policy tests, so a
broader or additional POST rule (e.g. POST **) that could enable push
via git-receive-pack fails the test instead of passing a substring
check. Add an e2e test that attaches the built-in github provider and
clones a public repo over HTTPS, exercising provider attachment,
effective-policy composition, TLS interception, and real git behavior.
Update the Providers V2 docs so the github.com git-transport endpoint
shows explicit clone/fetch rules instead of the stale read-only preset.
Refs #1769
Signed-off-by: Russell Bryant <rbryant@redhat.com>
* test(providers): isolate providers_v2 mutation in clone e2e
The clone e2e enables the gateway-global providers_v2_enabled setting.
Restore its exact prior value (or absence) captured via GetGatewayConfig
instead of unconditionally deleting it, and serialize the mutation
across xdist workers with an exclusive file lock on the run's shared
base temp dir, so a shared or pre-configured gateway is left untouched
and parallel workers cannot race the read-modify-restore.
Refs #1769
Signed-off-by: Russell Bryant <rbryant@redhat.com>
* test(providers): serialize providers_v2 mutation with a suite-wide guard
The clone e2e's per-fixture lock only coordinated fixtures that acquired
it; other xdist workers hit the same gateway without it and could
observe the transiently-enabled providers_v2_enabled global during their
own sandbox creation (CWE-362).
Add an autouse readers-writer guard in conftest: every test holds a
shared lock on the gateway config, and a test marked
exclusive_gateway_config holds an exclusive lock. Mark the clone test
exclusive so no other worker is mid-test while it enables and restores
the gateway-global setting. Exact prior-value restoration is retained.
Refs #1769
Signed-off-by: Russell Bryant <rbryant@redhat.com>
---------
Signed-off-by: Russell Bryant <rbryant@redhat.com>
* fix(helm): build chart dependencies before lint
Signed-off-by: Evan Lezar <elezar@nvidia.com>
* test(e2e): remove python gpu smoke test
Remove the Python GPU smoke test and its fixture. The e2e:k3s:gpu task only depended on e2e:python:gpu and did not have a separate k3s implementation, so remove that stale alias with the task it pointed at.
Signed-off-by: Evan Lezar <elezar@nvidia.com>
(cherry picked from commit 221a10378e188656c710560740cbc9463c002db6)
---------
Signed-off-by: Evan Lezar <elezar@nvidia.com>
Pass CLUSTER_GPU=1 inline in e2e:python:gpu's depends so that the
cluster is bootstrapped with --gpu when GPU e2e tests are run.
Add --gpu flag handling to cluster-bootstrap.sh and default
OPENSHELL_E2E_GPU_IMAGE to an empty string so the server resolves
the default sandbox image when no override is provided.
Signed-off-by: Evan Lezar <elezar@nvidia.com>
* feat(sandbox): add gpu sandbox scheduling support
Allow sandbox creation to request GPU resources explicitly or infer them from GPU image names. This wires GPU intent through bootstrap, validates gateway support, and adds dedicated GPU E2E coverage for follow-up cluster testing.
Closes#101
Add pytest-xdist for parallel e2e test execution with configurable
concurrency. Default to 5 workers; override via E2E_PARALLEL env var
(accepts a number or 'auto' for CPU-count matching). Make session-scoped
mock inference route fixtures worker-safe by incorporating the xdist
worker_id into route names and routing hints.
Co-authored-by: John Myers <johntmyers@users.noreply.github.com>
> **🏗️ build-from-issue-agent**
Closes#79
## Summary
Replaces the gateway-proxied inference model (`ProxyInference` RPC) with sandbox-local execution. The sandbox now resolves routes from either a standalone YAML route file or a cluster bundle fetched via the new `GetSandboxInferenceBundle` RPC, then forwards requests directly to inference backends using `navigator-router`.
### Benefits of removing the gRPC hop
Moving inference execution from the gateway to the sandbox eliminates the gRPC round-trip for every inference request:
- **No payload size limits**: Direct HTTP forwarding replaces protobuf-over-gRPC serialization.
- **Streaming-ready**: Preserves connection semantics for future SSE support (e.g., `stream: true`).
- **Lower latency**: Direct sandbox → backend path instead of sandbox → gateway → backend round-trip.
- **Reduced gateway load**: Gateway only serves the control-plane bundle delivery RPC.
- **Simpler error model**: HTTP status codes flow through directly without gRPC wrapping.
## Changes Made
- **Proto**: Removed `ProxyInference` RPC. Added `GetSandboxInferenceBundle` RPC for route bundle delivery.
- **Gateway**: Replaced `proxy_inference` with `get_sandbox_inference_bundle`. Removed `Router` and `navigator-router` dependency — gateway is now control-plane only for inference.
- **Router**: Switched config file format from TOML to YAML. Added `Debug` impl that redacts API keys.
- **Sandbox**: New `InferenceContext` with local `Router` + route cache. Routes loaded from `--inference-routes` file (standalone) or cluster bundle via gRPC, with 30s background refresh. New `--inference-routes` / `NAVIGATOR_INFERENCE_ROUTES` CLI arg.
- **Dev sandbox**: Added `inference-routes.yaml` at repo root with default NVIDIA NIM route. `mise run sandbox` now mounts it automatically and supports `-e` flag to forward env vars (e.g., `mise run sandbox`).
- **Example**: Added standalone example (`examples/inference/routes.yaml`) and rewrote README to cover both standalone and cluster workflows.
- **E2E tests**: Updated existing test names/docs. Added tests for Anthropic messages protocol and multi-route policy filtering.
## Tests Added
- **Unit (21 new):** `router_error_to_http` (6 variants), `load_from_file` YAML round-trip (5), `resolve_sandbox_inference_bundle` gRPC handler (6), `build_inference_context` route loading (4)
- **E2E (2 new):** Anthropic messages protocol routing, route filtering by `allowed_routes`
- **Total:** 188 unit tests passing across 3 crates
## Documentation Updated
- `architecture/gateway.md`: Removed router, updated Inference Service
- `architecture/sandbox.md`: Updated orchestration flow, InferenceContext, route loading
- `architecture/inference-routing.md`: Complete rewrite for sandbox-local architecture
## Verification
- [x] All unit tests passing (188 across 3 crates)
- [x] Pre-commit Rust checks passing (fmt, clippy, tests)
- [x] Architecture documentation updated
- [x] E2E tests updated with new test cases
- [x] Standalone example and documentation added
- [x] Default inference-routes.yaml with env passthrough for mise run sandbox
Closes#67
## Summary
Implements transparent inference interception and routing for sandboxed AI agents. The sandbox proxy intercepts outbound AI SDK calls (OpenAI, Anthropic) and reroutes them through the gateway to policy-controlled backends — enabling organizations to redirect inference traffic to local or self-hosted models without modifying agent code.
**Decision model** — a tri-state OPA evaluation for every CONNECT request:
1. Binary + endpoint explicitly allowed in `network_policies` → **allow** (pass through)
2. Not explicitly allowed + `inference.allowed_routes` configured → **inspect for inference** (TLS intercept, detect API patterns, route through gateway)
3. Otherwise → **deny**
No endpoint declarations or binary lists needed for inference routing. Just configure `inference.allowed_routes`.
## Key Changes
### Sandbox (interception)
- **OPA policy**: New `network_action` Rego rule with three outcomes (`allow`, `inspect_for_inference`, `deny`). New `NetworkAction` enum replaces `PolicyDecision.allowed` bool for the proxy's main decision path.
- **Proxy**: New `InspectForInference` path — TLS-terminates client, parses HTTP, detects inference API patterns (`POST /v1/chat/completions`, `/v1/completions`, `/v1/messages`), strips auth headers, forwards via gRPC.
- **New module**: `l7/inference.rs` — `InferenceApiPattern`, `detect_inference_pattern()`, HTTP request/response parsing.
- **gRPC client**: New `proxy_inference()` for sandbox→gateway forwarding.
- **Sandbox init**: Creates OPA engine when inference is configured, even without `network_policies`.
### Gateway (dispatch)
- **InferenceService**: `ProxyInference` RPC loads sandbox policy, resolves allowed routes, dispatches to router. Full CRUD for inference routes.
- **Proto**: `InferenceRoute`, `InferenceRouteSpec`, `ProxyInferenceRequest/Response`, Inference gRPC service.
### Router (backend proxying)
- **New crate**: `navigator-router` with `Router`, `proxy_with_candidates()`, protocol-based route selection, backend HTTP proxying with auth header rewriting.
- **Mock support** for testing (`mock://` scheme).
### CLI
- `nav inference create/update/delete/list` commands for route management.
### Python SDK
- Updated protobuf bindings. Removed old `inference.py` client (replaced by transparent interception).
### Documentation
- New `architecture/inference-routing.md` — end-to-end system documentation.
- Updated `architecture/sandbox.md` — proxy, OPA, and source index sections.
- Updated `architecture/README.md` — new subsystem overview and diagram.
## Addendum: Chunked Transfer Compatibility
This branch now also fixes intercepted SDK requests that send chunked request bodies:
- `inspect_for_inference` now accepts `Transfer-Encoding: chunked` and decodes chunked request bodies before forwarding to the gateway
- Removed the prior `411 Length Required` response for chunked intercepted requests
- Added request/response header sanitization for framing and hop-by-hop headers (`content-length`, `transfer-encoding`, `connection`, etc.) to keep forwarded requests and returned responses valid
- Added unit tests for chunked parsing and header sanitization
Note: this improves compatibility for streaming-style SDK request patterns; true token-by-token passthrough response streaming is still a separate follow-up.
## Minimal Policy for Inference Routing
```yaml
inference:
allowed_routes:
- local
```
Any outgoing connection from a binary not explicitly allowed in `network_policies` will be intercepted and checked for inference API patterns.
## Test Plan
- [x] `cargo test --workspace` — all tests pass
- [x] `mise run pre-commit` — all checks pass
- [x] E2E: OpenAI chat completions routed through gateway
- [x] E2E: Anthropic messages routed through gateway
- [x] E2E: Python OpenAI SDK from sandbox (`examples/inference/inference.py`)
- [x] E2E test: `e2e/python/test_inference_routing.py`
Closes#13
## Summary
- add Python sandbox execution APIs for command and callable workflows
- consolidate sandbox policy fixtures and expand e2e test coverage for policy and Python exec paths
- update CI/build config and images for sandbox e2e execution dependencies
## Test Plan
- mise run pre-commit