mirror of
https://github.com/NVIDIA/OpenShell.git
synced 2026-10-04 08:28:19 +08:00
windows
73
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
1860010850 |
feat(sdk): add lazy pagination pagers (#3256)
* feat(sdk): add lazy pagination pagers Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(sdk): harden pager edge cases Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(sdk): cover initial resume token Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(go): fix all-workspaces pager examples Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
33bbda3d33 |
refactor(persistence): adopt continuation-token pagination (#3249)
* refactor(persistence): adopt continuation-token pagination Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(pagination): address continuation review findings Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(tui): recover completed list refreshes Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(pagination): address review scalability findings Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(pagination): repair branch validation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(go): use page size in template example Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
90dbe5454b |
feat(api): add typed workspace selectors (#3245)
* feat(api)!: add typed workspace selectors Signed-off-by: Mrunal Patel <mrunalp@gmail.com> * fix(cli): preserve template workspace metadata Signed-off-by: Mrunal Patel <mrunalp@gmail.com> * test(e2e): migrate workspace request selectors Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(api): update public schema inventory Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> --------- Signed-off-by: Mrunal Patel <mrunalp@gmail.com> Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> Co-authored-by: John Myers <9696606+johntmyers@users.noreply.github.com> |
||
|
|
f4dc6be4b2 |
refactor(inference): remove managed inference routes (#3195)
* refactor(inference): remove managed inference routes Closes #3172 Remove the inference route control plane, inference.local data path, built-in router crate, and SDK surface. Move inference workloads to explicitly imported provider profiles and native endpoints, with migration cleanup and updated tests and documentation. Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(policy): preserve alternate upstream isolation Restore the provider policy activation guard so legacy OpenAI and Anthropic providers configured for alternate base URLs do not grant egress to the built-in public vendor endpoints. Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> --------- Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> |
||
|
|
592df3e014 |
feat(policy): preserve exact MCP revision allowlists (#3027)
* feat(mcp): add version-aware wire profile metadata Signed-off-by: Shiju <shiju@nvidia.com> * feat(policy): canonicalize MCP version allowlists Signed-off-by: Shiju <shiju@nvidia.com> * fix(policy): align MCP policy tests with current main Signed-off-by: Shiju <shiju@nvidia.com> * fix(policy): canonicalize supervisor protobuf ingress Materialize defaultable MCP revisions before ambiguity checks, OPA construction, and sidecar delivery. Reject invalid sidecar policies with bounded errors. Signed-off-by: Shiju <shiju@nvidia.com> --------- Signed-off-by: Shiju <shiju@nvidia.com> |
||
|
|
07df822090 |
feat(providers): make profiles authoritative (#2962)
* feat(providers): make profiles authoritative Closes #1988 Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * docs(providers): move profiles into provider navigation Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * docs(providers): clarify provider attachment lifecycle Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * fix(tui): scroll provider profile picker Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * fix(providers): honor profile credential semantics Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * fix(providers): prefer exact profile IDs Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * fix(providers): harden authoritative profile adoption Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * test(oidc): align provider fixtures with profiles Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * fix(providers): preserve authoritative profile lifecycle Signed-off-by: John Myers <johntmyers@users.noreply.github.com> --------- Signed-off-by: John Myers <johntmyers@users.noreply.github.com> Co-authored-by: John Myers <johntmyers@users.noreply.github.com> |
||
|
|
5206bc51b2 |
feat(sdk): add OAuth Client Credentials support to SDKs (#2907)
* feat(sdk): add renewable client credentials auth Implement lazy OAuth client-credentials acquisition and renewal for the Python, TypeScript, and Go SDK clients, with shared security conformance coverage and service-account documentation.\n\nCloses #2803 Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(sdk): address client credentials review Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(sdk): address client credentials review Signed-off-by: Seth Jennings <sjenning@redhat.com> --------- Signed-off-by: Seth Jennings <sjenning@redhat.com> |
||
|
|
8d67250a5d |
fix(providers): keep refresh credential handles stable (#2780)
* fix(providers): keep refresh credential handles stable Signed-off-by: Mrunal Patel <mrunalp@gmail.com> * fix(providers): protect refresh-owned credentials Signed-off-by: Mrunal Patel <mrunalp@gmail.com> --------- Signed-off-by: Mrunal Patel <mrunalp@gmail.com> |
||
|
|
0120535efc |
feat(proxy): bind static credentials to provider endpoints (#2510)
* feat(proxy): bind static credentials to provider endpoints Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * test(e2e): verify static credential endpoint isolation Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * docs(provider): explain static credential endpoint binding Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * fix(e2e): use valid endpoint isolation fixtures Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * docs(provider): explain static credential endpoint binding Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * fix(credentials): preserve binding identity across rotations Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(proxy): enforce bindings across request lifecycle Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(proxy): close credential relay gaps Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * docs(credentials): clarify binding failure behavior Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(credentials): hash selected provider profile scope Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(proxy): resolve credentials after request admission Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * docs(credentials): clarify binding failure diagnostics Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(proxy): align single-route credential denials Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(credentials): harden endpoint-bound rotation Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(credentials): enforce identity and authority binding Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(credentials): snapshot provider environment atomically Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(e2e): include authority port in query proxy requests Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(credentials): close credential revocation gaps Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * docs(proxy): explain authority mismatch diagnostics Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(credentials): enforce binding lifecycle invariants Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(provider): reject credential config collisions Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(network): capture credential scope atomically Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(network): distinguish origin and absolute targets Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(provider): isolate endpointless profile credentials Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(network): normalize IPv6 request authorities Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * docs(credentials): clarify endpointless profile isolation Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * feat(policy): bind endpointless provider credentials Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(credentials): use current GCP placeholder revision Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * docs(providers): explain policy credential bindings Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(credentials): cover endpointless fail-closed invariant Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(policy): expect ambiguity rejection at creation Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(server): authenticate rebased policy requests Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * refactor(proxy): share credential mismatch finding builder Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(credentials): cover malformed binding metadata Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(credentials): verify multi-key endpoint isolation Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(e2e): cover same-host credential path denial Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * docs(credentials): document serialized refresh contract Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * refactor(proxy): consolidate L7 log formatting Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * perf(credentials): precompile endpoint binding patterns Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * perf(credentials): share identity epoch revisions Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(proxy): require explicit request default ports Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(policy): validate SigV4 credential sources Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(credentials): preserve endpoint bindings for credential handles Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * feat(go-sdk): expose network credential bindings Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> --------- Signed-off-by: John Myers <johntmyers@users.noreply.github.com> Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> Co-authored-by: John Myers <johntmyers@users.noreply.github.com> |
||
|
|
905b554c7c |
refactor(network): consolidate proxy egress pipeline (#2373)
* refactor(network): introduce shared egress pipeline Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(network): cover shared proxy egress paths Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * refactor(network): make destination authorization explicit Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * refactor(network): pin proxy relay policy context Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(network): lock relay generation contracts Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(network): establish phase zero compatibility baseline Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * feat(policy): detect ambiguous network endpoints Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * feat(sandbox): fail closed on invalid policy updates Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * refactor(network): invalidate relays on policy changes Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * docs(policy): document validation failure posture Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(network): cover validation and middleware egress Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(config): move policy failure mode to gateway toml Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(network): name proxy contracts by behavior Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(network): align overlap validation with endpoint selection Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(network): expect hard loopback denial Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(network): match declared endpoint denial Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(policy): preserve path-specific endpoint overrides Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(network): respect hard-blocked host gateways Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(network): reconcile proxy refactor with main Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(network): preserve CONNECT policy generation Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(policy): cover runtime endpoint glob semantics Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(server): reject ambiguous policies before persistence Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * docs(policy): explain ambiguity preflight behavior Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(policy): compare body limits within protocol Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(proxy): avoid global tracing capture race Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * chore(server): format rebased provider tests Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(server): authenticate rebased policy requests Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(sandbox): retain runtime on middleware outage Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(server): preflight provider composition activation Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(sandbox): distinguish runtime failure transitions Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> --------- Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> Signed-off-by: John Myers <johntmyers@users.noreply.github.com> Co-authored-by: John Myers <johntmyers@users.noreply.github.com> |
||
|
|
d220d89468 |
feat(compute): negotiate gateway callback listeners (#2492)
* feat(compute): query gateway listener requirements Signed-off-by: Evan Lezar <elezar@nvidia.com> * feat(compute): add Podman listener requirements Signed-off-by: Evan Lezar <elezar@nvidia.com> * test(docker): use default gateway bind address Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(gateway): avoid wildcard primary listener Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(podman): validate callback listener discovery Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(server): support split dual-stack listeners Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(podman): support legacy rootless listener discovery Signed-off-by: Evan Lezar <elezar@nvidia.com> * test(e2e): accept loopback plaintext rejection Signed-off-by: Evan Lezar <elezar@nvidia.com> * docs(agent): add callback listener diagnostics Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(server): restrict compute callback listeners Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(podman): validate local callback port Signed-off-by: Evan Lezar <elezar@nvidia.com> * test(server): clarify callback listener contract Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(podman): require pasta for local callbacks Signed-off-by: Evan Lezar <elezar@nvidia.com> * docs(gateway): document RPM listener default Signed-off-by: Evan Lezar <elezar@nvidia.com> * refactor(server): keep listener provenance diagnostic-only Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(compute): preserve callback listener isolation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(e2e): remove Podman callback relay Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(packaging): preserve Podman callback loopback Signed-off-by: Drew Newberry <anewberry@nvidia.com> * ci(e2e): run VM smoke on nested-virt runner Signed-off-by: Drew Newberry <anewberry@nvidia.com> * ci(e2e): gate VM smoke on usable KVM Signed-off-by: Drew Newberry <anewberry@nvidia.com> * ci(e2e): probe KVM through VM driver Signed-off-by: Drew Newberry <anewberry@nvidia.com> * ci(e2e): tolerate hosted KVM denial Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(server): close traced futures before assertions Signed-off-by: Drew Newberry <anewberry@nvidia.com> * revert: remove tracing test stabilization Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Evan Lezar <elezar@nvidia.com> Signed-off-by: Drew Newberry <anewberry@nvidia.com> Co-authored-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
9c019a93f5 |
Wire authorization into workspace model (#2445)
* feat(auth): implement RFC 0011 Phase 2 workspace authorization Signed-off-by: Derek Carr <decarr@redhat.com> * fix(auth): address PR review feedback on workspace authorization - Docker e2e: add --health-port and switch readiness probe from `openshell status` to `curl /healthz`, fixing a false-positive readiness check in OIDC mode where the CLI exited 0 without actually contacting the gateway - ListWorkspaces: move membership filtering from post-query N+1 lookups into a SQL EXISTS subquery so pagination applies to the visible set, not the global ordering. Add generic list_with_membership to the persistence layer. - Descriptor validator: reject role/scope fields on unauthenticated and sandbox auth modes, and allow-list workspace_role as user/admin and global_role as platform_admin to catch typos at startup Signed-off-by: Derek Carr <decarr@redhat.com> * fix(server): use authed request in delete telemetry test The workspace authorization added by the Phase 2 auth changes requires a Principal on every delete request. The delete-telemetry test was still using a bare Request::new, so extract_principal failed before the handler could acquire the delete gate, causing a 5-second timeout flake. Signed-off-by: Derek Carr <decarr@redhat.com> * fix(auth): address gator review findings for workspace authorization - Inject unauthenticated-local-dev principal in no-auth gateway mode so handlers that call extract_principal() always find one. - Cap label-selector membership query at MAX_PAGE_SIZE instead of u32::MAX to bound the in-memory read. - Authorize workspace membership before resolving workspace existence in all sandbox RPCs to prevent workspace-name enumeration by non-members. - Remove dead_code allow on AuthorizedWorkspace.workspace now that callers use the normalized name from the authz result. Signed-off-by: Derek Carr <decarr@redhat.com> * fix(auth): close workspace-name oracle and label-selector truncation Swap authorize-before-resolve ordering in 27 handlers across provider.rs, service.rs, policy.rs, and workspace.rs to prevent CWE-203 workspace-name enumeration by non-members. Add combined membership+label SQL query (list_with_membership_and_selector) to both persistence backends so ListWorkspaces with label selectors no longer silently drops results beyond the first page of membership matches. Signed-off-by: Derek Carr <decarr@redhat.com> * test(auth): add non-member rejection and membership+label persistence tests Add comprehensive test coverage for workspace authorization changes: - Non-member rejection tests across all 44 workspace-scoped handlers (sandbox, provider, service, policy, workspace, inference) verifying PERMISSION_DENIED is returned instead of NOT_FOUND to prevent CWE-203 workspace-name oracle - Persistence test for list_with_membership_and_selector verifying SQL-level membership EXISTS + label filtering, multiple predicates, no-match cases, and pagination Signed-off-by: Derek Carr <decarr@redhat.com> * fix(auth): format merged import line in sandbox tests Signed-off-by: Derek Carr <decarr@redhat.com> * fix(auth): address gator re-review findings on workspace authorization - Fix TUI unconditionally setting providers_v2_enabled after provider refresh; read the actual gateway setting via GetGatewayConfig at startup instead - Fix SQLite json_extract with dotted label keys (e.g. example.com/env) by quoting the key in the JSON path - Add authed_request wrappers to upstream OCI identity tests that were missing a principal after rebase - Add test proving GetGatewayConfig is accessible without Platform Admin - Add test for dotted/prefixed Kubernetes-style label key filtering Signed-off-by: Derek Carr <decarr@redhat.com> * fix(auth): address second gator re-review findings - Loosen GetGatewayConfig from platform_admin to scope-only so workspace users can discover providers_v2_enabled during sandbox creation with inferred-provider commands; update proto descriptor, descriptor validation, and RFC 0011 access table - Add validate_label_selector to handle_list_workspaces and escape single quotes in SQLite json_extract interpolation (CWE-89 defense-in-depth) - Re-fetch providers_v2_enabled after TUI gateway switch so the new gateway's capability is reflected - Add e2e test for workspace user with inferred-provider command - Add persistence test for adversarial label keys with SQL injection attempts - Add handler test for invalid label selector rejection in ListWorkspaces Signed-off-by: Derek Carr <decarr@redhat.com> * fix(auth): address third gator review findings - Cap label selector pairs at 64 (CWE-400) to bound SQLite dynamic SQL - Add SCOPE_ONLY_METHODS allowlist for scope-without-role RPCs (CWE-863) - Normalize ID-based data-plane handlers to return NOT_FOUND for unauthorized sandboxes, closing the cross-workspace oracle (CWE-203) - Fix TUI provider profile cache lookup key mismatch for legacy providers with empty profile_workspace - Add whoami to CLI skill reference command tree - Update TUI skill doc with workspace, provider, and settings coverage - Document scope/workspace orthogonality on GetGatewayConfig proto Signed-off-by: Derek Carr <decarr@redhat.com> * fix(auth): extend CWE-203 normalization to policy.rs sandbox handlers GetSandboxConfig and GetSandboxLogs in policy.rs had the same fetch-before-authorize pattern that leaked cross-workspace sandbox existence. Promote fetch_and_authorize_sandbox to pub(super) and use it from both sandbox.rs and policy.rs handlers. Signed-off-by: Derek Carr <decarr@redhat.com> * test(auth): update assertions for CWE-203 sandbox ID normalization Cross-workspace sandbox access via ID-based handlers now returns NOT_FOUND instead of PERMISSION_DENIED to prevent existence inference. Update the unit test and OIDC e2e assertion to match. Signed-off-by: Derek Carr <decarr@redhat.com> * fix(auth): narrow CWE-203 error mapping and correct whoami output formats Only remap PERMISSION_DENIED to NOT_FOUND in fetch_and_authorize_sandbox and RevokeSshSession, letting INTERNAL and UNAUTHENTICATED propagate as-is. Fix whoami --output format values in cli-reference.md to match the actual CLI (table/json/yaml, not text/json). Signed-off-by: Derek Carr <decarr@redhat.com> * fix(ci): share network namespace with Keycloak in containerized CI In GitHub Actions job containers, Docker port publishing lands on the host, not inside the job container. Detect this environment and attach Keycloak to the job container's network namespace instead, with hardened defaults (cap-drop ALL, no-new-privileges, loopback-only listener). Signed-off-by: Derek Carr <decarr@redhat.com> --------- Signed-off-by: Derek Carr <decarr@redhat.com> |
||
|
|
5952a5a23f |
feat(workspace): add workspace resource model with scoping, membershi… (#2243)
* feat(workspace): implement workspace model (Phase 1 of RFC 0011) Implements workspace and membership model providing hard isolation boundaries for multi-player OpenShell deployments. Workspace CRUD with Kubernetes-style Terminating phase for graceful deletion. All resources scoped by workspace via ObjectMeta. Membership RPCs for workspace access control. Persistence migration shifts name uniqueness to (object_type, workspace, name). Provider profiles support platform and workspace scoping. Service routing uses workspace-prefixed DNS labels. Inference routes renamed and workspace-scoped with DeleteInferenceRoute RPC. Python SDK with WorkspaceClient, two-method list pattern (workspace-scoped and for_all_workspaces), and workspace parameter on all methods. CLI workspace flags, TUI workspace cycling. K8s driver filters unmanaged CRs and uses delete preconditions. Podman driver uses immutable container IDs. Label serialization fixed across all put_if call sites. Signed-off-by: Derek Carr <decarr@redhat.com> * fix(cli): delegate sandbox upload command to existing upload function The standalone `sandbox upload` command reimplemented upload logic inline with two bugs: it used `Path::exists()` which follows symlinks (rejecting dangling symlinks), and it ran git-aware filtering on symlink sources. The `run::sandbox_upload()` function already handles both cases correctly via `sandbox_upload_plan()`. Replace the inline logic with a call to the existing function. Signed-off-by: Derek Carr <decarr@redhat.com> * fix(e2e): shorten sandbox names and fix test compatibility Shorten the sandbox name in initial_sparse_policy_is_acknowledged_as_loaded from 'e2e-2159-sparse-enrich' (22 chars) to 'e2e-sparse-enrich' (17 chars) to comply with MAX_ROUTABLE_NAME_LEN (19 chars). Also capture stderr in create_keep_with_args so future sandbox creation failures include the actual CLI error instead of reporting empty output. Signed-off-by: Derek Carr <decarr@redhat.com> * test(workspace): add test coverage for workspace CRUD and persistence isolation Add unit tests for workspace create happy path, get round-trip, get not-found, get empty-name rejection, already-exists error, and resolve_workspace not-found. Add persistence test proving cross-workspace name uniqueness (same name in different workspaces produces separate records). Add workspace name max-length boundary tests. Fix e2e harness to include stderr in name-parse-failure error path. Align Python e2e test_workspace_crud with try/finally pattern. Document provider profile catalog workspace scoping gap in RFC 0011. Signed-off-by: Derek Carr <decarr@redhat.com> * fix(examples): update examples for workspace model compatibility Shorten sandbox names in demo scripts to fit the 19-character MAX_ROUTABLE_NAME_LEN limit: policy-demo prefix to pd-, multi-agent notepad derives a short SANDBOX_TAG from the run ID, governance interceptor uses gs-PID-RANDOM. Update vscode-remote-sandbox.md SSH host aliases from openshell-{name} to openshell-{name}.{workspace} format. Signed-off-by: Derek Carr <decarr@redhat.com> * feat(sdk): add workspace-scoped client and workspace CRUD Add WorkspaceScopedClient modeled after kube::Api::namespaced — captures workspace once and injects it into every sandbox request. Add workspace CRUD methods (create, get, list, delete) and list_sandboxes_all_workspaces on OpenShellClient. Extend SandboxRef with workspace field and add WorkspaceRef type. Include mock tests for all new operations. Signed-off-by: Derek Carr <decarr@redhat.com> * fix(lint): resolve clippy warnings in workspace test assertions Signed-off-by: Derek Carr <decarr@redhat.com> * fix(docs): convert indented code blocks to fenced in RFC 0011 Signed-off-by: Derek Carr <decarr@redhat.com> * fix(lint): resolve clippy warnings and apply cargo fmt across workspace Auto-format with cargo fmt and fix clippy warnings exposed by the reformat: unnecessary qualifications, map_unwrap_or, identical match arms, unused variable prefix, dead code annotations, and let-unit-value in e2e harness. Signed-off-by: Derek Carr <decarr@redhat.com> * fix(workspace): address workspace scoping issues from review - Add workspace field to settings JSON output (CLI) - Skip Podman containers missing workspace label instead of defaulting to empty string, matching K8s driver behavior - Add resource_version to list_by_scope SELECT in both SQLite and Postgres backends, with regression test - Gate PolicyLocalContext proposal/lookup routes on workspace readiness, returning 503 when workspace is not yet discovered - Block sandbox and provider creation in TUI all-workspaces mode - Clear workspace vectors in TUI reset_sandbox_state Signed-off-by: Derek Carr <decarr@redhat.com> * fix(workspace): make provider profile catalog workspace-aware Thread workspace through snapshot_catalog so the EffectiveProviderProfileCatalog enforces workspace boundaries on both read and write paths. UserProviderProfileSource now loads platform-scoped profiles (workspace "") plus the target workspace's profiles, preventing cross-workspace duplicate profile ID collisions that previously caused global catalog failures. Update RFC 0011 to reflect catalog scoping is implemented in Phase 1 rather than deferred to future work. Signed-off-by: Derek Carr <decarr@redhat.com> * fix(persistence): include workspace column in atomic policy revision INSERT put_policy_revision_atomic omitted the workspace column from the INSERT into the objects table in both SQLite and Postgres backends, causing atomically-written policy revisions to lose their workspace association. Add workspace field to AtomicPolicyRevisionWrite and thread it through both backend INSERT statements, matching the non-atomic put_policy_revision path which already included it. Signed-off-by: Derek Carr <decarr@redhat.com> * fix(proxy): skip ancestor walk when socket owner is the entrypoint collect_ancestor_identities walked the entire process tree above the entrypoint when the connecting process was the entrypoint itself, SHA256-hashing every ancestor binary (IDE, shell, container runtime). On dev machines with large binaries in the ancestor chain this exceeded the 30-second test timeout. When start_pid == stop_pid there are no intermediate ancestors to verify, so return an empty list immediately. Signed-off-by: Derek Carr <decarr@redhat.com> * fix(workspace): make provider profile catalog scope-aware Allow the same profile ID at platform and workspace scopes by introducing layered catalog entries where workspace profiles shadow platform profiles. Add source and scope fields to the ProviderProfile proto and CLI output. Migrate List/Get handlers to the catalog, fixing divergence with runtime profile resolution. Signed-off-by: Derek Carr <decarr@redhat.com> * fix(e2e): align podman e2e labels with centralized driver constants The podman driver moved its container labels to the centralized openshell.ai/ prefix, but the e2e test harness and cleanup script still referenced the old openshell.sandbox-* keys, causing the local_driver_token_restart test to fail on container lookup. Signed-off-by: Derek Carr <decarr@redhat.com> * fix(e2e): align python profile isolation test with scope-aware catalog Platform profiles are now visible in workspace listings as fallbacks per the layered catalog design. Update the assertion to match. Signed-off-by: Derek Carr <decarr@redhat.com> * fix(workspace): honor profile_workspace in runtime profile resolution Runtime profile lookups now consult provider.profile_workspace via get_type_profile_for_scope. Providers created with --global-profile (profile_workspace="") resolve to the platform profile even when a workspace profile shadows the same ID. All 6 runtime call sites updated; type-only call sites remain scope-agnostic. Signed-off-by: Derek Carr <decarr@redhat.com> --------- Signed-off-by: Derek Carr <decarr@redhat.com> |
||
|
|
9377e0d5fe |
fix(providers): allow git clone/fetch via default GitHub provider (#2317)
* fix(providers): allow git clone/fetch via default GitHub provider The github.com:443 git-transport endpoint used the read-only access preset, which expands to GET/HEAD/OPTIONS only. Git smart HTTP requires a POST to */git-upload-pack for clone and fetch, so the L7 proxy denied those operations and `gh repo clone` / `git clone https://...` failed. Replace the preset with explicit rules that permit the read-only methods plus POST */git-upload-pack, so clone/fetch work while push (git-receive-pack) stays blocked. Enabling push still requires an explicit policy proposal. Why allowing this POST is still read-only: in git's smart HTTP protocol POST is an RPC transport, not a write. A clone/fetch does GET */info/refs (ref discovery) followed by POST */git-upload-pack, whose body is only the client's want/have negotiation; the server responds with a packfile and nothing on the server is modified (data flows server -> client). The service names are from the server's perspective: git-upload-pack = the server uploads a pack to the client (a read/ download), while git-receive-pack = the server receives a pack from the client (the actual write/push). The new rule is scoped to */git-upload-pack only, so push (git-receive-pack) and arbitrary POSTs to github.com remain denied. Add a provider-profile regression test and a rego enforcement test covering ref discovery, upload-pack (allowed), and receive-pack (denied). Closes #1769 Signed-off-by: Russell Bryant <rbryant@redhat.com> * test(providers): strengthen git-transport regression and add clone e2e Pin the exact allowed rule set for the built-in github git-transport endpoint in both the provider-profile and composed-policy tests, so a broader or additional POST rule (e.g. POST **) that could enable push via git-receive-pack fails the test instead of passing a substring check. Add an e2e test that attaches the built-in github provider and clones a public repo over HTTPS, exercising provider attachment, effective-policy composition, TLS interception, and real git behavior. Update the Providers V2 docs so the github.com git-transport endpoint shows explicit clone/fetch rules instead of the stale read-only preset. Refs #1769 Signed-off-by: Russell Bryant <rbryant@redhat.com> * test(providers): isolate providers_v2 mutation in clone e2e The clone e2e enables the gateway-global providers_v2_enabled setting. Restore its exact prior value (or absence) captured via GetGatewayConfig instead of unconditionally deleting it, and serialize the mutation across xdist workers with an exclusive file lock on the run's shared base temp dir, so a shared or pre-configured gateway is left untouched and parallel workers cannot race the read-modify-restore. Refs #1769 Signed-off-by: Russell Bryant <rbryant@redhat.com> * test(providers): serialize providers_v2 mutation with a suite-wide guard The clone e2e's per-fixture lock only coordinated fixtures that acquired it; other xdist workers hit the same gateway without it and could observe the transiently-enabled providers_v2_enabled global during their own sandbox creation (CWE-362). Add an autouse readers-writer guard in conftest: every test holds a shared lock on the gateway config, and a test marked exclusive_gateway_config holds an exclusive lock. Mark the clone test exclusive so no other worker is mid-test while it enables and restores the gateway-global setting. Exact prior-value restoration is retained. Refs #1769 Signed-off-by: Russell Bryant <rbryant@redhat.com> --------- Signed-off-by: Russell Bryant <rbryant@redhat.com> |
||
|
|
a2cd5f8eda |
fix(gateway): honor tty flag for interactive exec (#2315)
* fix(gateway): honor tty flag for interactive exec Pass the requested TTY mode through the interactive SSH relay. Skip PTY allocation and resize forwarding when TTY is disabled, and add regression coverage for both modes. Signed-off-by: emonq <emonq@outlook.com> * test(gateway): improve `test_sandbox_interactive_exec_honors_tty` to test streamed stdin and stdout/stderr Signed-off-by: emonq <emonq@outlook.com> --------- Signed-off-by: emonq <emonq@outlook.com> |
||
|
|
ff9af8e320 |
fix(sandbox): acknowledge initial policy revision; expose SDK labels/selectors (#2170)
* fix(sandbox): acknowledge initial policy revision The supervisor loaded and enforced a sandbox-scoped policy but never told the gateway which revision it loaded. The policy poll loop seeded itself with the initial revision's hash on its first poll, so `policy_changed` was never true for that revision and `ReportPolicyStatus(LOADED)` — which only ran in the hot-reload branch — was never called. The revision stayed `Pending` and `current_policy_version` stayed 0 even though the sandbox was `Ready` and the policy was effective. This was most visible with sparse policies that get baseline-enriched into a new revision during startup. After the OPA engine is constructed, report the exact sandbox revision the supervisor loaded as LOADED, and seed the poll loop from that revision so it is not re-reported. Report FAILED with the original construction error if engine construction or conversion fails. Only sandbox-sourced revisions (version > 0) whose canonical content matches the loaded policy are acknowledged; global and local-file policies are untouched. Delivery uses the shared bounded retry, is non-fatal on transient failure, and a pending initial acknowledgement is delivered before any newer revision so policy history is never reordered. Signed-off-by: Kyle Zheng <kyzheng@nvidia.com> * feat(python): expose sandbox labels and selectors The gateway protobuf and CLI already support request-level sandbox labels (`CreateSandboxRequest.name`/`labels`) and selector-based listing (`ListSandboxesRequest.label_selector`), but the public Python SDK dropped them, so Python-created sandboxes could not be found via `openshell sandbox list --selector ...`. Add optional, source-compatible `name`/`labels` to `SandboxClient.create`, `create_session`, and the high-level `Sandbox`, and `label_selector` to `list`/`list_ids`. `SandboxRef` now carries the gateway labels as an immutable mapping (default empty, so `SandboxRef(id, name, status)` still works). Caller-provided label mappings are copied. Attaching the high-level `Sandbox` to an existing sandbox rejects `name`/`labels` since creation metadata cannot change on attach. Template labels remain a separate concept. No protobuf changes are required. Signed-off-by: Kyle Zheng <kyzheng@nvidia.com> * fix(python): keep SandboxRef hashable and copy high-level labels Excluding the new immutable `labels` field from SandboxRef equality/hash (`compare=False`) preserves the original (id, name, status) identity and keeps the frozen dataclass hashable — a MappingProxyType field would otherwise make `hash(SandboxRef(...))` raise. Also defensively copy caller-provided labels in the high-level `Sandbox` so later caller mutation cannot change what is sent. Signed-off-by: Kyle Zheng <kyzheng@nvidia.com> * fix(sandbox): bound initial-policy-ack retries The poll loop retried a pending initial acknowledgement before processing any newer revision, but retried unconditionally forever. A permanently undeliverable ack (e.g. the revision was superseded before it could be reported) would then stall all later policy hot-reloads and provider-env refreshes. Cap the retries; after the bound, give up and resume normal polling so the loop cannot livelock on a stuck acknowledgement. Signed-off-by: Kyle Zheng <kyzheng@nvidia.com> * test(sandbox): add sparse-policy revision-2 acknowledgement e2e Regression for #2159: create a sandbox with the network-only policy-advisor fixture, which the supervisor enriches with baseline filesystem paths during startup (creating revision 2, superseding revision 1). Assert the effective policy reaches revision 2 and no revision remains Pending once the supervisor acknowledges the load. Adds SandboxGuard::create_keep_with_args to create a kept sandbox with an initial --policy. Signed-off-by: Kyle Zheng <kyzheng@nvidia.com> * fix(ci): correct sandbox checks Signed-off-by: Kyle Zheng <kyzheng@nvidia.com> * test(cli): serialize mTLS environment access Signed-off-by: Kyle Zheng <kyzheng@nvidia.com> * fix(sandbox): address policy review feedback Signed-off-by: Kyle Zheng <kyzheng@nvidia.com> * fix(sandbox): preserve exact policy acknowledgements Signed-off-by: Kyle Zheng <kyzheng@nvidia.com> * fix(sandbox): preserve local policy overrides Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> --------- Signed-off-by: Kyle Zheng <kyzheng@nvidia.com> Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> Co-authored-by: John Myers <9696606+johntmyers@users.noreply.github.com> |
||
|
|
f23c2c8e84 |
test(e2e): remove python gpu smoke test (#1948)
* fix(helm): build chart dependencies before lint Signed-off-by: Evan Lezar <elezar@nvidia.com> * test(e2e): remove python gpu smoke test Remove the Python GPU smoke test and its fixture. The e2e:k3s:gpu task only depended on e2e:python:gpu and did not have a separate k3s implementation, so remove that stale alias with the task it pointed at. Signed-off-by: Evan Lezar <elezar@nvidia.com> (cherry picked from commit 221a10378e188656c710560740cbc9463c002db6) --------- Signed-off-by: Evan Lezar <elezar@nvidia.com> |
||
|
|
ec197a43ef | fix(e2e): correct return type of _stub_with_token (#1897) | ||
|
|
530aaf1360 |
feat(drivers): support docker and podman config mounts (#1785)
* feat(drivers): support docker and podman config mounts Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(drivers): trim mount docs Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(e2e): cover local driver volume mounts Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): satisfy linux clippy lint Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(drivers): gate bind mounts behind gateway config * docs(sandbox): simplify mount examples * cleanup * test(e2e): stabilize branch checks * fix(drivers): tighten local mount validation Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
7a0c444445 |
refactor!(auth): drop SSH handshake secret (#1274)
* refactor!(auth): drop SSH handshake secret in favor of mTLS The OPENSHELL_SSH_HANDSHAKE_SECRET / x-sandbox-secret mechanism was misnamed: it does not authenticate SSH (which flows over the RelayStream gRPC RPC and is gated by mTLS plus supervisor Unix-socket permissions). It only gated a small set of sandbox-to-gateway control-plane RPCs, and production deployments already enforce mTLS on that channel — so the shared secret was redundant. Replace the secret check with an mTLS-presence marker. Sandbox-class methods (ReportPolicyStatus, PushSandboxLogs, GetSandboxProviderEnvironment, SubmitPolicyAnalysis, GetSandboxConfig, GetInferenceBundle) accept callers without a Bearer token; the gRPC mTLS handshake is the trust boundary. Dual-auth methods treat Bearer-present as full-scope CLI access and Bearer-absent as sandbox-restricted scope via validate_sandbox_caller_update. Drops the secret from all drivers (K8s, Podman, VM), the sandbox gRPC interceptor, the Helm chart (values + pre-install hook + StatefulSet env), the RPM bootstrap script, the man pages, and the debug-openshell-cluster skill. Also removes the never-read ssh_handshake_skew_secs flag and config field. BREAKING CHANGE: --ssh-handshake-secret / OPENSHELL_SSH_HANDSHAKE_SECRET and --ssh-handshake-skew-secs / OPENSHELL_SSH_HANDSHAKE_SKEW_SECS are removed from the gateway, sandbox, and all driver binaries. The openshell-ssh-handshake K8s Secret is no longer managed by the chart; operators may delete the orphan. Deployments using --disable-gateway-auth must enforce caller authentication at the fronting proxy, since the gateway no longer validates a per-request secret on sandbox-class methods. Refs OS-174. * docs(auth): scrub residual SSH handshake secret references Sweep across docs, e2e scripts, the Podman driver README/NETWORKING notes, the RPM/Helm/setup guides, the gateway man page, and RFC 0003 to remove instructions and examples that still referenced OPENSHELL_SSH_HANDSHAKE_SECRET / --ssh-handshake-secret / ssh_handshake_skew_secs. The mechanism is gone; nothing should still suggest setting it. Negative-assertion regression tests are kept so the env var cannot silently be re-introduced. * test(drivers): drop SSH handshake secret negative-assertion tests The supporting code, env vars, CLI flags, and config plumbing are gone — these tests asserted absence of strings that no longer have any path to being set. Remove the guards from the Docker, Podman, Kubernetes, and VM driver test modules. |
||
|
|
1d3b741ee3 |
feat(providers): support sandbox provider attach lifecycle (#1242)
* feat(providers): support sandbox provider attach lifecycle Closes #1171 Adds sandbox provider list, attach, and detach API/CLI support while keeping provider policy and credential resolution derived from current sandbox attachments. * fix(providers): refresh sandbox provider credentials Adds provider environment revisions and generation-scoped sandbox credential snapshots so future SSH and exec launches pick up provider attach, detach, and credential updates without mutating already-running processes. Also blocks provider deletion while attached to prevent stale sandbox provider references. * fix(providers): serialize sandbox object mutations * test(providers): cover sandbox provider attach lifecycle * test(providers): accept versioned credential placeholders |
||
|
|
e4b4e923ae | test(e2e): run suites against docker gateway (#1153) | ||
|
|
a255ad9142 | fix(e2e): stabilize wildcard host DNS test (#1144) | ||
|
|
084505425b |
feat(auth): add OIDC/Keycloak authentication with RBAC and scope-based permissions (#935)
* feat(auth): add OIDC/Keycloak authentication with RBAC Add OAuth2/OIDC authentication to the gateway server with role-based access control, CLI login flows, and full deployment plumbing. Server: JWT validation against configurable OIDC issuer (oidc.rs), JWKS key caching with TTL and rotation handling, method classification (unauthenticated/sandbox-secret/dual-auth/bearer), identity extraction with provider-agnostic Identity type, and RBAC enforcement via AuthzPolicy with configurable admin/user roles and auth-only mode. CLI: browser-based Authorization Code + PKCE flow, Client Credentials flow for CI/automation, token storage with refresh, gateway add/login/ logout commands, OIDC bearer token injection over mTLS transport, discovery endpoint for auto-configuration. Security: sandbox-secret scope restriction on UpdateConfig (policy sync only), anti-spoofing header stripping, dual-auth fallthrough from sandbox-secret to Bearer token. Deployment: OIDC config wired through DeployOptions, Docker env vars, Helm values/templates, HelmChart manifest, cluster-entrypoint.sh, and bootstrap scripts. Keycloak dev server script with pre-configured realm (test users, roles, PKCE client, CI client). Tested with Keycloak. The roles claim path and role names are configurable to support other OIDC providers. * feat(auth): add OAuth2 scope-based fine-grained permissions Add opt-in scope enforcement on top of existing OIDC role-based access control. When --oidc-scopes-claim is set, the server extracts scopes from the JWT and checks them per-method against an exhaustive scope map. Scopes: sandbox:read, sandbox:write, provider:read, provider:write, config:read, config:write, inference:read, inference:write, and openshell:all (wildcard). Methods not in the scope map require openshell:all. Scopes layer on top of roles and cannot escalate privilege. Auth-only mode (empty role names) still enforces scopes when enabled. Server: scopes_claim in OidcConfig, scope extraction from JWT (space-delimited and JSON array formats), standard OIDC scope filtering, scope check in AuthzPolicy after role check. CLI: --oidc-scopes on gateway add/start stored in metadata and consumed by gateway login, --oidc-scopes-claim on gateway start forwarded to server, scopes parameter in browser and client credentials OAuth2 flows with openid deduplication. Deployment: oidc_scopes_claim wired through DeployOptions, docker.rs, Helm, bootstrap scripts, and cluster entrypoint. Keycloak: realm config updated with built-in OIDC scopes and 9 OpenShell client scopes as optional on openshell-cli and openshell:all as default on openshell-ci. * fix(auth): address branch review findings Add GetInferenceBundle to sandbox-secret methods so sandbox inference route refresh works under OIDC. Make GetSandboxConfig dual-auth so CLI users can read sandbox settings with Bearer tokens. Preserve OIDC gateway metadata on restart — a bare gateway start without --oidc-* flags no longer erases the stored OIDC registration. Document CI client ID requirement (openshell-ci vs openshell-cli) in the testing guide. Add security note about auth-only mode blast radius for GitHub Actions. * fix(auth): complete review findings for OIDC auth boundary Move OpenShell/GetSandboxConfig from sandbox-secret-only to dual-auth so CLI users can read sandbox settings with Bearer tokens while sandbox supervisors continue using the shared secret. Add sandbox secret interceptor to the inference bundle fetch path so GetInferenceBundle works under OIDC-enabled gateways. Extract shared interceptor constructor to avoid duplication. Add GetSandboxConfig to the config:read scope map so scope enforcement applies consistently when scopes are enabled. Refactor OIDC metadata preservation into apply_oidc_gateway_metadata() with explicit resume semantics — only preserve existing OIDC metadata on real resume paths, not on fresh deployments. Update architecture docs and testing guide to reflect the corrected method classifications and add new test coverage for interceptor injection, scope requirements, metadata preservation, and dual-auth classification. * refactor(auth): use oauth2 crate for CLI OIDC flows Replace hand-written PKCE generation, authorization URL construction, token exchange, client credentials, and token refresh with the oauth2 crate's typed API. Eliminates sha2, hex, and getrandom dependencies from the CLI. The custom urlencoded() helper and manual form POST logic are replaced by BasicClient methods with proper type-state safety. Discovery and the callback server remain custom since the oauth2 crate does not provide OIDC discovery or a localhost redirect listener. * refactor(auth): move server auth modules into auth/ directory Group oidc.rs, authz.rs, identity.rs, and the auth HTTP endpoints under src/auth/ module directory. No behavioral changes. auth/mod.rs — module root, re-exports HTTP router auth/oidc.rs — JWT validation, JWKS caching, method classification auth/authz.rs — role and scope authorization policy auth/identity.rs — provider-agnostic Identity type auth/http.rs — /auth/connect and /auth/oidc-config endpoints * fix(auth): use RequestBody auth type for client credentials flow The oauth2 crate defaults to BasicAuth (HTTP Basic header) but Keycloak and most OIDC providers expect client_secret_post (credentials in the request body). Set AuthType::RequestBody explicitly to match the pre-refactor behavior. Also re-export Identity, IdentityProvider, and JwksCache from the auth module so ServerState's public API remains nameable by external consumers. * fix(auth): forward OPENSHELL_OIDC_SCOPES through cluster bootstrap Pass --oidc-scopes to gateway start so the metadata includes requested scopes after cluster bootstrap. Without this, users had to manually edit metadata.json to set scopes for gateway login. Usage: OPENSHELL_OIDC_SCOPES="openshell:all" mise run cluster * test(auth): add OIDC e2e tests for RBAC, scopes, and client credentials Add 10 end-to-end tests covering OIDC authentication against a live K3s cluster with Keycloak: RBAC (5 tests): admin can create providers, user cannot, user can list sandboxes, unauthenticated requests rejected, health probe works without auth. Scopes (4 tests): sandbox-scoped token can list sandboxes but not providers, openshell:all grants full access, no-scopes token denied. Client credentials (1 test): CI token via client_credentials grant. Tests are opt-in via OPENSHELL_E2E_OIDC=1 and OPENSHELL_E2E_OIDC_SCOPES=1 env vars. They derive the Keycloak URL from gateway metadata to match the server's configured issuer. Run with: OPENSHELL_E2E_OIDC=1 OPENSHELL_E2E_OIDC_SCOPES=1 \ PYTHONPATH=python uv run pytest e2e/python/oidc/ -v * fix(docs): fix markdown lint errors in OIDC architecture docs Add blank lines before lists and fenced code blocks to satisfy markdownlint MD031 and MD032 rules. |
||
|
|
d414e69a20 |
refactor(server): unify policy persistence in objects table (#972)
* refactor(server): unify policy persistence in objects table Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(server): clean sandbox-owned records on reconcile delete * refactor(server): use protos for stored policy records * refactor(server): move policy persistence into policy_store * fix(server): restore compute runtime merge compatibility * fix(server): validate draft chunk sandbox ownership --------- Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> |
||
|
|
5e28ea3a4b |
feat(server): add object meta convention to top-level objects (#919)
- adds filterable label selectors on resources Closes #864 Signed-off-by: Derek Carr <decarr@redhat.com> |
||
|
|
87f50f5e5d |
fix(e2e): add /dev/urandom to provider test sandbox policy (#948)
Python runtime requires /dev/urandom access during initialization to seed the hash randomizer. The _default_policy() in provider tests was missing this path, causing exec_python tests to fail with: 'Fatal Python error: _Py_HashRandomization_Init: failed to get random numbers to initialize Python'. Add /dev/urandom to read_only paths to match the policy used in test_sandbox_policy.py, allowing Python to initialize successfully. Signed-off-by: Derek Carr <decarr@redhat.com> |
||
|
|
29a3b1cacd |
fix(sandbox): two-phase Landlock to fix privilege ordering and add enforcement tests (#810)
* fix(sandbox): add parent-side Landlock availability probe logging Landlock status was only logged from inside the pre_exec child process where the tracing/OCSF pipeline is non-functional after fork. This made Landlock failures completely invisible in sandbox logs. Add a probe_availability() function that issues the raw landlock_create_ruleset syscall to check kernel support, and call it from the parent process before fork in all three spawn paths (entrypoint, SSH PTY, SSH pipe). Uses std::sync::Once to emit exactly once per sandbox lifetime. WIP - addresses logging gap from #803. * test(e2e): add Landlock filesystem enforcement tests Verify Landlock availability logging and enforcement in e2e: - OCSF probe event appears in sandbox logs - Read-only paths block writes, allow reads - Read-write paths allow both - Paths outside policy are denied entirely - User-owned paths outside policy are still blocked (proves Landlock enforces independently of Unix DAC permissions) Requires Linux host with Landlock support (GitHub Actions runners, Docker Desktop linuxkit). Related to #803. * test(e2e): add xfail tests proving #803 privilege ordering bug Two strict xfail tests that demonstrate the root cause of #803: - PathFd::new() runs as uid 998 after drop_privileges, so root-only paths (mode 700) silently fail and Landlock degrades - When ALL paths fail, best_effort silently drops Landlock entirely These tests will pass after the two-phase Landlock fix (open PathFds as root before drop_privileges, restrict_self after). * fix(test): fix Landlock e2e tests based on test run results - Remove log-reading tests: the Landlock probe logs to the supervisor's container stdout, not the in-sandbox file appender at /var/log/openshell*.log* - Replace broken xfail test (checked for child-process log messages that never reach the file appender) with a stat()-based test that verifies /root is actually in the Landlock allowlist - Keep enforcement tests (all passing) and the root-only policy xfail test (correctly proves #803 bug) * fix(test): remove mixed-policy xfail test Landlock doesn't restrict stat(), so the test passed unexpectedly. In the mixed policy case where /root is silently skipped, Landlock still applies (using other paths) and blocks /root even harder (not in allowlist = denied). The observable security degradation only occurs when ALL paths fail, which the existing xfail test already covers. * fix(sandbox): two-phase Landlock to fix privilege ordering (#803) Split Landlock apply into prepare() and enforce(): - prepare() runs as root before drop_privileges: opens PathFds, creates ruleset, adds rules. Root-only paths (mode 700) now succeed instead of silently failing as uid 998. - enforce() runs after drop_privileges: calls restrict_self() which does not require root. This fixes the root cause of #803 where drop_privileges() ran before sandbox::apply(), causing PathFd::new() to fail on root-only paths. In best_effort mode this silently dropped all Landlock restrictions. The fix applies to all three spawn paths: entrypoint (process.rs), SSH PTY (ssh.rs), and SSH pipe exec (ssh.rs). Removes xfail marker from e2e test that now passes. * fix(sandbox): address PR review feedback - enforce() now respects best_effort: if restrict_self() fails and policy is best_effort, log and degrade instead of aborting startup - log_sandbox_readiness distinguishes best_effort (degraded) from hard_requirement (will fail) in OCSF messages |
||
|
|
2ca553a4a0 |
fix(sandbox): validate always-blocked IPs at load time, enrich denial logs, and filter un-fixable proposals (#814) (#815)
Policies with allowed_ips entries targeting loopback, link-local, or unspecified ranges now fail at connection time instead of being silently blocked at runtime. The shorthand log format for DENIED events includes a [reason:...] suffix so operators can distinguish 'allowlist miss' from 'structurally un-allowable'. The mechanistic mapper skips proposals for always-blocked destinations, preventing the infinite TUI notification loop. The gateway validates proposed rules on approval as defense-in-depth. - Extract shared IP helpers (is_always_blocked_ip, is_always_blocked_net, is_internal_ip) to openshell_core::net - Reject always-blocked entries in parse_allowed_ips with hard error - Skip implicit allowed_ips synthesis for always-blocked literal IP hosts - Add status_detail to HttpActivityBuilder for denial reason propagation - Enrich NET and HTTP shorthand with [reason:...] for DENIED events - Add engine: tag to HTTP shorthand (consistency with NET shorthand) - Filter always-blocked proposals in mechanistic mapper generate_proposals - Add validate_rule_not_always_blocked server-side defense-in-depth - Update architecture docs, published docs, and E2E test assertions |
||
|
|
b7779bdefa |
feat(sandbox): integrate OCSF structured logging for sandbox events (#720)
* feat(sandbox): integrate OCSF structured logging for all sandbox events WIP: Replace ad-hoc tracing calls with OCSF event builders across all sandbox subsystems (network, SSH, process, filesystem, config, lifecycle). - Register ocsf_logging_enabled setting (defaults false) - Replace stdout/file fmt layers with OcsfShorthandLayer - Add conditional OcsfJsonlLayer for /var/log/openshell-ocsf.log - Update LogPushLayer to extract OCSF shorthand for gRPC push - Migrate ~106 log sites to OCSF builders (NetworkActivity, HttpActivity, SshActivity, ProcessActivity, DetectionFinding, ConfigStateChange, AppLifecycle) - Add openshell-ocsf to all Docker build contexts * fix(scripts): attach provider to all smoke test phases to avoid rate limits GitHub's unauthenticated API rate limit (60/hour) causes flaky 403s for Phases 1, 2, and 4. Fix by attaching the provider to all sandboxes and upgrading the Phase 1 policy to L7 so credential injection works. Phase 4 (tls:skip) cannot inject credentials by design, so relax the assertion to accept either 200 or 403 from upstream -- both prove the proxy forwarded the request. * fix(ocsf): remove timestamp from shorthand format to avoid double-timestamp The display layer (gateway logs, TUI, sandbox logs CLI) already prepends a timestamp. Having one in the shorthand output too produces redundant double-timestamps like: 15:49:11 sandbox INFO 15:49:11.649 I NET:OPEN ALLOWED ... Now the shorthand is just the severity + structured content: 15:49:11 sandbox INFO I NET:OPEN ALLOWED ... * refactor(ocsf): replace single-char severity with bracketed labels Replace cryptic single-character severity codes (I/L/M/H/C/F) with readable bracketed labels: [LOW], [MED], [HIGH], [CRIT], [FATAL]. Informational severity (the happy-path default) is omitted entirely to keep normal log output clean and avoid redundancy with the tracing-level INFO that the display layer already provides. Before: sandbox INFO I NET:OPEN ALLOWED ... After: sandbox INFO NET:OPEN ALLOWED ... Before: sandbox INFO M NET:OPEN DENIED ... After: sandbox INFO [MED] NET:OPEN DENIED ... * feat(sandbox): use OCSF level label for structured events in log push Set the level field to 'OCSF' instead of 'INFO' for OCSF events in the gRPC log push. This visually distinguishes structured OCSF events from plain tracing output in the TUI and CLI sandbox logs: sandbox OCSF NET:OPEN [INFO] ALLOWED python3(42) -> api.example.com:443 sandbox OCSF NET:OPEN [MED] DENIED python3(42) -> blocked.com:443 sandbox INFO Fetching sandbox policy via gRPC * fix(sandbox): convert new Landlock path-skip warning to OCSF PR #677 added a warn!() for inaccessible Landlock paths in best-effort mode. Convert to ConfigStateChangeBuilder with degraded state so it flows through the OCSF shorthand format consistently. * fix(sandbox): use rolling appender for OCSF JSONL file Match the main openshell.log rotation mechanics (daily, 3 files max) instead of a single unbounded append-only file. Prevents disk exhaustion when ocsf_logging_enabled is left on in long-running sandboxes. * fix(sandbox): address reviewer warnings for OCSF integration W1: Remove redundant 'OCSF' prefix from shorthand file layer — the class name (NET:OPEN, HTTP:GET) already identifies structured events and the LogPushLayer separately sets the level field. W2: Log a debug message when OCSF_CTX.set() is called a second time instead of silently discarding via let _. W3: Document the boundary between OCSF-migrated events and intentionally plain tracing calls (DEBUG/TRACE, transient, internal plumbing). W4: Migrate remaining iptables LOG rule failure warnings in netns.rs (IPv4 TCP/UDP, IPv6 TCP/UDP) to ConfigStateChangeBuilder for consistency with the IPv4 bypass rule failure already migrated. W5: Migrate malformed inference request warn to NetworkActivity with ActivityId::Refuse and SeverityId::Medium. W6: Use Medium severity for L7 deny decisions (both CONNECT tunnel and FORWARD proxy paths) to match the CONNECT deny severity pattern. Allows and audits remain Informational. * refactor(sandbox): rename ocsf_logging_enabled to ocsf_json_enabled The shorthand logs are already OCSF-structured events. The setting specifically controls the JSONL file export, so the name should reflect that: ocsf_json_enabled. * fix(ocsf): add timestamps to shorthand file layer output The OcsfShorthandLayer writes directly to the log file with no outer display layer to supply timestamps. Add a UTC timestamp prefix to every line so the file output matches what tracing::fmt used to provide. Before: CONFIG:VALIDATED [INFO] Validated 'sandbox' user exists in image After: 2026-04-01T15:49:11.649Z CONFIG:VALIDATED [INFO] Validated ... * fix(docker): touch openshell-ocsf source to invalidate cargo cache The supervisor-workspace stage touches sandbox and core sources to force recompilation over the rust-deps dummy stubs, but openshell-ocsf was missing. This caused the Docker cargo cache to use stale ocsf objects from the deps stage, preventing changes to the ocsf crate (like the timestamp fix) from appearing in the final binary. Also adds a shorthand layer test verifying timestamp output, and drafts the observability docs section. * fix(ocsf): add OCSF level prefix to file layer shorthand output Without a level prefix, OCSF events in the log file have no visual anchor at the position where standard tracing lines show INFO/WARN. This makes scanning the file harder since the eye has nothing consistent to lock onto after the timestamp. Before: 2026-04-01T04:04:13.065Z CONFIG:DISCOVERY [INFO] ... After: 2026-04-01T04:04:13.065Z OCSF CONFIG:DISCOVERY [INFO] ... * fix(ocsf): clean up shorthand formatting for listen and SSH events - Fix double space in NET:LISTEN, SSH:LISTEN, and other events where action is empty (e.g., 'NET:LISTEN [INFO] 10.200.0.1' -> 'NET:LISTEN [INFO] 10.200.0.1') - Add listen address to SSH:LISTEN event (was empty) - Downgrade SSH handshake intermediate steps (reading preface, verifying) from OCSF events to debug!() traces. Only the final verdict (accepted/denied) is an OCSF event now, reducing noise from 3 events to 1 per SSH connection. - Apply same spacing fix to HTTP shorthand for consistency. * docs(observability): update examples with OCSF prefix and formatting fixes Align doc examples with the deployed output: - Add OCSF level prefix to all shorthand examples in the log file - Show mixed OCSF + standard tracing in the file format section - Update listen events (no double space, SSH includes address) - Show one SSH:OPEN per connection instead of three - Update grep patterns to use 'OCSF NET:' etc. * docs(agents): add OCSF logging guidance to AGENTS.md Add a Sandbox Logging (OCSF) section to AGENTS.md so agents have in-context guidance for deciding whether new log emissions should use OCSF structured logging or plain tracing. Covers event class selection, severity guidelines, builder API usage, dual-emit pattern for security findings, and the no-secrets rule. Also adds openshell-ocsf to the Architecture Overview table. * fix: remove workflow files accidentally included during rebase These files were already merged to main in separate PRs. They got pulled into our branch during rebase conflict resolution for the deleted docs-preview-pr.yml file. * docs(observability): use sandbox connect instead of raw SSH Users access sandboxes via 'openshell sandbox connect', not direct SSH. * fix(docs): correct settings CLI syntax in OCSF JSON export page The settings CLI requires --key and --value named flags, not positional arguments. Also fix the per-sandbox form: the sandbox name is a positional argument, not a --sandbox flag. * fix(e2e): update log assertions for OCSF shorthand format The E2E tests asserted on the old tracing::fmt key=value format (action=allow, l7_decision=audit, FORWARD, L7_REQUEST, always-blocked). Update to match the new OCSF shorthand (ALLOWED/DENIED, HTTP:, NET:, engine:ssrf, policy:). * feat(sandbox): convert WebSocket upgrade log calls to OCSF PR #718 added two log calls for WebSocket upgrade handling: - 101 Switching Protocols info → NetworkActivity with Upgrade activity. This is a significant state change (L7 enforcement drops to raw relay). - Unsolicited 101 without client Upgrade header → DetectionFinding with High severity. A non-compliant upstream sending 101 without a client Upgrade request could be attempting to bypass L7 inspection. |
||
|
|
77e55ea989 |
test(e2e): replace flaky Python live policy update tests with Rust (#742)
Remove test_live_policy_update_and_logs and test_live_policy_update_from_empty_network_policies from the Python e2e suite. Both used a manual 90s poll loop against GetSandboxPolicyStatus that flaked in CI with 'Policy v2 was not loaded within 90s'. Add e2e/rust/tests/live_policy_update.rs with two replacement tests that exercise the same policy lifecycle (version bumping, hash idempotency, policy list history) through the CLI using the built-in --wait flag for reliable synchronization. |
||
|
|
1c659c1c12 |
fix(sandbox/bootstrap): GPU Landlock baseline paths and CDI spec missing diagnosis (#710)
* fix(sandbox): add GPU device nodes and nvidia-persistenced to landlock baseline Landlock READ_FILE/WRITE_FILE restricts open(2) on character device files even when DAC permissions would otherwise allow it. GPU sandboxes need /dev/nvidiactl, /dev/nvidia-uvm, /dev/nvidia-uvm-tools, /dev/nvidia-modeset, and per-GPU /dev/nvidiaX nodes in the policy to allow NVML initialization. Additionally, CDI bind-mounts /run/nvidia-persistenced/socket into the container. NVML tries to connect to this socket at init time; if the directory is not in the landlock policy, it receives EACCES (not ECONNREFUSED), which causes NVML to abort with NVML_ERROR_INSUFFICIENT_PERMISSIONS even though nvidia-persistenced is optional. Both classes of paths are auto-added to the baseline when /dev/nvidiactl is present. Per-GPU device nodes are enumerated at runtime to handle multi-GPU configurations. |
||
|
|
151fca9dc5 |
fix(server): return already_exists for duplicate sandbox names (#695)
Check for existing sandbox name before persisting, matching the provider-creation pattern. The CLI now surfaces a clear hint instead of a raw UNIQUE constraint error. Closes #691 |
||
|
|
e8950e624c |
feat(sandbox): add L7 query parameter matchers (#617)
* feat(sandbox): add L7 query parameter matchers Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(sandbox): decode + as space in query params and validate glob syntax Three improvements from PR #617 review: 1. Decode + as space in query string values per the application/x-www-form-urlencoded convention. This matches Python's urllib.parse, JavaScript's URLSearchParams, Go's url.ParseQuery, and most HTTP frameworks. Literal + should be sent as %2B. 2. Add glob pattern syntax validation (warnings) for query matchers. Checks for unclosed brackets and braces in glob/any patterns. These are warnings (not errors) because OPA's glob.match is forgiving, but they surface likely typos during policy loading. 3. Add missing test cases: empty query values, keys without values, unicode after percent-decoding, empty query strings, and literal + via %2B encoding. * fix(sandbox): add missing query_params field in forward proxy L7 request info * style(sandbox): fix formatting in proxy L7 query param parsing --------- Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> Co-authored-by: John Myers <johntmyers@users.noreply.github.com> |
||
|
|
256f7fc884 |
fix(sandbox,server): fix chunk merge duplicates and OPA variable collision with overlapping policies (#571)
* fix(sandbox,server): fix chunk merge duplicates and OPA variable collision with overlapping policies
Two related bugs triggered when a draft rule approval creates a second
policy entry for the same host:port:
1. merge_chunk_into_policy looked up existing rules by chunk.rule_name
(auto-generated as allow_{host}_{port}), which never matched the
user's original rule name. Now scans all network_policies entries
for a host:port endpoint match before falling back to insertion,
and merges allowed_ips into the existing endpoint.
2. The Rego allow_request rule and _matching_endpoint_configs
comprehension used 'some ep; ep := policy.endpoints[_]' which
caused regorus to error with 'duplicated definition of local
variable ep' when multiple policies covered the same host:port.
Refactored to isolate endpoint iteration inside helper functions
(_policy_allows_l7, _policy_endpoint_configs) so variables are
scoped per-policy evaluation.
Refs: #567
* test(e2e): add overlapping policy tests and update FWD-2 for implicit allowed_ips
- Update FWD-2 (test_forward_proxy_denied_without_allowed_ips ->
test_forward_proxy_allows_private_ip_host_without_allowed_ips):
literal IP host no longer requires explicit allowed_ips, expects 200.
- Add OVL-1: overlapping L4 policies for same host:port must not crash
OPA and should allow forward proxy connections.
- Add OVL-2: overlapping L7 policies for same host:port must not crash
OPA and should allow CONNECT tunnel establishment.
Refs: #567
* style: apply cargo fmt formatting
* test(e2e): update SSRF-3 and SSRF-6 for implicit allowed_ips behavior
SSRF-6: Private IP with literal IP host now gets implicit allowed_ips
from PR #570, so CONNECT returns 200 instead of 403.
SSRF-3: Loopback is still blocked but via the always-blocked path
(implicit allowed_ips is synthesized, then resolve_and_check_allowed_ips
catches it). Log message says 'always-blocked' instead of 'internal
address'.
* fix(e2e): use negative assertion for SSRF-6 when nothing listens on target port
When the SSRF check passes but nothing listens on the target port,
recv() returns empty bytes. Use 'assert 403 not in' (matching SSRF-4
pattern) instead of 'assert 200 in'.
* fix(e2e): update provider tests for redacted credential values
PR #569 changed credential redaction from clearing the map to
replacing values with 'REDACTED'. Update e2e assertions to expect
credential keys with REDACTED values instead of an empty map.
|
||
|
|
1a9eea5351 |
feat(tasks): wire e2e:gpu to bootstrap cluster with GPU support (#547)
Pass CLUSTER_GPU=1 inline in e2e:python:gpu's depends so that the cluster is bootstrapped with --gpu when GPU e2e tests are run. Add --gpu flag handling to cluster-bootstrap.sh and default OPENSHELL_E2E_GPU_IMAGE to an empty string so the server resolves the default sandbox image when no override is provided. Signed-off-by: Evan Lezar <elezar@nvidia.com> |
||
|
|
834f8aa184 |
fix: security hardening batch 1 (SEC-002 through SEC-010) (#548)
* fix(l7): reject ambiguous HTTP framing in REST proxy (SEC-009) Harden the L7 REST proxy HTTP parser against request smuggling: - Reject requests containing both Content-Length and Transfer-Encoding headers per RFC 7230 Section 3.3.3 (CL/TE ambiguity) - Replace String::from_utf8_lossy with strict UTF-8 validation to prevent interpretation gaps with upstream servers - Reject bare LF line endings (require CRLF per HTTP spec) - Validate HTTP version string (HTTP/1.0 or HTTP/1.1 only) * fix(server): harden shell_escape and command construction (SEC-002) Harden the gRPC exec handler against command injection via the structured-to-shell-string conversion: - Reject null bytes and newlines/carriage returns in shell_escape() - Add input validation in exec_sandbox: reject control characters in command args, env values, and workdir - Enforce size limits: max 1024 args, 32 KiB per arg/value, 4 KiB workdir, 256 KiB total assembled command string - Change shell_escape and build_remote_exec_command to return Result so callers must handle validation failures * fix(server): add command validation at SSH transport boundary (SEC-003) Add defense-in-depth validation in run_exec_with_russh before sending the command to the sandbox SSH server: - Reject null bytes in command string at transport boundary - Enforce max command length (256 KiB) at transport boundary - Enhance stream_exec_over_ssh logging with command length, stdin length, and truncated command preview for audit trail * fix(sandbox): validate port range and extend loopback check in SSH (SEC-007) Harden the SSH direct-tcpip channel handler: - Validate port_to_connect <= 65535 before u32-to-u16 cast to prevent port truncation (e.g., 65537 becoming port 1) - Replace string-literal loopback check with is_loopback_host() that covers the full 127.0.0.0/8 range, IPv4-mapped IPv6 (::ffff:127.x), bracketed IPv6, and case-insensitive localhost - Remove #[allow(clippy::cast_possible_truncation)] since the cast is now proven safe by the preceding range check * fix(sandbox): sanitize inference error messages returned to sandbox (SEC-008) Replace verbatim internal error strings in router_error_to_http with generic messages to prevent information leakage to sandboxed code. Upstream URLs, internal hostnames, TLS details, and file paths are no longer exposed. Full error context is still logged server-side at warn level by the caller for debugging. * fix(server): block internal IPs in SSH proxy target validation (SEC-006) Add IP validation in start_single_use_ssh_proxy to prevent SSRF if a sandbox status record were poisoned: - Resolve DNS before connecting and validate the resolved IP - Block loopback (127.0.0.0/8) and link-local (169.254.0.0/16, covers cloud metadata endpoint) addresses - Block IPv4-mapped IPv6 variants of the same ranges - Connect to the validated SocketAddr directly to prevent TOCTOU - Add debug logging of resolved target IP for audit * fix(sandbox): add resource limits to chunked body parser (SEC-010) Harden parse_chunked_body in the inference interception path: - Replace all unchecked +2 additions with checked_add for consistent overflow safety across all target architectures - Add MAX_CHUNKED_BODY (10 MiB) to cap decoded body size - Add MAX_CHUNK_COUNT (4096) to prevent CPU exhaustion via tiny chunks - Early-reject chunk sizes larger than remaining buffer space * fix(cli): double-escape command for SSH path, validate host and name (SEC-004) Harden doctor_exec against command injection in the SSH remote path: - Apply shell_escape to inner_cmd in the SSH path so it survives the double shell interpretation (SSH remote shell + sh -lc). This also fixes a correctness bug where multi-word commands were silently broken in the SSH path. - Add validate_gateway_name to reject shell metacharacters in gateway names before use in container_name - Add validate_ssh_host to reject metacharacters in remote_host loaded from metadata.json * fix(sandbox): add CIDR breadth warning and control-plane port blocklist (SEC-005) Defense-in-depth for the allowed_ips feature: - Log a warning when a CIDR entry has a prefix length < /16, as overly broad ranges may unintentionally expose control-plane services - Block K8s API (6443), etcd (2379/2380), and kubelet (10250/10255) ports unconditionally in resolve_and_check_allowed_ips, even when the resolved IP matches an allowed_ips entry * test(e2e): update assertion for sanitized inference error message (SEC-008) The SEC-008 fix changed the error message from 'no compatible route for source protocol ...' to 'no compatible inference route available'. Update the E2E assertion substring to match. |
||
|
|
bbcaed2ea7 | refactor(proto): rename UpdateSettings to UpdateConfig for consistency with read path (#515) | ||
|
|
a831a8921b |
feat(settings): gateway-to-sandbox runtime settings channel (#474)
* feat(gateway/sandbox): add global and sandbox runtime settings flow |
||
|
|
c0cdd665b6 | fix(gateway): allow first live network policy update (#493) | ||
|
|
de9dcaa44b | fix(e2e): update log-reading helpers for rolling file appender (#480) (#481) | ||
|
|
241e95dc39 |
feat(policy): support host wildcards and multi-port endpoints (#366)
* feat(policy): support host wildcards and multi-port endpoints Add glob-style host wildcards to endpoints[].host using OPA's glob.match with "." as delimiter — *.example.com matches a single DNS label, **.example.com matches across labels. Validation rejects bare * and requires *. prefix; warns on broad patterns like *.com. Add repeated uint32 ports field to NetworkEndpoint for multi-port support. Backwards compatible: existing port scalar is normalized to ports array. Both the proto-to-JSON and YAML-to-JSON conversion paths emit a ports array; Rego always references endpoint.ports[_]. Fix OpaEngine::reload() to route through the full preprocessing pipeline instead of bypassing L7 validation and port normalization. Closes #359 * fix(policy): reject configs with both port and ports set --------- Co-authored-by: johntmyers <johntmyers@users.noreply.github.com> |
||
|
|
647b7947f6 |
fix: security hardening from aardvark/codex scanner findings (#352)
* fix(sandbox): prevent overread request smuggling in L7 REST parser The HTTP/1.1 parser used a 1024-byte read buffer that could capture bytes from a pipelined second request. Those overflow bytes were forwarded upstream as body overflow without L7 policy evaluation, enabling request smuggling that bypasses per-request method/path enforcement. Replace multi-byte read with byte-at-a-time read_u8 that stops exactly at the CRLFCRLF header terminator. Add regression test proving two pipelined requests are parsed independently. Refs: #350 * fix(server): harden sandbox TLS secret volume permissions to 0400 Kubernetes secret volumes default to 0644, allowing the unprivileged sandbox user to read the mTLS client private key via the Landlock baseline /etc read set. A compromised sandbox could use the key to impersonate the control-plane client. Set defaultMode to 256 (octal 0400, owner-read only) on both the default and custom pod template paths. The supervisor reads TLS materials as root before forking, so this does not affect normal operation. Refs: #350 * fix(sandbox): stop using cmdline paths for binary policy matching /proc/pid/cmdline is fully attacker-controlled (argv[0] can be set to any string via execve) and had no integrity verification, unlike exec.path and ancestors which are kernel-managed and get TOFU/SHA256 checks. The Rego policy used cmdline_paths as a grant-access signal, allowing any sandboxed process to claim the identity of an allowed binary and bypass network restrictions. Remove the cmdline exact-match rule and exclude cmdline_paths from the glob-match rule. Only exec.path and exec.ancestors (from /proc/pid/exe) are now used for binary identity. cmdline_paths remain in the OPA input for deny-reason diagnostics only. Refs: #350 * fix(sandbox): reject symlink and non-dir read_write paths before chown The supervisor runs as root and calls chown on each read_write path. Since chown follows symlinks, a malicious container image could place a symlink (e.g. /sandbox -> /etc/shadow) to trick the supervisor into transferring ownership of arbitrary files to the sandbox user. Add symlink_metadata (lstat) check before chown to reject symlinks and non-directory entries. The TOCTOU window is not exploitable because no untrusted child process has been forked yet at this point. Refs: #350 * fix(server): redact provider credentials in gRPC CRUD responses Provider CRUD RPCs (create, get, list, update) returned full Provider objects including plaintext credentials (API keys, secrets). Any authenticated client -- including sandbox workloads running untrusted code -- could read credentials for all providers. Add redact_provider_credentials helper that clears the credentials map before returning. Internal server paths (inference routing, sandbox env injection) read from the store directly and are unaffected. Update tests to verify redaction and assert persistence via direct store reads. Refs: #350 * fix(sandbox): enforce non-root fallback when process user unset drop_privileges silently returned Ok(()) when both run_as_user and run_as_group were None, even when running as root. In local/dev mode policies are loaded from disk without passing through the server-side ensure_sandbox_process_identity normalization, so child processes could retain root and all capabilities (SYS_ADMIN, NET_ADMIN). When running as root with no process identity configured, fall back to sandbox:sandbox instead of no-oping. Non-root runtimes are unaffected. Refs: #350 * fix(sandbox): deny forward proxy for L7-configured endpoints The forward proxy path only performed L4 (endpoint) and allowed_ips checks. If an endpoint had L7 rules (method/path restrictions), a sandboxed process could bypass them by using HTTP_PROXY with plain http:// requests instead of CONNECT tunneling, since L7 inspection only runs in the CONNECT path. Add a guard in handle_forward_proxy that queries the endpoint's L7 config and returns 403 if any L7 rules are present, forcing traffic through the CONNECT path where per-request inspection happens. Refs: #350 * test(e2e): add regression test for forward proxy L7 bypass Verifies that the forward proxy path (plain http:// via HTTP_PROXY) returns 403 for endpoints with L7 rules configured, preventing sandboxed processes from bypassing per-request method/path enforcement by avoiding the CONNECT tunnel. Refs: #350 * fix: update tests for cmdline path and TLS volume changes Update OPA tests to verify cmdline_paths no longer grant access (regression tests for the security fix). Fix formatting in TLS volume test. Refs: #350 * test(e2e): update provider e2e tests for credential redaction Provider CRUD gRPC responses no longer include credential values. Update three e2e tests to assert credentials are empty in responses. Provider functionality is verified by existing e2e tests that check env var injection into sandboxes (which read from the store directly). Refs: #350 * chore: reformat for Rust 1.94 assert! macro style * fix(sandbox): make drop_privileges tests root-aware for CI CI runs as root but has no 'sandbox' user. The security fix correctly errors when running as root with no process identity and no fallback user available -- this is the intended behavior (refuse to run as root). Update tests to expect that error in root-without-sandbox-user environments instead of unconditionally asserting Ok. Refs: #350 * fix(sandbox): allow non-directory read_write entries like /dev/null The symlink guard incorrectly rejected all non-directory entries. Character devices like /dev/null are legitimate read_write paths used in sandbox policies. Only reject symlinks, which are the actual attack vector for the chown privilege escalation. Refs: #350 |
||
|
|
f6ae1da12d | chore: remove remaining navigator and nemoclaw references (#279) | ||
|
|
72e0268028 |
feat(sandbox): add gpu sandbox scheduling support (#257)
* feat(sandbox): add gpu sandbox scheduling support Allow sandbox creation to request GPU resources explicitly or infer them from GPU image names. This wires GPU intent through bootstrap, validates gateway support, and adds dedicated GPU E2E coverage for follow-up cluster testing. |
||
|
|
fbd93a4632 | refactor: rename navigator- crate prefix to openshell- (#277) | ||
|
|
7b0a243304 | ci: remove sandbox docker build from publish and e2e workflows (#275) | ||
|
|
b9d10861b6 | refactor(sandbox): move secrets to supervisor placeholders (#192) | ||
|
|
756950140c | refactor(python): rename navigator module to openshell and migrate config to gateway paths (#220) | ||
|
|
984d1a6e5c | chore: rename project from NemoClaw to OpenShell (#198) |