mirror of
https://github.com/NVIDIA/OpenShell.git
synced 2026-10-04 00:23:53 +08:00
windows
84
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
ab64e84bfc |
fix(core): enforce owner-only Windows ACLs on sensitive files and dirs (#3495)
* fix(core): enforce owner-only Windows ACLs on sensitive files and dirs set_dir_owner_only/set_file_owner_only were unconditional no-ops on Windows, so the CLI's mTLS client private key, OIDC/edge tokens, cached SSH keys, and the gateway's key-encryption key relied entirely on inherited NTFS ACLs with no OpenShell-applied restriction. Apply an owner-only DACL via SetEntriesInAclW/SetNamedSecurityInfoW with PROTECTED_DACL_SECURITY_INFORMATION to strip inherited ACEs, matching the 0700/0600 guarantee already provided on Unix. is_file_permissions_too_open now also works on Windows instead of being Unix-only, closing the detection gap alongside the prevention gap. Signed-off-by: Prashant Khodade <pkhodade@nvidia.com> (cherry picked from commit 71560e947f85819efbcddf70ddda94befab62b0b) * fix(core): treat a NULL DACL as too open in is_file_permissions_too_open has_foreign_trustee conflated a NULL DACL with an unreadable/invalid ACL and returned Some(false) (not too open) for both. Per the Win32 contract, a NULL DACL means the object grants full access to everyone -- the most permissive state possible -- so it must be flagged as too open. Split the null and invalid-ACL branches: null now returns Some(true), invalid ACL keeps the existing unreadable-ACL fallback (None, which the caller maps to false via unwrap_or). Adds a regression test that constructs a real NULL DACL via a SetNamedSecurityInfoW helper confined to the windows_acl module, consistent with the existing unsafe-FFI confinement in that module. Found by CodeRabbit review on MR !113. Signed-off-by: Prashant Khodade <pkhodade@nvidia.com> (cherry picked from commit 46e635a4ef1d6937cdb088f46aa85baee3d6ad28) * fix(core): close three false-negative gaps in the Windows ACL audit restrict_to_current_user() updated only the DACL, leaving a foreign owner's implicit WRITE_DAC right intact -- they could later replace the DACL we just set. Query OWNER_SECURITY_INFORMATION and take ownership in the same SetNamedSecurityInfoW call; if the caller can't (a genuinely foreign-owned object), the call now fails instead of silently leaving the object insecure. is_file_permissions_too_open() mapped every Win32 inspection failure (missing READ_CONTROL, an invalid ACL, a token-query failure) to "not too open" via unwrap_or(false). Fail closed instead: an inspection failure is a security false-negative risk, not a green light. has_foreign_trustee()'s ACE loop only recognized plain ACCESS_ALLOWED_ACE_TYPE and treated every other type as non-granting. Windows also defines access-allowed object, callback, and callback-object ACE variants that can grant rights to a foreign trustee; this audit doesn't parse their wider layouts, so their mere presence is now conservatively flagged as too open instead of silently skipped. Also updates architecture/gateway.md, which still described the SQLite file-tightening behavior only in terms of Unix mode 0o600, to distinguish it from the owner-only DACL behavior on Windows. Addresses review comments on PR #3495. Signed-off-by: Prashant Khodade <pkhodade@nvidia.com> * fix(core): conditional owner claim and audit owner in Windows ACL helpers restrict_to_current_user: query the current owner before calling SetNamedSecurityInfoW. Include OWNER_SECURITY_INFORMATION only when the path has a foreign owner -- requesting it unconditionally fails with ACCESS_DENIED (0x80070005) on standard credentials even when the current user is already the owner, because WRITE_OWNER is not implied by object ownership. A foreign-owned path still triggers an ownership claim and fails hard if the claim is denied, preserving the security contract. has_foreign_trustee: request OWNER_SECURITY_INFORMATION alongside DACL_SECURITY_INFORMATION and reject paths with a foreign owner immediately, before inspecting the DACL. A foreign owner has implicit WRITE_DAC rights and can replace any DACL we set, so a clean DACL is not sufficient evidence of safety on a foreign-owned object. architecture/gateway.md: clarify that the Windows path-hardening behavior sets mode 0o600 on Unix and applies a protected owner-only DACL on Windows, with conditional ownership claim and fail-hard semantics for foreign-owned objects. Signed-off-by: Prashant Khodade <pkhodade@nvidia.com> --------- Signed-off-by: Prashant Khodade <pkhodade@nvidia.com> |
||
|
|
39cf4823f7 |
feat(api): add structured gateway errors and SDK decoding (#3313)
* feat(api): expose structured gateway errors across SDKs Refs #3051. Add standard validation, conflict, and retry details; preserve raw transport status in Rust, Go, TypeScript, and Python; document status and recovery guidance. This is the structured-error foundation only. Mutation result shapes, allow_missing, durable request deduplication, and exec retry semantics remain follow-up work. Signed-off-by: Mrunal Patel <mrunalp@gmail.com> * fix(python): preserve wrapped RPC cleanup handling Inspect the original gRPC call when handling missing sandboxes during deletion waits and managed cleanup. Add intercepted cleanup regressions and clarify the error-wrapper migration contract. Addresses the cleanup review on #3313; part of #3051. Signed-off-by: Mrunal Patel <mrunalp@gmail.com> --------- Signed-off-by: Mrunal Patel <mrunalp@gmail.com> |
||
|
|
fd3fd9cf74 |
feat(sandbox): explain failed calls to external tool servers (#3207)
Show configured tool server addresses and their last observed connection results together in sandbox status. Keep sandbox lifecycle readiness separate so an external connection failure does not mark the sandbox unready. Expose direct endpoint records through the CLI and SDKs, with plain-language failure explanations and gateway acceptance times. Keep observation tracking, runtime reporting, and gateway validation in dedicated endpoint status modules. Preserve bounded reporting, request attribution, retry ordering, and configuration and supervisor authority checks. Clear obsolete observations while retaining the configured addresses, and document the distinction between an observed HTTP response, current availability, and tool success. Signed-off-by: Shiju <shiju@nvidia.com> |
||
|
|
02b664bb0d |
refactor(config): normalize and enforce gateway schema v2 (#2814)
* refactor(config): normalize compute driver field names Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * refactor(config): introduce canonical gateway fields Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * refactor(config): enforce gateway schema version 2 Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): preserve compute driver runtime guarantees Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): address schema v2 review regressions Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): complete schema v2 migration safeguards Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(config): expand schema v2 regression coverage Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(config): add schema v2 parity manifest Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): correct parity manifest inventory Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * docs(config): record schema v2 intentional changes Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * docs(config): disposition schema v2 parity gaps Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): add dual schema parity harness Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): establish compute lifecycle parity baseline Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): preserve gateway option compatibility Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): record gateway option parity Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * docs(config): close gateway-wide parity gaps Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(podman): apply configured pids limit Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): validate Podman option parity Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): add Kubernetes option parity harness Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): record Kubernetes option parity Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): disposition VM parity lanes Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): add external driver parity lane Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(e2e): preserve external driver pull policy Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): attest parity artifacts and launches Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): require clean parity build sources Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): bind parity runtime artifacts Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(e2e): use isolated supervisor tags Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(e2e): qualify parity image tags Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(e2e): serve parity supervisor locally Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): isolate parity podman services Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): harden parity evidence provenance Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): pin parity sandbox artifacts Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): attest parity runtime inputs Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): bind parity runtime evidence Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): record compute boundary parity Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): disposition cross-cutting parity lanes Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(packaging): preflight gateway config upgrades Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): preserve rebase integration guarantees Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(ci): isolate temporary git signing config Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): update remaining schema v2 consumers Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(ci): provide e2fs tools to VM tests Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): align preflight with gateway startup Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(vm): preserve rootfs tar configuration Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * chore(config): adopt duration unit constructors Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(packaging): preflight RPM gateway config Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): address driver review findings Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(e2e): require fresh semantic parity evidence Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(docker): update tests for renamed sandbox label Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(gateway): preserve selective driver coverage after rebase Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> Signed-off-by: Drew Newberry <anewberry@nvidia.com> Co-authored-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
33bbda3d33 |
refactor(persistence): adopt continuation-token pagination (#3249)
* refactor(persistence): adopt continuation-token pagination Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(pagination): address continuation review findings Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(tui): recover completed list refreshes Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(pagination): address review scalability findings Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(pagination): repair branch validation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(go): use page size in template example Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
90dbe5454b |
feat(api): add typed workspace selectors (#3245)
* feat(api)!: add typed workspace selectors Signed-off-by: Mrunal Patel <mrunalp@gmail.com> * fix(cli): preserve template workspace metadata Signed-off-by: Mrunal Patel <mrunalp@gmail.com> * test(e2e): migrate workspace request selectors Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(api): update public schema inventory Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> --------- Signed-off-by: Mrunal Patel <mrunalp@gmail.com> Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> Co-authored-by: John Myers <9696606+johntmyers@users.noreply.github.com> |
||
|
|
f4dc6be4b2 |
refactor(inference): remove managed inference routes (#3195)
* refactor(inference): remove managed inference routes Closes #3172 Remove the inference route control plane, inference.local data path, built-in router crate, and SDK surface. Move inference workloads to explicitly imported provider profiles and native endpoints, with migration cleanup and updated tests and documentation. Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(policy): preserve alternate upstream isolation Restore the provider policy activation guard so legacy OpenAI and Anthropic providers configured for alternate base URLs do not grant egress to the built-in public vendor endpoints. Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> --------- Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> |
||
|
|
48c449d8c8 |
chore(deps): replace ring with AWS-LC (#3243)
* chore(deps): replace ring with AWS-LC Signed-off-by: Simon Scatton <sscatton@nvidia.com> * fix(lint): address warnings after dependency upgrades Signed-off-by: Simon Scatton <sscatton@nvidia.com> * fix(tls): limit provider initialization to reqwest clients Signed-off-by: Simon Scatton <sscatton@nvidia.com> --------- Signed-off-by: Simon Scatton <sscatton@nvidia.com> |
||
|
|
6e6b3c8905 |
refactor(cli): remove local Dockerfile image builds (#3214)
Signed-off-by: Evie Howard <evhoward@redhat.com> |
||
|
|
b8162822d5 |
chore(example): refresh content guard lockfile (#3226)
Signed-off-by: Gordon Sim <gsim@redhat.com> |
||
|
|
e4369adcd0 |
chore(deps): replace serde_yml with noyalib (#3031)
* chore(deps): replace serde_yml with noyalib Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(providers): annotate generic YAML test values Signed-off-by: Evan Lezar <elezar@nvidia.com> --------- Signed-off-by: Evan Lezar <elezar@nvidia.com> |
||
|
|
039b265096 |
feat(middleware): define HTTP response pre-return interface (#3073)
* feat(middleware): define HTTP response pre-return interface Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(middleware): clarify HTTP response interface Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * refactor(middleware): align response result actions Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * refactor(middleware): expose response reason codes Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * refactor(middleware): share session end reasons Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * feat(middleware)!: finalize HTTP response pre-return contract Replace the separate body_end event with HttpResponseBodyUnit.end_of_stream. Every body-inspecting stage receives exactly one flagged unit, which may be empty; a zero-byte body is one empty flagged unit and OpenShell never reads ahead to set the flag. Defer response trailers from V1 and reserve their field numbers. HTTP/1.0 clients and Content-Length bodies cannot carry trailers and that behavior was undefined. Add HttpResponsePreflight.permitted_body_modes, computed once from the original upstream head so every stage sees the same list, and make an unlisted selection a failure rather than a downgrade. Add the block_delivery preflight action as a successful decision enforced regardless of on_error. Expose Content-Length, Content-Encoding, and Content-Range read-only in preflight. Cap STREAM_BYTES input units at half of max_payload_bytes and permit deferring bytes across replacements only for fail_closed bindings, surfaced as deferral_permitted. Split PEER_DISCONNECT into DOWNSTREAM_DISCONNECT and UPSTREAM_DISCONNECT and attribute WebSocket relay failures by direction instead of a generic peer error. Compile the content-guard example in lint and branch checks so proto renames cannot break it silently. BREAKING CHANGE: WebSocketSessionEndReason and WebSocketSessionEnd are replaced by the shared MiddlewareSessionEndReason and MiddlewareSessionEnd. NORMAL_CLOSE is now NORMAL, UPSTREAM_REJECTED is now UPSTREAM_FAILURE, and PEER_DISCONNECT is split into DOWNSTREAM_DISCONNECT and UPSTREAM_DISCONNECT. Enum numbers are unchanged so binary wire compatibility is preserved; generated symbols and JSON names change. Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(middleware): describe skip as opting out of inspection Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * feat(middleware): add body-phase block_delivery and skip_remaining actions Body results may now stop delivery or opt out of inspecting the rest of the response after a prefix. One HttpResponseBlockDelivery message is shared by preflight and body results and documents the difference between blocking before and after head commitment. Drop the field reservations, since nothing in this contract has shipped, and renumber session_end to close the gap. Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * refactor(middleware): share HTTP body leaf messages across directions HttpBodyUnit, HttpBodyPassThrough, HttpBodyTransform, HttpBodySkipRemaining, and HttpBodyMode carry no response-specific semantics, so name them for reuse by the streaming request hook. Envelopes, results, preflight, and block_delivery stay response-specific because commitment semantics differ. Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * refactor(middleware): keep HTTP body leaf messages response-specific Reverts the shared HttpBody* naming. A direction-specific payload such as a response-only semantic mode would otherwise add unreachable variants to the other direction or force a source-breaking fork after 0.1.0. The streaming request hook defines its own HttpRequestBody* messages and copies the shape; SDKs present a direction-neutral body handler over both. Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(middleware): simplify response proto comments Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(middleware): reject undispatched response bindings Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(middleware): simplify phase field comment Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * refactor(middleware): rename HTTP response preflight result Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * feat(middleware): add HTTP response trailer results Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(middleware): define response block delivery Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(middleware): define stage-local response body modes Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(middleware): define final response body units Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(middleware): define streaming response deferral Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(middleware): defer whole-body accumulation timeout Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(middleware): align response result diagnostics Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * test(middleware): cover upstream WebSocket disconnect Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(middleware): keep response streams unit-local Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(middleware): trim disconnect compatibility note Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> --------- Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> |
||
|
|
c93b2fa7da |
docs(gateway-config): fix stale community sandbox image path (#2800)
* docs(gateway-config): fix stale community sandbox image path Signed-off-by: Yuedong Wu <dwcn22@outlook.com> * docs(sandbox-image): purge remaining stale image references Rebasing onto main surfaced four more instances of the same dead ghcr.io/nvidia/openshell/sandbox path, introduced by commits merged after this branch was opened: three test fixtures (driver-docker, openshell-ocsf, compute::mod) and one user-facing default in the SPIFFE token-exchange Podman demo README. Correct all four to ghcr.io/nvidia/openshell-community/sandboxes/base, consistent with the rest of this fix. Signed-off-by: Yuedong Wu <dwcn22@outlook.com> --------- Signed-off-by: Yuedong Wu <dwcn22@outlook.com> |
||
|
|
857af42a16 |
feat(vm): support corporate HTTP forward proxy egress for microVM sandboxes (#3090)
* feat(vm): support corporate HTTP forward proxy egress for microVM sandboxes The corporate forward proxy machinery from #1792 is driver-agnostic and already merged: openshell-supervisor-network implements CONNECT chaining, NO_PROXY matching, credentials, https:// proxies and corporate CA trust, and openshell-sandbox exposes it as six argv-only flags. Podman gained the driver half in #2245/#2512 and Kubernetes in #2633; the VM driver had none of it, so VM sandboxes on proxy-only networks could not reach any destination requiring the proxy even when policy allowed it. The blocking piece was not proxy logic but delivery: the VM guest init script runs as PID 1 and execs a fixed supervisor command line, and libkrun's krun_set_exec receives an empty argv, so there was no channel for driver-owned supervisor arguments. The supervisor's proxy flags deliberately have no environment fallback, and build_guest_environment merges user-supplied environment, so the guest env is not a safe transport either. Add a driver-authored argument file, mirroring the existing init.d manifest: the driver writes /opt/openshell/supervisor-args into the overlay upperdir on every launch and the guest reads it verbatim, one argument per line, appending it to every supervisor exec. It is written even when empty, which is what makes the channel unforgeable -- the upperdir always shadows the read-only image layer, so an image can neither supply its own arguments nor disable the operator's by omitting the file. Because both launch backends exec the same init script, this covers libkrun and QEMU without touching either. A microVM has no bind mounts or container secrets, so the credential and CA bundle are staged into the per-sandbox overlay the way the gateway JWT already is: credential root-only at 0600, CA at 0644, both rewritten every launch so a removed setting clears prior material, and both deleted with the sandbox state directory. This places the credential at rest in the overlay image on the gateway host, which differs from the Podman secret model and is documented as an explicit security consideration. Validation is fail-closed and shared: a new openshell_core::driver_utils::validate_upstream_proxy_settings holds the pairing rules the Podman driver established, and both the gateway and the driver call it so an invalid table names the offending key instead of surfacing as an opaque driver-readiness timeout. Guest egress leaves through gvproxy, so a proxy on the gateway host's loopback is reachable only through host.openshell.internal; the guest to gateway callback is unaffected. Closes #3088 Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(vm): bound the proxy CA read and scope the host-loopback recipe Two review findings on the corporate forward proxy support for microVM sandboxes. The driver read the operator's proxy_ca_bundle with an unbounded fs::read and accepted it on a substring match for the PEM BEGIN CERTIFICATE marker. A special file such as /dev/zero therefore grew driver memory without bound on every authorized sandbox create, and a PEM block holding invalid DER passed the host check but contributes no trust anchor in the guest, so every supervisor would fail after boot with an error attributed to the sandbox rather than to the setting. Move the read into openshell-core as read_upstream_proxy_ca_bundle_file: it reuses the credential reader's bounded-read path (non-regular files rejected on fstat, size capped, read bounded even if the file grows), then requires at least one anchor that RootCertStore::add_parsable_certificates accepts. The supervisor's own reader now delegates to it, so host acceptance and guest acceptance are the same function and cannot drift. The published host-loopback recipe was written for libkrun only. gvproxy NATs host.openshell.internal to the gateway host's 127.0.0.1, but GPU sandboxes run on the QEMU/TAP backend where that name resolves to the TAP host address and the driver's own nftables input chain accepts only the gateway port from the guest — no proxy on the gateway host is reachable there at any bind address, so an operator following the generic recipe lost all proxy-required egress while configuration validation succeeded. Scope the recipe to libkrun in every reference and reject a gateway-host proxy URL when a launch plan resolves to QEMU, naming the reason, instead of booting a sandbox whose policy-approved CONNECTs all time out. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(vm): match the QEMU proxy preflight to the selected TAP host The gateway-host proxy guard added for the QEMU/TAP backend classified the wrong set of addresses in both directions. It ran at the top of configure_qemu_launch_plan, before the subnet allocation that settles plan.host_ip, so it could not compare against the address the guest actually reaches the host on. An operator pointing https_proxy at the sandbox's own TAP host address, such as 10.0.128.1, passed the check, and the driver's nftables input chain — which accepts only the gateway port from the guest — then dropped every policy-approved CONNECT, which is exactly the silent timeout the guard exists to prevent. In the other direction it rejected 192.168.127.254 unconditionally. That address is special only to libkrun/gvproxy; on QEMU/TAP it is an ordinary address that may be routable through the guest's masqueraded egress, so the guard refused a working configuration. Run the check after the launch plan's network allocation, on both the freshly-allocated and already-complete paths, and compare IP literals with that sandbox's selected TAP host. Loopback literals, localhost, and the documented host aliases that write_host_gateway_aliases seeds to the TAP host still classify as the gateway host, and the failure names the address. The gvproxy host-loopback constant returns to being a documentation anchor. Signed-off-by: Philippe Martin <phmartin@redhat.com> --------- Signed-off-by: Philippe Martin <phmartin@redhat.com> |
||
|
|
9ca19e6c80 |
refactor(compute): decouple gateway driver composition (#2823)
* refactor(compute): decouple gateway driver composition Move first-party composition and VM process ownership into openshell-gateway, leaving openshell-server backend-independent. Update packaging and build references with the new crate, simplify the compiled-driver boundary, and keep the driver-free gateway path buildable with bundled Z3 tooling. Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(telemetry): bound compute driver categories Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(core): keep runtime transport generic Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(compute): complete server driver decoupling Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(compute): preserve driver integrations after rebase Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> * fix(compute): preserve docker tracing after decoupling Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> * fix(compute): preserve driver behavior after extraction Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(compute): remove MXC policy side channel Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(compute): separate policy delivery from readiness Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> |
||
|
|
07df822090 |
feat(providers): make profiles authoritative (#2962)
* feat(providers): make profiles authoritative Closes #1988 Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * docs(providers): move profiles into provider navigation Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * docs(providers): clarify provider attachment lifecycle Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * fix(tui): scroll provider profile picker Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * fix(providers): honor profile credential semantics Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * fix(providers): prefer exact profile IDs Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * fix(providers): harden authoritative profile adoption Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * test(oidc): align provider fixtures with profiles Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * fix(providers): preserve authoritative profile lifecycle Signed-off-by: John Myers <johntmyers@users.noreply.github.com> --------- Signed-off-by: John Myers <johntmyers@users.noreply.github.com> Co-authored-by: John Myers <johntmyers@users.noreply.github.com> |
||
|
|
40d1b48666 |
feat(provider): support for SPIFFE backed token exchange (#1970)
* feat(provider): add ability to request token exchange instead of client credentials as OAuth grant_type Signed-off-by: Gordon Sim <gsim@redhat.com> * test(proxy): add further tests for token exchange Signed-off-by: Gordon Sim <gsim@redhat.com> * test(provider): add runnable example for token exchange Signed-off-by: Gordon Sim <gsim@redhat.com> * test(e2e): cover Podman token exchange grants Signed-off-by: Gordon Sim <gsim@redhat.com> * refactor(oauth): extract duplicated functionality from server and supervisor Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(provider): evict nearest-to-expiry entry from intermediate token cache Signed-off-by: Gordon Sim <gsim@redhat.com> * doc(supervisor): add podman example for token exchange Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(provider): withhold token-exchange subject credentials Signed-off-by: Gordon Sim <gsim@redhat.com> --------- Signed-off-by: Gordon Sim <gsim@redhat.com> |
||
|
|
40f822906c |
feat(compute): add standalone first-party drivers (#2822)
* feat(compute): add standalone first-party drivers Build Docker, Podman, Kubernetes, and VM drivers as external binaries and exercise each through the public compute-driver API. Keep the external E2E setup complete at introduction, including VM image selection, Kubernetes post-renderer isolation, supervisor reuse, and scoped Podman coverage. External Kubernetes endpoints support shared and managed workspace modes. Operator mode remains restricted to the in-process driver because gateway authentication and the driver must share a dynamic namespace allowlist. Signed-off-by: Drew Newberry <anewberry@nvidia.com> * ci(e2e): run managed and external drivers independently Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(compute): cover external driver socket contract Signed-off-by: Evan Lezar <elezar@nvidia.com> * ci(e2e): install bundled Z3 build dependency Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(server): gate in-tree driver tracing Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> Signed-off-by: Evan Lezar <elezar@nvidia.com> Co-authored-by: Evan Lezar <elezar@nvidia.com> |
||
|
|
b2ea81822b |
feat(network): enable Docker and Podman policy DNS and transparent TCP (#2723)
* feat(network): enable Docker transparent TCP egress Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(e2e): cover Docker transparent TCP egress Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * feat(network): correlate transparent TCP audit events Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * docs(examples): add transparent TCP Redis demo Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * docs(examples): demonstrate blocked TCP connections Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * docs(examples): focus Redis demo audit output Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(network): close transparent TCP policy bypasses Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(sandbox): reject unsupported TCP policy reloads Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(ci): satisfy Linux transparent TCP lints Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * feat(podman): enable transparent TCP egress Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(podman): permit policy DNS port binding Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(e2e): use qualified transparent TCP hostname Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(podman): preserve exact policy DNS names Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(podman): route policy DNS over TCP Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(dns): serve multiple TCP queries per connection Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * docs(network): explain native DNS and TCP egress Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(sandbox): reconcile runtime reload with upstream Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(network): harden transparent DNS capture Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(podman): preserve resolver behavior for native tcp Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * docs(network): clarify native tcp runtime constraints Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(network): remove unused transparent tcp pin Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(network): admit redirected transparent tcp Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(network): restore podman transparent networking Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(podman): permit alpine busybox binaries Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(podman): use portable alpine keepalive Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(podman): build musl networking fixture Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(podman): isolate musl DNS probe Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(podman): keep privileged port capability dropped Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(network): preserve transparent TCP port 53 Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(network): report synthetic pool pressure by family Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(podman): bind tcp fixtures before readiness Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(podman): grant fixture low-port bind Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> --------- Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> |
||
|
|
4d7f402ce2 |
feat(policy): establish direct TCP egress foundation (#2711)
* feat(policy): accept explicit tcp endpoint protocol Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * refactor(network): snapshot authoritative egress decisions Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * docs(policy): document explicit tcp protocol Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * docs(policy): defer transparent TCP release guidance Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * chore(go): regenerate sandbox protobuf bindings Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(network): complete tcp egress foundation Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * docs(policy): document explicit tcp contract Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(network): fail closed on authorization errors Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(podman): fence delayed exit events before restart Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(policy): validate network endpoint destinations Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(providers): opt in tcp credential fixture Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(policy): require dns host for transparent tcp Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> --------- Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> |
||
|
|
44bf0df485 |
feat(middleware): inspect WebSocket text messages (#2477)
* feat(middleware): inspect websocket text messages Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(middleware): address websocket review feedback Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(network): bound websocket message assembly Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(network): harden websocket upgrade lifecycle Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * refactor(middleware): unify in-process and remote transports Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * feat(middleware): support regex websocket redaction Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(middleware): bound persistent streaming sessions Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(middleware): accept websocket sequence gaps Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * refactor(middleware): refine websocket introspection contract Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(middleware): clarify websocket preflight lifecycle Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(middleware): clarify websocket coverage semantics Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(network): type websocket frame failures Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(network): return 503 when middleware admission is exhausted Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(middleware): align streaming API contract Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(middleware): clarify WebSocket event result scope Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(rfc): simplify middleware revision history Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * feat(examples): add WebSocket content guard support Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(middleware): unify binding payload limits Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * refactor(middleware): align payload limit terminology Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(middleware): address websocket review feedback Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(network): address websocket review findings Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * test(network): allow Linux handler setup in preflight regression Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(network): harden websocket relay finalization Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(network): inspect compressed websocket messages Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * test(network): stabilize compressed websocket regressions Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(go-sdk): regenerate middleware protobuf binding Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(middleware): clarify websocket skip lifecycle Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(middleware): address WebSocket review feedback Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> --------- Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> |
||
|
|
bdabb54cb3 |
fix(security): authenticate extension services (#2638)
* fix(security): authenticate extension services Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * feat(extension-core): verify gateway JWTs Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * refactor(extension-core): keep inbound verification external This should become an extension SDK package. Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(security): harden the extension authentication contract Follow-up hardening on the alpha extension authentication mechanism. Claim contract: - Extension tokens carry an explicit `typ` of `openshell-ext+jwt`. They share a signing key with sandbox-to-gateway admission tokens and were otherwise separated by audience alone, so a verifier that neglects to check `aud` could accept a gateway credential. The header is a second, independent discriminator. - Publish OIDC-shaped discovery at `/.well-known/openid-configuration` so a service configured with only the gateway URL can learn the exact expected issuer and the JWKS location. It is shaped, not compliant: `issuer` is the gateway identity, not the serving URL. Audience agreement: - `MiddlewareManifest` and `InterceptorManifest` gain `expected_audience`. The audience is otherwise configured independently on each side of the boundary, where a mismatch surfaces only as an opaque authentication failure on every call. OpenShell now compares the two and fails at startup. An empty field keeps the check off for existing services. Compatibility: - Add `allow_insecure_transport` per registration. Enabling gateway JWT signing previously made any plaintext endpoint a hard startup failure, including the endpoint form used in our own documentation. The opt-out attaches no credential, is refused by the gateway if a supervisor asks for one, and warns at every startup. - Make the transport requirement kind-aware. A middleware endpoint must be reachable from every sandbox supervisor, so only interceptors may use a gateway-local Unix socket. Credential lifecycle: - Replace the process-global slot map with a supervisor-owned `ExtensionCredentialStore` shared explicitly across the gateway connections the supervisor opens, removing test-order coupling. - Rotate only when a credential is missing or has passed four fifths of its lifetime. Configuration polling ran every ten seconds against fifteen-minute credentials, so each poll re-ran gateway effective-policy resolution and re-minted the gateway token. - Bound credential minting per sandbox, since each request resolves the caller's effective policy. Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs: record alpha extension authentication in RFC appendices Restore the RFC 0009 and 0010 bodies to their accepted text and move every extension-authentication update into appendices instead. An RFC records a decision at a point in time; superseding detail belongs alongside it rather than rewritten into it. RFC 0009's appendix carries the shared contract: claims, authorization, key distribution, the `allow_insecure_transport` replacement for the body's `allow_insecure`, and residual risks. RFC 0010's records only what differs for interceptors and links to it. The existing protocol-extensions appendix, which parked the phase 2 transport question, now points forward to what was built. Also document the audience handshake, the discovery endpoint, the `typ` requirement, and `jti` replay guidance in the extensibility and gateway configuration pages, and correct the middleware transport guidance: middleware endpoints must be reachable from sandbox supervisors, so Unix sockets are not an option there. Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * refactor(extension-core): abstract extension server trust Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(core): update middleware manifest example Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(extension-auth): preserve unsigned gateway compatibility Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(extension-auth): reject cross-domain token replay Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> --------- Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> |
||
|
|
bc14018cad |
feat(sandbox): use policy-first OCI image identity (#2509)
* feat(sandbox): use policy-first OCI image identity Closes #2331 Preserve per-field policy omission, derive Docker and Podman fallbacks from the inspected immutable image, and resolve the final numeric identity before starting agent children. Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> * fix(sandbox): preserve declared process identities Keep explicit policy values and OCI-declared names intact, defer passwd lookup until a primary GID is required, and refresh stale policy examples. Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> * fix(supervisor): reuse resolved OCI identity Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> * fix(supervisor): allow Linux pre-exec arguments Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> * fix(kubernetes): protect resolved sandbox identity Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> * fix(sandbox): prepare workspace for OCI identity Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> * refactor(sandbox): own only workspace root Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> * fix(sandbox): harden partial identity drops Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> * test(sandbox): scope OCI image e2e to Docker Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> * fix(sandbox): narrow OCI identity fallback scope Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> * test(podman): cover OCI identity launch Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> * fix(podman): exercise OCI fallback in E2E Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> --------- Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> |
||
|
|
2b7f04fe0a |
feat(examples): add supervisor middleware content guard (#2169)
* feat(examples): add supervisor middleware content guard Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(examples): refine middleware preview warning Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(examples): add middleware policy version Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * refactor(supervisor-middleware): simplify service endpoints Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(examples): adapt content guard to middleware enums Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(examples): align content guard with merged middleware Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * feat(examples): add content guard smoke flow Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * chore(examples): remove smoke launcher test Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(examples): align content guard smoke with main Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(examples): address content guard review feedback Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * refactor(examples): parse cargo metadata with jq Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(examples): prioritize longest content matches Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(examples): use GitHub warning alert Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(examples): merge overlapping content matches Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(examples): render preview warning on GitHub Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> --------- Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> |
||
|
|
75d24688e1 |
fix(examples): add missing workspace fields to governance interceptor (#2436)
The ListSandboxesRequest proto gained workspace and all_workspaces fields in the workspace feature but the governance interceptor example was not updated, causing a build failure. Set all_workspaces to true so the interceptor's policy reload covers sandboxes across all workspaces. Assisted-By: Claude (Anthropic AI) <noreply@anthropic.com> Signed-off-by: Pavel Anni <panni@redhat.com> |
||
|
|
5952a5a23f |
feat(workspace): add workspace resource model with scoping, membershi… (#2243)
* feat(workspace): implement workspace model (Phase 1 of RFC 0011) Implements workspace and membership model providing hard isolation boundaries for multi-player OpenShell deployments. Workspace CRUD with Kubernetes-style Terminating phase for graceful deletion. All resources scoped by workspace via ObjectMeta. Membership RPCs for workspace access control. Persistence migration shifts name uniqueness to (object_type, workspace, name). Provider profiles support platform and workspace scoping. Service routing uses workspace-prefixed DNS labels. Inference routes renamed and workspace-scoped with DeleteInferenceRoute RPC. Python SDK with WorkspaceClient, two-method list pattern (workspace-scoped and for_all_workspaces), and workspace parameter on all methods. CLI workspace flags, TUI workspace cycling. K8s driver filters unmanaged CRs and uses delete preconditions. Podman driver uses immutable container IDs. Label serialization fixed across all put_if call sites. Signed-off-by: Derek Carr <decarr@redhat.com> * fix(cli): delegate sandbox upload command to existing upload function The standalone `sandbox upload` command reimplemented upload logic inline with two bugs: it used `Path::exists()` which follows symlinks (rejecting dangling symlinks), and it ran git-aware filtering on symlink sources. The `run::sandbox_upload()` function already handles both cases correctly via `sandbox_upload_plan()`. Replace the inline logic with a call to the existing function. Signed-off-by: Derek Carr <decarr@redhat.com> * fix(e2e): shorten sandbox names and fix test compatibility Shorten the sandbox name in initial_sparse_policy_is_acknowledged_as_loaded from 'e2e-2159-sparse-enrich' (22 chars) to 'e2e-sparse-enrich' (17 chars) to comply with MAX_ROUTABLE_NAME_LEN (19 chars). Also capture stderr in create_keep_with_args so future sandbox creation failures include the actual CLI error instead of reporting empty output. Signed-off-by: Derek Carr <decarr@redhat.com> * test(workspace): add test coverage for workspace CRUD and persistence isolation Add unit tests for workspace create happy path, get round-trip, get not-found, get empty-name rejection, already-exists error, and resolve_workspace not-found. Add persistence test proving cross-workspace name uniqueness (same name in different workspaces produces separate records). Add workspace name max-length boundary tests. Fix e2e harness to include stderr in name-parse-failure error path. Align Python e2e test_workspace_crud with try/finally pattern. Document provider profile catalog workspace scoping gap in RFC 0011. Signed-off-by: Derek Carr <decarr@redhat.com> * fix(examples): update examples for workspace model compatibility Shorten sandbox names in demo scripts to fit the 19-character MAX_ROUTABLE_NAME_LEN limit: policy-demo prefix to pd-, multi-agent notepad derives a short SANDBOX_TAG from the run ID, governance interceptor uses gs-PID-RANDOM. Update vscode-remote-sandbox.md SSH host aliases from openshell-{name} to openshell-{name}.{workspace} format. Signed-off-by: Derek Carr <decarr@redhat.com> * feat(sdk): add workspace-scoped client and workspace CRUD Add WorkspaceScopedClient modeled after kube::Api::namespaced — captures workspace once and injects it into every sandbox request. Add workspace CRUD methods (create, get, list, delete) and list_sandboxes_all_workspaces on OpenShellClient. Extend SandboxRef with workspace field and add WorkspaceRef type. Include mock tests for all new operations. Signed-off-by: Derek Carr <decarr@redhat.com> * fix(lint): resolve clippy warnings in workspace test assertions Signed-off-by: Derek Carr <decarr@redhat.com> * fix(docs): convert indented code blocks to fenced in RFC 0011 Signed-off-by: Derek Carr <decarr@redhat.com> * fix(lint): resolve clippy warnings and apply cargo fmt across workspace Auto-format with cargo fmt and fix clippy warnings exposed by the reformat: unnecessary qualifications, map_unwrap_or, identical match arms, unused variable prefix, dead code annotations, and let-unit-value in e2e harness. Signed-off-by: Derek Carr <decarr@redhat.com> * fix(workspace): address workspace scoping issues from review - Add workspace field to settings JSON output (CLI) - Skip Podman containers missing workspace label instead of defaulting to empty string, matching K8s driver behavior - Add resource_version to list_by_scope SELECT in both SQLite and Postgres backends, with regression test - Gate PolicyLocalContext proposal/lookup routes on workspace readiness, returning 503 when workspace is not yet discovered - Block sandbox and provider creation in TUI all-workspaces mode - Clear workspace vectors in TUI reset_sandbox_state Signed-off-by: Derek Carr <decarr@redhat.com> * fix(workspace): make provider profile catalog workspace-aware Thread workspace through snapshot_catalog so the EffectiveProviderProfileCatalog enforces workspace boundaries on both read and write paths. UserProviderProfileSource now loads platform-scoped profiles (workspace "") plus the target workspace's profiles, preventing cross-workspace duplicate profile ID collisions that previously caused global catalog failures. Update RFC 0011 to reflect catalog scoping is implemented in Phase 1 rather than deferred to future work. Signed-off-by: Derek Carr <decarr@redhat.com> * fix(persistence): include workspace column in atomic policy revision INSERT put_policy_revision_atomic omitted the workspace column from the INSERT into the objects table in both SQLite and Postgres backends, causing atomically-written policy revisions to lose their workspace association. Add workspace field to AtomicPolicyRevisionWrite and thread it through both backend INSERT statements, matching the non-atomic put_policy_revision path which already included it. Signed-off-by: Derek Carr <decarr@redhat.com> * fix(proxy): skip ancestor walk when socket owner is the entrypoint collect_ancestor_identities walked the entire process tree above the entrypoint when the connecting process was the entrypoint itself, SHA256-hashing every ancestor binary (IDE, shell, container runtime). On dev machines with large binaries in the ancestor chain this exceeded the 30-second test timeout. When start_pid == stop_pid there are no intermediate ancestors to verify, so return an empty list immediately. Signed-off-by: Derek Carr <decarr@redhat.com> * fix(workspace): make provider profile catalog scope-aware Allow the same profile ID at platform and workspace scopes by introducing layered catalog entries where workspace profiles shadow platform profiles. Add source and scope fields to the ProviderProfile proto and CLI output. Migrate List/Get handlers to the catalog, fixing divergence with runtime profile resolution. Signed-off-by: Derek Carr <decarr@redhat.com> * fix(e2e): align podman e2e labels with centralized driver constants The podman driver moved its container labels to the centralized openshell.ai/ prefix, but the e2e test harness and cleanup script still referenced the old openshell.sandbox-* keys, causing the local_driver_token_restart test to fail on container lookup. Signed-off-by: Derek Carr <decarr@redhat.com> * fix(e2e): align python profile isolation test with scope-aware catalog Platform profiles are now visible in workspace listings as fallbacks per the layered catalog design. Update the assertion to match. Signed-off-by: Derek Carr <decarr@redhat.com> * fix(workspace): honor profile_workspace in runtime profile resolution Runtime profile lookups now consult provider.profile_workspace via get_type_profile_for_scope. Providers created with --global-profile (profile_workspace="") resolve to the platform profile even when a workspace profile shadows the same ID. All 6 runtime call sites updated; type-only call sites remain scope-agnostic. Signed-off-by: Derek Carr <decarr@redhat.com> --------- Signed-off-by: Derek Carr <decarr@redhat.com> |
||
|
|
8cf2673c0b |
docs: bump stated Rust MSRV from 1.88 to 1.90 (#2276)
* docs: bump stated Rust MSRV from 1.88 to 1.90 Cargo.toml sets rust-version = "1.90" (rust-toolchain.toml pins 1.95.0), so building with the previously documented 1.88 fails Cargo's MSRV check. * docs: bump e2e/rust MSRV to 1.90 * fix: align remaining Rust version fields to 1.90 examples/governance-interceptor/Cargo.toml still had rust-version 1.88. Also bump e2e/rust's prost dependency to 0.14 to match the workspace, since it was on 0.13 in an otherwise standalone crate. |
||
|
|
aa483ecb9a |
feat(providers): AWS STS AssumeRole refresh strategy and aws-s3 profile (#1782)
Add gateway-managed AWS STS credential refresh (provider-v2, #1576). The gateway calls sts:AssumeRole and writes three short-lived credentials (AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, AWS_SESSION_TOKEN) to the provider record; the proxy re-signs requests with SigV4. Adds the aws and aws-s3 provider profiles and a declarative multi-output refresh model (additional_outputs) so one AssumeRole co-mints all three credentials. Signed-off-by: Russell Bryant <rbryant@redhat.com> |
||
|
|
83003e80fc |
feat(interceptors): initial gateway interceptor implementation and reference example (#2005)
* feat(gateway): add descriptor-driven interceptors Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(gateway): add service-reflected interceptors Signed-off-by: Drew Newberry <anewberry@nvidia.com> * wip * fix(gateway): harden interceptor evaluation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(interceptors): label metrics and harden governance smoke Signed-off-by: Drew Newberry <anewberry@nvidia.com> * remove on_error: ignore * feat(gateway-interceptors): emit log annotations Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(examples): govern provider profiles in interceptor Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(gateway): preserve update config annotations Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(providers): support interceptor profile catalogs Signed-off-by: Drew Newberry <anewberry@nvidia.com> * wip * feat(governance-interceptor): sign provider profiles Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(providers): use configured profile sources for refresh updates Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(gateway-interceptors): add phase-specific evaluation payloads Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(providers): compose provider profile sources Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(gateway-interceptors): preserve committed responses Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(gateway-interceptors): reject ambiguous protobuf oneofs Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(gateway-interceptors): validate patch candidates per binding Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(gateway-interceptors): use reflected protobuf codec Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(governance-example): canonicalize signed protobuf hashes Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(gateway): commit policy provenance atomically Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(gateway): close signed governance bypasses Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(gateway): isolate interceptor secrets and authority Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(gateway): snapshot provider profiles per request Signed-off-by: Drew Newberry <anewberry@nvidia.com> * chore(gateway): resolve server clippy warnings Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(server): satisfy provider source clippy lint Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(server): initialize policy test annotations Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(gateway-interceptors): require explicit route allowlist Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(proto): clarify update annotation semantics Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
a72711697d |
chore: remove deprecated --keep flag from docs, scripts, and e2e tests (#2126)
* docs: remove deprecated --keep flag from tutorials and examples The --keep flag is deprecated, hidden, and a no-op since sandboxes are kept by default. Remove references from tutorial docs and example READMEs that explain it as a real feature. - Remove --keep from sandbox create commands - Remove --keep explanation text - Clarify that sandboxes are kept by default Signed-off-by: Ignas Baranauskas <ibaranau@redhat.com> * chore: remove deprecated --keep usage from scripts and e2e tests The --keep flag is a deprecated no-op since sandboxes are kept by default. Stop passing it in internal scripts, e2e test scripts, and example demo scripts. Signed-off-by: Ignas Baranauskas <ibaranau@redhat.com> --------- Signed-off-by: Ignas Baranauskas <ibaranau@redhat.com> |
||
|
|
9c14de7b83 | docs: fix article before OpenShell in sync-files (#2133) | ||
|
|
6461677c32 |
feat(policy): accept numeric UIDs for sandbox process identity (#1973)
* feat(policy): accept numeric UIDs in sandbox process identity validation Allow run_as_user and run_as_group to be either the literal 'sandbox' or a numeric UID/GID within [1000, 2_000_000_000]. This removes the hard dependency on a baked-in 'sandbox' user in container images, enabling compute drivers to inject resolved UIDs at sandbox creation. Phase 1 of #1959. Signed-off-by: Seth Jennings <sjenning@redhat.com> * feat(supervisor): accept numeric UIDs for process identity dropping Allow run_as_user and run_as_group to be numeric UIDs/GIDs, removing the hard dependency on a baked-in 'sandbox' user in container images. Changes: - validate_sandbox_user(): accepts numeric UIDs without passwd lookup (logs OCSF event); keeps passwd check for "sandbox" name; rejects non-numeric non-sandbox strings that fail passwd lookup - prepare_filesystem(): passes numeric UIDs/GIDs directly to chown() instead of requiring a passwd entry - drop_privileges(): resolves numeric UIDs/GIDs directly via UID::from_raw / Gid::from_raw; skips initgroups when target uid matches current euid; uses guard conditions before setgid/setuid calls - session_user_and_home(): falls back to ("{uid}", "/sandbox") for numeric UIDs, avoiding a passwd lookup that will fail Re-exports MIN_SANDBOX_UID and MAX_SANDBOX_UID from openshell-policy so callers have consistent range constants. Phase 2 of #1959. Signed-off-by: Seth Jennings <sjenning@redhat.com> * feat(driver-kubernetes): resolve sandbox UID/GID from config or OpenShift SCC annotations Phase 3 of the numeric-UID plan: allow operators to specify explicit sandbox_uid/sandbox_gid in Kubernetes driver config, auto-detect from OpenShift SCC namespace annotations, and propagate resolved values to supervisor container env vars and PVC init container securityContext. Changes: - Add sandbox_uid/sandbox_gid fields to KubernetesComputeConfig - Add SANDBOX_UID/SANDBOX_GID env var constants to openshell-core - Implement resolve_sandbox_identity() to fetch namespace annotations and auto-detect OpenShift SCC UID ranges (sa.scc.uid-range) - Pass resolved UID/GID through SandboxPodParams to pod spec builder - Inject SANDBOX_UID/SANDBOX_GID env vars into supervisor container - Update PVC init container securityContext with resolved UID/GID instead of hard-coded root - Add comprehensive unit tests for resolution logic and annotation parsing (resolve_sandbox_uid, resolve_sandbox_gid, OpenShift SCC annotation parsing) Signed-off-by: Seth Jennings <sjenning@redhat.com> * feat(driver-vm): add configurable sandbox UID/GID and update docs/examples Phase 4 of the numeric-UID plan: replace hardcoded SANDBOX_UID (10001) in VM rootfs preparation with configurable sandbox_uid/sandbox_gid fields. Changes: - Add sandbox_uid/sandbox_gid to VmDriverConfig with serde derives - Pass resolved UID/GID through prepare_sandbox_rootfs_from_image_root to ensure_sandbox_guest_user which writes /etc/passwd/group/gshadow - Update BYOC Dockerfile: remove groupadd/useradd, document runtime UID injection and the ability to skip baked-in sandbox user - Update gateway-config.mdx: document sandbox_uid/sandbox_gid for both Kubernetes (with OpenShift SCC autodetection) and VM drivers - Update sandbox-compute-drivers.mdx: add Sandbox User Identity section explaining numeric UID support across all compute drivers - Update rootfs tests to use non-default UIDs, verify config passthrough Signed-off-by: Seth Jennings <sjenning@redhat.com> * code review changes * fix(supervisor): harden tests for restricted CI container environments Guard tests against CI-specific constraints: root without CAP_SETPCAP, UIDs with no /etc/passwd entry, and restricted /proc access. Signed-off-by: Seth Jennings <sjennings@nvidia.com> Signed-off-by: Seth Jennings <sjenning@redhat.com> --------- Signed-off-by: Seth Jennings <sjenning@redhat.com> Signed-off-by: Seth Jennings <sjennings@nvidia.com> |
||
|
|
702cbc4f63 |
feat(providers): support SPIFFE-backed token grants (#1784)
* feat(providers): support SPIFFE-backed token grants Add provider profile token_grant metadata and expand endpoint-specific dynamic credentials so sandbox supervisors can request SPIFFE JWT-SVIDs, exchange them with an OAuth-style token endpoint, cache returned access tokens, and inject bearer tokens into matching HTTP requests. Wire Kubernetes and Helm deployments to mount the provider SPIFFE Workload API socket into sandbox pods for token grant exchange. Signed-off-by: Taylor Mutch <taylormutch@gmail.com> Signed-off-by: Gordon Sim <gsim@redhat.com> * test(examples): add SPIFFE token grant demo Add a reusable alpha/beta demo that deploys a SPIFFE-verifying token issuer and protected services, imports a token-grant provider profile, creates a sandbox, and verifies endpoint-specific bearer tokens. The script leaves Kubernetes workloads in place, deletes sandboxes through openshell unless KEEP_SANDBOX=1, and prints protected service logs as proof of life. Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(providers): harden SPIFFE token grants * fix(providers): harden dynamic token grants Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(providers): harden token grant handling --------- Signed-off-by: Taylor Mutch <taylormutch@gmail.com> Signed-off-by: Gordon Sim <gsim@redhat.com> |
||
|
|
8bf667f377 | fix: update RFC link in agent-driven-policy-management README (#1677) | ||
|
|
ae5127f14d |
fix: correct example paths in local-inference README (#1676)
* fix: correct example paths in local-inference README * fix: correct example paths in local-inference routes.yaml |
||
|
|
e98ea3ee93 | feat(policy): add agentic approval loop (#1528) | ||
|
|
603b3e27fa | docs: update NemoClaw/OpenClaw references (#1529) | ||
|
|
c5d1d76d94 |
refactor(sandbox): replace iptables with nftables for network policy enforcement (#1401)
Migrate all sandbox and VM driver network policy enforcement from iptables to nftables. nftables provides atomic ruleset loading, a cleaner rule syntax, and is the standard netfilter interface in modern kernels. Sandbox bypass enforcement (openshell-sandbox): - Replace iptables chain of individual rule insertions with a single atomic nftables ruleset load via nft -f - New nft_ruleset module with pure functions for ruleset generation and unit tests - Combine log and reject rules in one inet family table (handles both IPv4 and IPv6 in a single ruleset) - Fall back to reject-only ruleset when kernel lacks nft_log support - Enable net.netfilter.nf_log_all_netns so log rules work from non-init network namespaces - Use temp file for nft ruleset loading instead of stdin for compatibility with minimal VM guest environments VM TAP networking (openshell-driver-vm): - Replace iptables NAT/forwarding rules with nftables equivalents - New nft_ruleset module for TAP network rule generation with unit tests - Atomic table-per-TAP-device lifecycle (create/destroy) - Host-side rules provide NAT infrastructure and defense-in-depth isolation (input chain restricts VM to gateway port only, forward chain blocks unsolicited inbound); primary security enforcement happens inside the VM guest via the sandbox supervisor's own rules VM init script: - Load nft kernel modules at sandbox init - Enable nf_log_all_netns sysctl for bypass detection logging OCSF / docs: - Update firewall rule engine references from iptables to nftables - Document host firewall interaction model and two-layer enforcement architecture in VM driver README and compute drivers reference Closes #1335 Signed-off-by: Russell Bryant <rbryant@redhat.com> |
||
|
|
1c317646c6 |
docs: replace --sync with --upload . in sync-files example (#1366)
The --sync flag was never implemented. The correct flag is --upload with a path argument. Updated examples and workflow snippet to reflect the actual CLI interface. Fixes #1339 Signed-off-by: Mesut Oezdil <versusfinem@gmail.com> |
||
|
|
ea2fddbe2d |
feat(policy): agent-driven policy management — the agent half (#1323)
* feat(policy): plumb chunk_ids and rejection_reason through proposal pipeline Prereq plumbing for the agent revise-and-resubmit loop. Two narrow additive proto changes unblock the upcoming /wait endpoint (#1092), prover validation badge (#1097), and reject --guidance surfaces (#1098). - SubmitPolicyAnalysisResponse: add accepted_chunk_ids so the in-sandbox agent gets handles to watch its proposals. Surfaced through the typed grpc_client wrapper and policy.local's POST /v1/proposals 202 body. Closes #1094. - PolicyChunk + StoredDraftChunk + DraftChunkPayload: add validation_result (gateway prover verdict, populated by #1097) and rejection_reason (operator free-form text). Both plain strings; no enums, no parsing on the read path. Closes #1096. - RejectDraftChunk now persists the existing reason field into the chunk's rejection_reason so it round-trips back to the agent via GetDraftPolicy. UndoDraftChunk clears it on the way back to pending so consumers cannot read a stale guidance string from a prior reject -> re-approve -> undo cycle. Whole surface stays gated behind agent_policy_proposals_enabled. Two focused tests cover the round-trip and the undo-clears guarantee. Signed-off-by: Alexander Watson <zredlined@gmail.com> * feat(sandbox): add /v1/proposals/{id} and /wait long-poll to policy.local The agent feedback channel back from policy.local. Two new routes let the in-sandbox agent learn its proposal's outcome on a single blocking HTTP call — zero LLM tokens during the wait. - GET /v1/proposals/{chunk_id} returns the chunk's current state in one gateway call. - GET /v1/proposals/{chunk_id}/wait?timeout=<s> blocks until the chunk transitions out of pending. Default 60s, clamped [1, 300]. Agent re-issues on timeout to extend. Response carries the chunk's status plus the two feedback fields shipped in the prereq commit: rejection_reason (free-form reviewer text) and validation_result (gateway prover verdict, empty until #1097). On timeout: same shape with timed_out: true so the agent can disambiguate without parsing. Wait handler short-polls GetDraftPolicy every 1s inside the request with a tokio::time::Instant deadline. One gateway connection is opened per request and reused across all polls, so a 60s wait does one TLS handshake instead of sixty. A future commit can swap the loop body for a tokio::sync::broadcast driven by a watcher task — the agent-visible contract (URL, query, response shape) is independent of the polling implementation. All routes stay behind agent_policy_proposals_enabled. Closes #1092. Signed-off-by: Alexander Watson <zredlined@gmail.com> * docs(sandbox): teach policy_advisor skill the wait + redraft loop The agent-facing instructions for the feedback loop. The endpoints exist; this is the doc that makes them usable. policy_advisor.md gains: - API entries for GET /v1/proposals/{chunk_id} and /wait?timeout=<s>, including the field semantics (status, rejection_reason, validation_result, timed_out). - A note on the submit response's accepted_chunk_ids / rejection_reasons split so the agent handles partial acceptance. - Step 6 saves the chunk_ids and addresses any submit-time rejections before waiting. - Step 7 walks the four wait outcomes: approved (retry, with the honest "may still fail" caveat), rejected (read rejection_reason AND validation_result; address whichever has content), still-pending with timed_out (re-call), non-2xx (surface, do not retry). skills.rs gains two assertions on the skill content so a future edit cannot drop the wait endpoint or the rejection_reason directive silently. Closes #1095. Signed-off-by: Alexander Watson <zredlined@gmail.com> * test(policy-advisor): add end-to-end smoke for the agent feedback loop A focused smoke that exercises the new policy.local /wait endpoint on a live gateway + sandbox, separate from the existing no-LLM regression harness (which still drives the OLD retry-with-bash-loop recovery pattern). Two flows: - Flow A — approve-and-retry: agent submits, /wait blocks, host runs `openshell rule approve`, /wait returns status=approved. Confirms the happy path round-trip latency. - Flow B — reject-with-guidance: agent submits, /wait blocks, host runs `openshell rule reject --reason "..."`, /wait returns status=rejected with the exact reviewer text in rejection_reason. Confirms the free-form guidance contract round-trips through the agent feedback channel. No GitHub credentials needed — proposals are synthetic and never trigger outbound traffic. Both flows expect agent_policy_proposals_enabled=true and a running gateway. Adds three cases to sandbox-runner.sh: submit-test-proposal (no GH deps), proposal-status, proposal-wait. The existing put-file and submit-proposal cases used by test.sh are untouched. Signed-off-by: Alexander Watson <zredlined@gmail.com> * fix(policy-advisor): surface real CLI errors from wait-smoke preflight The preflight piped openshell's stderr to /dev/null and relied on jq to default the missing setting key to "<unset>", but under `set -euo pipefail` a non-zero exit from openshell makes the whole pipeline fail and the command substitution exits the script silently before the intended fail() message can print. Capture stderr explicitly, check the CLI exit code, and surface the real error plus the expected fix (port-forward + gateway add + select) when the CLI cannot reach the gateway. Signed-off-by: Alexander Watson <zredlined@gmail.com> * fix(policy-advisor): pass --json to settings get in wait-smoke preflight `openshell settings get --global` defaults to a human-readable table; jq cannot parse it and the preflight died with a numeric-literal error. Pass --json so jq gets actual JSON. Also touched up the suggested recovery commands in the preflight error to match the real CLI shape (`gateway add <endpoint> --name <name>` and the env-var override warning). Signed-off-by: Alexander Watson <zredlined@gmail.com> * fix(policy): dedup draft chunks only in mechanistic mode; return effective id The smoke harness for the agent feedback loop caught a real bug in the gateway: SubmitPolicyAnalysis's response carried a chunk_id that was never persisted whenever the SQL ON CONFLICT path fired. Two failure modes, both load-bearing: - Agent-authored proposals targeting the same host/port/binary (e.g. the redraft-after-rejection loop) silently folded into one row and any RejectDraftChunk by the new chunk_id failed with "chunk not found." Latent since #1151, surfaced by #1094 returning chunk_ids. - Mechanistic mode had the same class of bug — the dedup fold-in is the intended behavior there, but the response still advertised the newly-generated UUID instead of the existing row's id. Less visible because no current caller reads mechanistic chunk_ids back, but the proto contract was violated either way. Fix in three parts: - put_draft_chunk now takes Option<&str> dedup_key explicitly and returns the effective row id (via RETURNING). None binds NULL to the dedup_key column, which bypasses the partial-index ON CONFLICT path entirely. Caller-decides semantics replace store-side magic. - handle_submit_policy_analysis picks dedup_key per chunk using an allowlist (only "mechanistic" dedups) and pushes the returned effective_id to accepted_chunk_ids. New modes default to no-dedup so a misconfigured caller cannot silently lose proposals. - The two-copy draft_chunk_dedup_key helper consolidated to one observation_dedup_key in policy_store.rs with a doc comment. Tests: - agent_authored_submits_for_same_endpoint_do_not_dedup pins the redraft-loop contract: two intentional submissions with the same host/port/binary get distinct chunk_ids, both findable via GetDraftPolicy, both rejectable by id. - mechanistic_submits_for_same_endpoint_dedup_into_one_chunk locks in the observation-mode dedup AND asserts both submits return the same effective_id — would have caught the deeper bug. Proto: SubmitPolicyAnalysisRequest.analysis_mode doc updated to describe the actual semantics (mechanistic dedups, agent_authored and unknown modes do not). Signed-off-by: Alexander Watson <zredlined@gmail.com> * docs(examples): retarget policy-management demo at the /wait endpoint The narrated demo (examples/agent-driven-policy-management) has been the public face of this feature since #1151. Its agent prompt told Codex to retry the original PUT every few seconds for up to 120 seconds — a polling workaround for the missing /wait endpoint that this branch shipped. Update the demo to exercise /wait so the canonical reading of the feature reflects the actual UX win. - agent-task.md: step 4 is now "call /wait, branch on status" with the three outcomes spelled out (approved → retry once; rejected → read rejection_reason and revise or stop; pending+timed_out → re-issue /wait once, do NOT busy-loop or shorten the timeout). Also makes explicit that the demo submits one rule per proposal so accepted_chunk_ids[0] is the safe single id to wait on. - demo.sh: header docstring rewritten as a six-step loop that mirrors the README. narrate_sandbox_workflow drops its parallel numbering and uses bullets (the runtime narration is the agent's sub-actions, not a separate decomposition of the loop). Approve step header and success message now reference /wait waking the agent, not "policy hot-reload retry." - README.md: top-of-file flow expanded from 5 to 6 steps to include the /wait call and chunk_id capture; "Going further" section now describes both regression scripts and the boundary between them (real-GitHub retry vs. synthetic /wait wire test). Slow-path qualifier corrected from "image pull on first run" to "sandbox cold-start (SSH bring-up plus Codex install)". - wait-smoke.sh header rewritten to make it unambiguous this is a regression, NOT a tutorial, with explicit prereq commands instead of prereq descriptions, and a pointer at demo.sh for the narrated story. No code paths change; this is the readability pass. Signed-off-by: Alexander Watson <zredlined@gmail.com> * fix(examples): pass --yes on demo.sh's global setting writes Global setting updates require explicit confirmation in non-interactive mode; demo.sh's enable_agent_proposals and the cleanup restore path were missing --yes and hard-failed the preflight. Pre-existing issue that surfaced now that more of the demo runs through this path. No other behavior change. Signed-off-by: Alexander Watson <zredlined@gmail.com> * feat(sandbox): emit OCSF audit events for policy proposal lifecycle The demo's policy decision trace previously showed only the proxy enforcement story (HTTP:PUT DENIED, CONFIG:LOADED, HTTP:PUT ALLOWED). It was silent about who proposed what or who decided what — the audit-trail receipts for the agent feedback loop were missing. policy.local now emits sandbox-side OCSF events at the observation moments, into the same stream as the existing CONFIG:LOADED: - CONFIG:PROPOSED on submit_proposal acceptance. Per accepted chunk: the message names the chunk_id, target endpoint, L7 method/path, and binary so the trace correlates against the inbox card via chunk_id. - CONFIG:APPROVED on /wait observation of approved status. - CONFIG:REJECTED on /wait observation of rejected status. Carries the reviewer's free-form rejection_reason in the message AND as an unmapped field, both sanitized (control chars stripped, capped at 200 chars with an ellipsis marker). The agent still reads the raw text via GET /v1/proposals/{id}; sanitization is audit-side only, per AGENTS.md's no-secrets-in-OCSF rule. The submit path defends the audit_summaries / accepted_chunk_ids index pairing against a future gateway change that compresses past rejected chunks (the proto doesn't promise 1:1 ordering with the request). Today client-side validation makes the lengths always match; if they don't, the pairing falls back to a generic per-id event rather than mis-attribute. The wait handler's emit site fires once per terminal-status observation. Multiple concurrent waiters on the same chunk would each emit one event; acceptable for single-waiter-per-chunk demos and the right place to dedup is the SIEM. demo.sh's trace filter now surfaces the four CONFIG: events alongside HTTP:PUT, so the trace at the end of every run tells the full story from deny to allow via propose -> approve. wait-smoke.sh's prereq notes recommend redirecting kubectl port-forward output so its "Handling connection for 8090" lines don't bleed into demo narration. Three new unit tests on the sandbox-side helpers — summary builder happy path, fallback, and the rejection_reason sanitizer. Signed-off-by: Alexander Watson <zredlined@gmail.com> * feat(policy): /wait awaits local policy reload; demo auto-approves redrafts Three things in one commit, all surfaced by running the demo end-to-end against a real gateway and finding the agent had to draft a broader second proposal. 1. /wait race fix. Previously /wait returned `approved` the moment it observed the gateway's chunk status flip, but the local supervisor reloads policy on its own poll cycle (~10s in practice). The agent's retry would race the reload and hit the still-old policy, getting denied. Codex then drafted a broader rule and re-submitted — sound agent behavior, but not what /wait should provoke. Now /wait captures the local policy version at start, and after observed-approved waits for the supervisor to load a strictly-newer version before returning. Bounded by the caller's deadline; best-effort return if the deadline elapses without the version bumping. Two new unit tests pin the happy path and the deadline-clamped fallback. 2. demo.sh auto-approve loop. Replaces approve_when_pending + wait_for_agent with one approve_pending_until_agent_exits function that keeps watching for pending chunks and approving them until the agent process exits (or the configured timeout). Defense in depth against future redraft scenarios for any reason; today (post-fix #1) the agent should only submit one proposal per task, but we don't want to hang silently if it does submit more. 3. UX. Step headers now carry "[t+1.2s]" relative timestamps so reading the run output makes latency visible (the demo's whole point is the wait is cheap — surface that). A spin_wait helper renders an ASCII spinner during the watch loop so the demo never looks frozen on a TTY. Falls back to plain sleep on non-TTY contexts. Closes the race condition diagnosed from the trace timing where the gateway approved at t+0, sandbox observed at t+0.3s, but the supervisor didn't load v2 until t+9.4s — well after the agent had already retried and been denied. Signed-off-by: Alexander Watson <zredlined@gmail.com> * fix(sandbox): /wait detects policy reload by content, not the schema version The previous attempt at the /wait-after-approve race fix compared `SandboxPolicy.version` between /wait start and the policy reload — but that field is the *schema* version (constant 1), not a revision counter. Every comparison was `current(1) > baseline(1) == false`, so the wait blocked until the agent's 300s timeout regardless of whether the supervisor had actually reloaded. The demo SSH connection then timed out around the 240s mark. Diagnosed from a live run's OCSF trace: supervisor pulled v2 at +8.5s after approval (CONFIG:LOADED), but the sandbox-side CONFIG:APPROVED that my /wait emits didn't fire until +304s — exactly at the 300s deadline. Fix: compare the whole policy via prost's derived PartialEq. Any field change (network_policies map being the only one that actually mutates today) flips equality. A clone-per-200ms-tick on a few-KB proto is cheap inside the bounded wait window. Tests rewritten to match the new contract: the supervisor-reload fixture now keeps `version: 1` constant and changes `network_policies` contents, mirroring the exact failure mode from the live run. Signed-off-by: Alexander Watson <zredlined@gmail.com> * fix(examples): redact tokens with python literal-string replace, not sed The sed-based redact_log in demo.sh broke when one of the auth tokens contained a character that conflicted with sed's pattern parser ("unterminated substitute pattern" on the Codex JWT). The whole log tail then blanks on failure, hiding the very failure context we're trying to surface. Switch to a python subprocess that takes the tokens via argv and does literal str.replace. No regex, no delimiter games, no truncation. Signed-off-by: Alexander Watson <zredlined@gmail.com> * fix(sandbox): scope /wait reload check to the approved rule Reviewer (John Myers) flagged two failure modes in the prior whole-policy fingerprint approach used by policy.local /wait: - False sleep: when the supervisor reloads between two /wait calls (the skill tells the agent to re-issue on timed_out), the new call snapshots the already-updated policy as baseline and burns the full timeout waiting for a change that never comes. - False wakeup: any unrelated reload (other agent's approval, settings change) flips the diff, but the chunk's actual rule may not be loaded yet — the agent retries and hits policy_denied for no real signal. Replace the diff with rule-coverage. New public helper openshell_policy::policy_covers_rule reuses endpoints_overlap (so it matches add_rule's merge semantics, including the fold-into-existing-key case) plus an L7 allow check on method/path (so an existing endpoint that doesn't yet contain the proposed method doesn't signal coverage). Add policy_reloaded: true|false to the /wait response on approve, with a 500ms floor on the reload-wait phase so approvals arriving near the deadline still get a fair shot at reloaded=true. Update the policy_advisor skill to branch on it: reloaded=true → retry; reloaded=false → re-issue /wait once with timeout=30, then surface to user. Don't loop tightly. Tests: - 9 new unit tests in openshell-policy pinning coverage semantics (L4-only, L7 method gap, fold-into-existing-key, empty binaries). - 4 new tokio tests in policy_local mirroring John's exact scenarios. - wait-smoke.sh asserts policy_reloaded=true on Flow A. Signed-off-by: Alexander Watson <zredlined@gmail.com> * fix(server): make GetDraftPolicy dual-auth so /wait works under OIDC policy.local calls GetDraftPolicy from inside the sandbox supervisor via the sandbox gRPC client, which authenticates with the shared x-sandbox-secret. GetDraftPolicy was listed only in the Bearer-auth scope table (config:read) and was not in SANDBOX_SECRET_METHODS or DUAL_AUTH_METHODS, so OIDC-enabled gateways rejected those calls and the /wait long-poll surfaced gateway_lookup_failed. Local/no-OIDC setups happened to work because the auth check is short-circuited. Add GetDraftPolicy to DUAL_AUTH_METHODS, matching the existing GetSandboxConfig pattern (called by both CLI reviewer surfaces with Bearer and the sandbox supervisor with x-sandbox-secret). Dual-auth short-circuits the scope check for sandbox-secret callers, so the config:read entry in authz.rs continues to gate Bearer-only flows. Mirror the openshell_get_sandbox_config_is_dual_auth assertion for GetDraftPolicy. Note: ssh_handshake_secret is server-wide, not per-sandbox, so a sandbox-secret caller can today name any sandbox in a SubmitPolicyAnalysis request — and now in a GetDraftPolicy request. The exposure is symmetric with the existing SANDBOX_SECRET_METHODS pattern. Filed as a follow-up: per-sandbox secret binding, tracked separately. Signed-off-by: Alexander Watson <zredlined@gmail.com> * fix(ci): address rebased check failures Signed-off-by: Alexander Watson <zredlined@gmail.com> --------- Signed-off-by: Alexander Watson <zredlined@gmail.com> |
||
|
|
1c79b2131f |
feat: agent-driven policy management MVP (#1151)
* docs(rfc): add agent-driven policy management * docs(rfc): switch policy MVP to local API * docs(rfc): clarify policy advisor skill and local logs * feat(sandbox): add agent-driven policy proposal loop * test(examples): add codex policy dogfood loop * refactor(examples): make policy demo agent-agnostic * refactor(examples): colocate policy validation harness * docs(examples): add policy demo env sample * docs(examples): use placeholder env example * feat(sandbox): wire policy.local denials to OCSF JSONL log Wires GET /v1/denials?last=N on the sandbox-local policy advisor API to read recent OCSF JSONL events from /var/log/openshell-ocsf.YYYY-MM-DD.log, filter to network/L7 denials (action_id=2, class_uid 4001/4002), and return a compact summary newest-first. Default limit is 10, capped at 100. Ran inside spawn_blocking so file I/O does not block the policy.local handler. Other cleanup: - POST /v1/proposals now uses the typed grpc_client wrapper instead of raw_client, so accepted/rejected counts surface to the agent uniformly. Wrapper return type extended to the response struct. - Drop the 'add_rule' snake_case alias in the proposal JSON; canonical form is camelCase 'addRule', matching the PolicyMergeOperation convention used elsewhere. - skills/policy_advisor.md updated to match: documents the now-real /v1/denials?last=10 endpoint and uses 'addRule' consistently. - skills.rs test asserts on the canonical 'addRule' phrase rather than the removed 'PolicyMergeOperation' substring. * feat(cli): show L7 protocol/method/path in rule get output format_endpoint() previously rendered only host:port, dropping protocol, access, and the L7 rules array. That made openshell rule get text output unable to distinguish a broad L4 grant from a method/path-scoped L7 REST rule -- exactly the distinction a developer needs at approval time. New rendering tags each endpoint with its enforcement layer and surfaces allow/deny rules: bare L4: api.example:443 [L4] L7 read-only: api.example:443 [L7 rest, access=read-only] L7 method/path: api.example:443 [L7 rest, allow PUT /v1/foo/bar] Pure display change: no proto, gateway, or behavior changes. Unit test covers all three rendering cases with synthetic fixtures. * refactor(examples): rewrite policy demo as Codex-default loop Re-shape examples/agent-driven-policy-management/ to be a single, clean end-to-end demonstration of the agent-driven policy loop. A Codex agent inside an OpenShell sandbox attempts a GitHub Contents API write, hits a structured 403 from the L7 proxy, reads the policy_advisor skill, drafts a narrow addRule proposal via http://policy.local/v1/proposals, the host auto-approves, the sandbox hot-reloads policy, and the agent's retry succeeds. Whole loop runs in roughly two minutes. Demo cleanup: - Drop .env file ceremony. Defaults resolve from gh: owner via 'gh api user --jq .login', repo defaults to 'openshell-policy-demo', token from gh auth token / GITHUB_TOKEN / GH_TOKEN. With gh auth login and codex login already done, 'bash demo.sh' Just Works. - Codex-specific. Bootstraps ~/.codex/auth.json from credentials injected by the OpenShell provider, runs codex exec --sandbox danger-full-access (OpenShell is the actual security boundary; bwrap nesting cannot create user namespaces inside the sandbox container). - Tighter narrative output: a single 'Preflight' step, a run summary banner before launch, an inline narration of what's happening inside the sandbox while we poll for the proposal (including the literal structured 403 body the agent acts on), and an OCSF trace at the end filtered to the three events that tell the story (DENY, RELOAD, ALLOW). - Replace Python heredoc templating with sed; uploads use the single-flag pattern (--upload "${PAYLOAD_DIR}:/sandbox") with files referenced at the basename-prefixed path that #952 / #1028 established. - README documents the trust model honestly: structured rule is the contract, agent rationale is a hint, prover validation badge in progress per RFC 0001. Move the deterministic no-LLM regression harness out of examples/ into e2e/policy-advisor/ -- it was a parallel demo, not an example. Same loop without the LLM, useful for iterating on the proxy and policy.local API. * style(sandbox,cli): apply rustfmt Whitespace-only fixups caught by mise run pre-commit. No functional change. * perf(examples): cap Codex reasoning at 'low' in policy demo The demo task is mechanical (one HTTP request, parse a structured 403, post a JSON proposal, retry). Codex's default high-effort reasoning roughly doubles the demo's wall time without improving outcomes; running at 'low' lands the same minimal L7 grant in roughly half the time. Override with DEMO_CODEX_REASONING=medium (or higher) to compare runs. * fix(sandbox): harden policy.local denials endpoint Three changes addressing review feedback before merging the agent-driven policy management MVP: - Distinguish "OCSF JSONL enabled, no denials" from "OCSF JSONL disabled, nothing to read." The endpoint now returns a `log_available` flag and an explanatory `note` when the log file is missing, so the in-sandbox agent can give the developer an accurate hint instead of a misleading empty list. - Stop echoing the OCSF `message` field in the per-denial summary. The proxy's denial messages can include the request path with query string (e.g., `?access_token=...`); the structured `host`/`port`/`method`/ `path`/`binary` fields carry everything the agent needs to draft a proposal, and `path` is sourced from `http_request.url.path` which already excludes the query string. - Cap `read_request_body` at a 15s timeout. Bounds slowloris-style stalls from a misbehaving in-sandbox process. The proxy listener only accepts loopback connections so practical impact is small, but this is cheap defense-in-depth. New tests cover the missing-log signal and the message-redaction guarantee. * fix(examples): redact tokens in agent log tail and validate DEMO_FILE_DIR Two small hardening passes on the policy management demo: - `fail()` now pipes the agent log tail through a redactor that masks the GitHub token and Codex credential triple before printing. Codex itself is well-behaved about not echoing the token, but a misbehaving tool call could leak it; this is a final safety net before the log hits the developer's terminal (and any clipboard or chat history that follows). - `validate_env` now regex-checks DEMO_FILE_DIR with the same allow-list the other path-shaped variables use. The value is interpolated through sed with `|` as the delimiter when rendering the agent task; rejecting unsupported characters keeps the templating predictable and stops a user-supplied value from breaking out into a shell context. * refactor(sandbox): centralize policy.local routes and skill path Addresses review feedback that the deny body's `next_steps` array and the route table could drift apart. The route paths and skill location now live as `pub const`s in `policy_local.rs` and feed both: - the dispatcher in `route_request` that matches against them - a new `agent_next_steps()` helper that builds the JSON the L7 deny body embeds `l7/rest.rs::deny_response_body` calls `policy_local::agent_next_steps()` instead of inlining the array, so adding or renaming a route is a one-line change in `policy_local.rs` and the agent contract follows automatically. * feat(sandbox): switch /v1/denials to shorthand log pass-through Previously /v1/denials parsed `/var/log/openshell-ocsf.*.log` (OCSF JSONL) and returned structured per-event objects. JSONL is opt-in via `ocsf_json_enabled`, so the endpoint returned an empty list with a "log not enabled" hint by default — agents had to navigate a setup step before the inspect-recent-denials guidance was useful. Switch to reading the shorthand log at `/var/log/openshell.*.log`, which is always-on and the same human-readable format `openshell logs` displays. The endpoint now returns raw shorthand lines (newest first) — the agent reads them directly, no field parsing. Tradeoffs: - Removes the JSONL-on-by-default debate: shorthand is already on, no defaults change. - Updating shorthand is a single-file change in this repo; no schema rev needed when we want to add fields. Implementation: - `read_recent_denial_lines` walks shorthand log files newest-first, filters lines with ` OCSF ` AND ` DENIED ` (the OCSF action label, uppercase, space-bounded). - `collect_shorthand_log_files` matches `openshell.<date>.log`; the trailing dot in `SHORTHAND_LOG_PREFIX = "openshell."` excludes `openshell-ocsf.<date>.log` so JSONL-on doesn't bleed into responses. - 4096-byte cap per surfaced line as defense against pathological inputs. - Skill doc updated to reflect that `/v1/denials` returns raw shorthand lines, not structured fields. Defense-in-depth on query-string secrets: - `redact_query_strings` strips `?<query>` to `?[redacted]` from each surfaced line. The L7 relay path emits OCSF events using `redacted_target` (secret-placeholder redaction), but the FORWARD deny path in `proxy.rs` populates `OcsfUrl::new("http", host, path, port)` and `.message(...)` with the raw request path — query string included. Stripping queries at the consumer guards `/v1/denials` regardless of whether the upstream emit sites are tightened. The on-disk log is not rewritten by this change; that is a separate hardening task tracked for the FORWARD path emit sites in proxy.rs. - `truncate_at_char_boundary` is UTF-8 safe; redaction runs before truncation so a cut cannot slice mid-secret. Tests: - `recent_denials_returns_newest_first_from_shorthand_lines` covers the happy path with mixed allowed/denied/non-OCSF lines. - `recent_denials_skips_jsonl_log_files` confirms JSONL files don't surface even if present. - `recent_denials_truncates_pathological_lines` covers the cap. - `is_ocsf_denial_line_filters_correctly` covers the line-level filter. - `redact_query_strings_removes_query_from_url_token` and `redact_query_strings_removes_query_in_reason_tag` cover the redaction in both URL token and `[reason:...]` contexts. - `truncate_at_char_boundary_does_not_panic_on_multibyte` covers the UTF-8 safety. * chore(sandbox): align proto inits with main's L7 GraphQL additions Post-rebase fixups after #1083 (GraphQL L7 inspection) landed on main and introduced new fields on the proto types this branch constructs: - `crates/openshell-sandbox/src/l7/relay.rs`: two `deny_with_redacted_target` call sites (REST and GraphQL relay deny paths) now pass the `DenyResponseContext` argument that `rest::send_deny_response` expects. Both sites pass `host`, `port`, and `binary` from the existing `L7EvalContext`, matching the pattern used at the primary deny site. - `crates/openshell-sandbox/src/policy_local.rs`: `L7Allow`, `L7DenyRule`, and `NetworkEndpoint` proto initializers now populate the new GraphQL and path-scoping fields with empty defaults. Agent-authored proposals via `policy.local` target REST/SQL/L4 today; GraphQL operation matching is set on the gateway side or via direct YAML, so empty defaults are correct here. No behavior change. `cargo test -p openshell-sandbox --lib` (650 tests) and `cargo clippy -p openshell-sandbox --lib --tests -- -D warnings` clean. * feat(sandbox): gate agent policy proposals behind opt-in feature flag The agent-driven policy proposal surface delivered by this PR (skill install, `policy.local` API, `next_steps` array on L7 deny bodies) is now opt-in via the new `agent_policy_proposals_enabled` setting. Default false. Same shape as `providers_v2_enabled`: registered in `openshell-core::settings`, sandbox-level, hot-toggleable via the existing settings poll loop. Why: the surface is a novel agent-controlled mutation point in every sandbox. The per-proposal developer approval gate is a correctness control, but it doesn't address "should this sandbox have an agent-authoring API at all" — compliance teams may want that question closed. The flag is the second gate. Implementation: - New registry entry + `AGENT_POLICY_PROPOSALS_ENABLED_KEY` constant in `openshell-core::settings`. - `lib.rs`: process-wide `OnceLock<Arc<AtomicBool>>` mirroring the `OCSF_CTX` pattern. `agent_proposals_enabled()` is the single read point. - Initial settings fetch added to `run_sandbox` so skill install honors the flag at startup (not just on the poll loop's first tick). - Skill install in `run_sandbox` is gated on the flag. - `policy_local::route_request` returns `404 feature_disabled` for all routes when the flag is off — including the otherwise-public `current_policy` and `denials` routes. When the surface is off it's off entirely. - `policy_local::agent_next_steps` returns an empty array when the flag is off so deny bodies don't advertise routes that 404. - Poll loop updates the atomic on each tick, lazily installs the skill on a false→true transition (no claw-back on true→false; stale skill on disk is harmless because route + next_steps gate on the live atom). Tests: - Shared `test_helpers::ProposalsFlagGuard` mutex+atomic guard for the process-wide flag, used across `policy_local::tests` and `l7::rest::tests`. - New: `agent_next_steps_returns_empty_when_flag_off`, `agent_next_steps_returns_full_array_when_flag_on`, `route_request_returns_feature_disabled_when_flag_off`. - Updated existing tests that exercise the deny body or the route dispatcher to set the flag on first. - Full sandbox lib test suite: 653 pass, clippy clean. Demo and e2e: - `examples/agent-driven-policy-management/demo.sh` and `e2e/policy-advisor/test.sh` now snapshot the prior global value of the setting, set it to true before sandbox creation (so the supervisor's initial poll picks it up), and restore on exit (delete if previously unset, otherwise write the prior value back). Docs: - RFC 0001 MVP-implementation note documents the flag, default, and intended soft-launch posture. * test(policy-advisor): require proposal opt-in for e2e * refactor(sandbox): group policy poll loop state * test(e2e): isolate Kubernetes user namespace test --------- Co-authored-by: John Myers <9696606+johntmyers@users.noreply.github.com> |
||
|
|
316c788eac |
fix(helm): derive grpcEndpoint from chart context (#1241)
* fix(helm): derive grpcEndpoint from chart context The chart hardcoded server.grpcEndpoint to https://openshell.openshell.svc.cluster.local:8080, which only matched the in-cluster Service DNS for the standard release name and namespace. A new helper now builds <scheme>://<fullname>.<namespace>.svc.cluster.local:<port> from chart context, picking the scheme from server.disableTls. An explicit server.grpcEndpoint override is passed through verbatim. * chore(scripts): validate k3d cluster name length early helm-k3s-local.sh derives the cluster name from the current branch suffix. Long branch names produced names exceeding k3d's 32-char cap and failed deep inside k3d cluster create with a confusing validation error. cmd_create now bails out before invoking docker/k3d with a copy-pasteable HELM_K3S_CLUSTER_NAME override hint. Status, start, stop, delete, and help remain unaffected so an over-long derived name does not block diagnostics. |
||
|
|
70a0f6c547 | refactor(cli): remove gateway lifecycle management (#1221) | ||
|
|
142a3a3d31 |
fix(examples): harden multi-agent notepad 409 retry and improve docs (#1166)
* fix(examples): harden multi-agent notepad 409 retry and improve docs Add decorrelated jitter to the GitHub Contents API retry loop so racing workers don't retry in lockstep, and widen the retry set to 5xx for transient server errors. Without jitter, N concurrent workers trip the same backoff schedule and re-collide on each beat. Also: - Extract the inline 220-line sandbox runner from demo.sh into a separate runner.sh. demo.sh stays as host orchestration; runner.sh is the per-sandbox program. demo.sh shrinks 460 to 240 lines. - Rewrite README to lead with "Why GitHub as a shared notepad?" and add a "Memory architecture variants" section covering pile, append-journal, and indexed-memory patterns. - Switch Quick Start to use \`gh auth token\` as the primary path. - Pass the worker slice arg explicitly instead of reading \$9 from the outer dispatcher. - Add .gitignore for personal run_demo.sh helpers. * docs updates |
||
|
|
f56c09c7df | docs: update gateway deployment architecture (#1108) | ||
|
|
25c4fdecd1 | fix(examples): repair multi-agent notepad uploads (#1152) | ||
|
|
2e0afeabe1 | feat(vm): derive guest rootfs from sandbox images (#957) | ||
|
|
ebbd9dee5a |
docs(examples): add multi-agent notepad demo (#991)
Signed-off-by: Alexander Watson <zredlined@gmail.com> |
||
|
|
de9dce0436 |
fix(e2e): use high UID range to avoid host user conflicts (#978)
Change sandbox user UID from 1000 to 1000660000 in custom image examples and E2E tests. Using a high UID range (1000000000+) prevents conflicts with host users when running without user namespace remapping, where container UIDs map directly to host UIDs. This resolves fork failures caused by RLIMIT_NPROC enforcement when the host user already has many threads running. Signed-off-by: Derek Carr <decarr@redhat.com> |
||
|
|
df38d1f66f | feat(ci): add Markdown and Mermaid linting (#933) |