mirror of
https://github.com/NVIDIA/OpenShell.git
synced 2026-10-02 07:34:45 +08:00
main
21
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
1374672967 |
fix(helm): restore Kubernetes e2e chart rendering (#3692)
Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
52cb8ecee7 |
fix(kubernetes): remove NetworkPolicy acknowledgement (#3677)
Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
c1f2e7189f |
feat(isolation): implement the RFC 0012 sandbox architecture (#2942)
* feat(isolation): add RFC 0012 backend contract Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> * refactor(isolation): name the interface crate explicitly Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): expose trusted host gateway Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(agents): inventory the MXC driver Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): add mediated DNS transport Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): tighten interface error and digest contracts Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(isolation): remove unrelated driver inventory Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): define capability-free launch contract Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): seal confirmed boundary state Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): validate confirmation for external backend implementations Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): clarify mediated DNS identity Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(isolation): generalize loopback connector Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(isolation): unify typed network mediation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): bind launches to sandbox sessions Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(mxc): initialize extended sandbox status Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): add boundary protocol and Linux primitives Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): harden signals and separate process status from transport Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): validate remote confirmation through public contract Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): validate wire state and propagate snapshot failures Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(isolation): import owned agent specification explicitly Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(isolation): describe mediated DNS channel Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): bound mediation attach without nested retries Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(isolation): generalize loopback protocol Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): add transport-neutral session authentication Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(isolation): separate sandbox backend protocol Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): harden runtime boundary controls Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): add terminal boundary operation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): split supervisor and sandbox runtimes Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): harden boundary isolation and lifecycle ownership Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): reject private root redirects and adopt typed errors Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): preserve accept thread ownership on musl Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(sandbox): isolate credential probes from filtered threads Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): return retained exec exit status to independent waiters Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): bound network mediation and preserve socket authorization Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): bound control admission and retire stale mediation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * ci(e2e): select migrated drivers per stack layer Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(sandbox): implement loopback connector Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): authenticate the Sandbox Protocol Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(supervisor): rotate launch-scoped authentication Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(sandbox): consume dedicated backend crate Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(sandbox): align topology session fixture Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): align projected bootstrap bundle Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(auth): validate refreshed credentials before rotation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): fail closed across supervisor disconnects Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): repair rebased sandbox CI Signed-off-by: Drew Newberry <anewberry@nvidia.com> * build(runtime): publish separate sandbox and supervisor images Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(config): configure the sandbox runtime image Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(ci): validate sandbox binary linkage Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(isolation): use backend and runtime terminology Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(sandbox): use a scratch runtime image Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(ci): refresh schema and dependency policy Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): bind reconnects to supervisor process Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs: align runtime split operational guidance Signed-off-by: Drew Newberry <anewberry@nvidia.com> * chore(security): document Kubernetes runtime RBAC Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): enforce runtime lifecycle invariants Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(compute): identify sandbox start generations Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(server): restore sandbox launch sessions Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): support authenticated runtime replacement Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(auth): bind sandbox session successors Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(auth): retry pending sandbox successors Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(vm): run the supervisor outside the guest workload Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(vm): use sandbox backend protocol Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): repair rebase integration Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): use unified build toolchain Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(vm): use sandbox runtime terminology Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(vm): own guest network bootstrap Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): expose guest init version Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): select native supervisor artifacts Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): guard guest init Linux symbols Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): scope Linux test imports Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): avoid guest interface casts Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): reconcile admitted sandbox identity Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): share resolved sandbox identity Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): surface host supervisor failures Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): include guest logs on supervisor exit Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): rotate and clean runtime generations Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): make sandbox starts generation-aware Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): rotate restored sandbox sessions Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(vm): keep shared paths in the base layer Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): bind sandbox session lineage Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(docker): isolate workloads behind the host supervisor Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(docker): rotate launch-scoped authentication Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(docker): use sandbox backend protocol Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(docker): use host networking for supervisor Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): preserve host gateway alias resolution Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(docker): use separate sandbox and supervisor images Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): restore startup validation after rebase Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(docker): name the sandbox runtime directly Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): narrow supervisor CA runtime storage Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): close companion isolation gaps Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(docker): align mediated network expectations Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(docker): exercise mediated network paths Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): attach supervisor to managed network Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): defer supervisor recovery until gateway is ready Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): make sandbox starts generation-aware Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): rotate restored sandbox sessions Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): preserve workloads during session rotation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(docker): remove unrelated configuration RFC changes Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): bind sandbox session lineage Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(kubernetes): add proxy-pod isolation topology Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(kubernetes): use sandbox backend protocol Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): use stable sandbox service authority Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(kubernetes): split sandbox and supervisor images Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): adapt proxy pods to current runtime APIs Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(kubernetes): describe the single runtime placement Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(kubernetes): simplify sandbox orchestration Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): validate deployment prerequisites Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): update Trivy Helm profile inventory Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(kubernetes): update Trivy scan inventory count Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): reuse preloaded runtime images in e2e Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): type and clean runtime resources Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): make sandbox restarts recoverable Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): rotate restored sandbox sessions Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): preserve supervisor egress Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): bind sandbox session lineage Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(podman): adopt isolated sandbox and supervisor containers Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): stage bootstrap archives at named volume destinations Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(podman): rotate launch-scoped authentication Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(podman): use sandbox backend protocol Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(podman): use host networking for supervisor Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(podman): split sandbox and supervisor images Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): repair rebase integration Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(podman): name the sandbox runtime directly Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): provision supervisor CA runtime storage Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): address isolation review findings Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): inspect Debian supervisor provenance Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): use libpod-compatible tmpfs options Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): bind verified sandbox runtime binary Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): provide external driver data directory Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): start sandbox before joining user namespace Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): separate supervisor user namespace Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): make sandbox starts generation-aware Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): rotate restored sandbox sessions Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): bind sandbox session lineage Signed-off-by: Drew Newberry <anewberry@nvidia.com> * perf(isolation): add TCP and DNS benchmark harnesses Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(perf): align benchmark timing and supported protocols Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(perf): report TCP benchmark metrics accurately Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(perf): cancel failed worker startup Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): build matching local supervisor image Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): make local sandbox smoke test runnable Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): wire local sandbox runtime image Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): narrow sandbox service RBAC Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(ci): validate split runtime artifacts Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): harden runtime session handling Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(supervisor): add standalone network proxy role Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(rfc): remove implementation companion notes Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(vm): standardize runtime release name Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): pin renamed runtime artifacts Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(auth): persist sandbox runtime identity Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(runtime): restore branch validation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): reconcile main after rebase Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(network): close unframed HTTP 1.0 responses Signed-off-by: Drew Newberry <anewberry@nvidia.com> * chore(isolation): preserve upstream OCSF updates Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(security): close credential and TLS replay paths Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(auth): make sandbox refresh retries idempotent Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
cc4ded2088 |
feat(helm): split gateway and workspace charts (#2643)
* feat(helm): split gateway and workspace charts Signed-off-by: Dhiraj Bokde <dbokde@nvidia.com> * fix(helm): preserve split chart upgrade compatibility Keep workspace manifests valid after value validation and default legacy reused values to the combined resource topology. * fix(ci): preserve VM runtime for E2E The Rust cache restores target/ after VM runtime artifacts are staged, overwriting target/vm-runtime-compressed before openshell-driver-vm is built. Stage the compressed runtime outside target and pass that location through OPENSHELL_VM_RUNTIME_COMPRESSED_DIR so build.rs can embed the supervisor. Also locate the Helm split-ownership test repository root from the script path rather than git rev-parse. The test runs in a container where the GitHub checkout can be owned by a different UID and rejected as dubious ownership. Signed-off-by: Dhiraj Bokde <dbokde@nvidia.com> * fix(ci): install yq for Helm ownership test The split-chart ownership regression uses yq to inspect rendered YAML, but the Helm CI container installs only tools declared in mise. Declare and lock yq so mise install --locked provides the test dependency. Signed-off-by: Dhiraj Bokde <dbokde@nvidia.com> --------- Signed-off-by: Dhiraj Bokde <dbokde@nvidia.com> |
||
|
|
d0dfb22baf |
feat(kubernetes): export driver traces over OTLP (#2958)
Mirror the VM, Podman, and Docker driver tracing setup for Kubernetes. Export standalone driver spans through OTLP/gRPC as the distinct openshell-driver-kubernetes service, preserve gateway trace context, record lifecycle operations and gRPC failures, and flush spans on shutdown. Kubernetes currently runs in-process when selected as a built-in gateway driver. Use the temporary server-boundary shim shared with Podman and Docker so traces retain the shape they will have when Kubernetes moves to a separate process. Move the common ComputeDriver RPC tracing layer into openshell-otel to keep all drivers aligned. Propagate the active W3C context through the controller-reserved Sandbox annotation and enable Agent Sandbox OTLP export in the local k3s workflow. This connects asynchronous controller reconciliation spans to the originating OpenShell create trace. Expose gateway OTLP configuration through Helm and add an Aspire collector to the local k3s workflow. Extend helm:k3s:forward with OTLP ingest and trace UI forwarding for Kubernetes and local container gateway development. Signed-off-by: Kris Hicks <khicks@nvidia.com> |
||
|
|
c399342649 |
feat(dev): unify local Kubernetes gateway workflow (#2914)
Make the local k3s gateway workflow match the Docker and Podman flows by registering and selecting successful plaintext Skaffold deployments with the OpenShell CLI. Derive the registration name from the worktree-specific k3d cluster name so parallel worktrees retain independent gateway metadata. Add helm:k3s:forward as the standard way to expose the Kubernetes gateway on localhost:8090, and update the development and debugging guidance to use the active registered gateway instead of one-off endpoint flags. Signed-off-by: Kris Hicks <khicks@nvidia.com> |
||
|
|
3be2cd8a29 |
fix(helm): preflight Agent Sandbox APIs (#2867)
* fix(helm): preflight Agent Sandbox APIs Signed-off-by: Evan Lezar <elezar@nvidia.com> * test(kubernetes): share Agent Sandbox setup Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(e2e): wait for Agent Sandbox CRD status Signed-off-by: Evan Lezar <elezar@nvidia.com> * ci(canary): sparse-checkout sandbox helper Signed-off-by: Evan Lezar <elezar@nvidia.com> --------- Signed-off-by: Evan Lezar <elezar@nvidia.com> |
||
|
|
2f96c53b8c |
feat(gateway,cli): windows compilation support (#2496)
* chore(windows): gate Unix-only workspace code for MSVC Signed-off-by: Shailendra Singh <shailendras@nvidia.com> Signed-off-by: Akber Raza <akberr@nvidia.com> * feat(windows): stub unsupported compute drivers Signed-off-by: Shailendra Singh <shailendras@nvidia.com> Signed-off-by: Akber Raza <akberr@nvidia.com> * ci(windows): add MSVC mise build lane Signed-off-by: Shailendra Singh <shailendras@nvidia.com> Signed-off-by: Akber Raza <akberr@nvidia.com> * docs(windows): document MSVC build-only design Signed-off-by: Shailendra Singh <shailendras@nvidia.com> Signed-off-by: Akber Raza <akberr@nvidia.com> * docs(agent): add Windows MSVC build skill Signed-off-by: Shailendra Singh <shailendras@nvidia.com> Signed-off-by: Akber Raza <akberr@nvidia.com> * feat(windows): add Windows build support Signed-off-by: Akber Raza <akberr@nvidia.com> * refactor(windows): consolidate Windows-specific dependencies and improve build logic Signed-off-by: Akber Raza <akberr@nvidia.com> * feat(windows): add libclang path resolution and update cargo commands with bundled Z3 features Signed-off-by: Akber Raza <akberr@nvidia.com> * chore(tooling): lock Windows tool artifacts Signed-off-by: Giedrius Burachas <gburachas@nvidia.com> Signed-off-by: Akber Raza <akberr@nvidia.com> * feat(windows): enhance libclang path resolution to support architecture-specific subdirectories Signed-off-by: Akber Raza <akberr@nvidia.com> * Fix Windows dependency gating after sync merge Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(z3): update Z3 header path requirements in Windows build documentation and scripts Signed-off-by: Akber Raza <akberr@nvidia.com> * docs(windows): relocate Windows MSVC build design to architecture/ Why: windows-msvc-build-design.mdx is a design document ("design decisions for the native Windows MSVC build lane"), but it lived in the published, user-facing docs/reference/ tree. Per AGENTS.md (Documentation) and architecture/README.md ("rfc/ vs architecture/"), design content belongs in architecture/ (or rfc/), not in published reference. It also shared Fern sidebar "position: 6" with the MXC compute-driver design page, colliding in the Reference nav ordering. What: - Move docs/reference/windows-msvc-build-design.mdx -> architecture/windows-msvc-build.md. - Strip the Fern publish frontmatter and add a plain H1, matching the other architecture docs. - Register it in the architecture doc index in architecture/README.md. - Repoint the inbound references (build-openshell-mxc-windows skill + reference, implement-openshell-mxc-driver skill) to the new path. With both design pages moved out of docs/reference/, the duplicate position-6 sidebar collision is resolved. Signed-off-by: Akber Raza <akberr@nvidia.com> * remove openshell-supervisor-network from unsupported driver package test exclusion list Signed-off-by: Akber Raza <akberr@nvidia.com> # Conflicts: # tasks/scripts/windows-msvc.ps1 * fix(interceptors): gate unix-only imports so the crate builds on Windows openshell-gateway-interceptors failed to compile on Windows (E0432: no UnixStream in tokio::net), breaking any Windows build of openshell-server (which depends on it unconditionally). The connect_unix_endpoint fn was already #[cfg(unix)]-gated, but the imports it uses (UnixStream, TokioIo, Uri, service_fn) were left ungated. Gate those four imports with #[cfg(unix)] too. No behavior change on unix; Windows now compiles (no errors, no unused-import warnings). Signed-off-by: Akber Raza <akberr@nvidia.com> * feat(windows): add native ARM64 test support Signed-off-by: Shailendra Singh <shailendras@nvidia.com> Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(mise): skip Skaffold on Windows Signed-off-by: Shailendra Singh <shailendras@nvidia.com> Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(windows): harden ARM64 toolchain discovery Signed-off-by: Shailendra Singh <shailendras@nvidia.com> Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(windows): scope ARM64 toolchain preflight Signed-off-by: Shailendra Singh <shailendras@nvidia.com> Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(windows): restore compatibility after GitHub sync Signed-off-by: Shailendra Singh <shailendras@nvidia.com> Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(windows): avoid rate-limited Z3 source lookup Signed-off-by: Shailendra Singh <shailendras@nvidia.com> Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(mise): skip Helm checks on Windows Signed-off-by: Shailendra Singh <shailendras@nvidia.com> Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(windows): support repository pre-commit checks Signed-off-by: Shailendra Singh <shailendras@nvidia.com> Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(windows): stabilize native MSVC validation Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(windows): harden shared Z3 source cache Signed-off-by: Shailendra Singh <shailendras@nvidia.com> * fix(windows): avoid leaking MSVC flags into clang-cl Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(windows): complete ARM64 migration audit Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(windows): restore ARM64 Ninja discovery Signed-off-by: Akber Raza <akberr@nvidia.com> * refactor(windows): separate platform crate roots Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(windows): restore proto include cfg gating Signed-off-by: Akber Raza <akberr@nvidia.com> * refactor: address lint errors * fix(windows): add preflight check for proxy auth file path * docs(windows): update GitHub checkout guidance Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(windows): restore CI after dependency updates Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(mise): repair Windows sccache lock entry Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(windows): reconcile validation after rebase Signed-off-by: Akber Raza <akberr@nvidia.com> * refactor(server): exclude unsupported drivers on Windows Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * refactor(server): isolate platform driver config Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(windows): repair unsupported driver contract test Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(sandbox): remove stale dependencies Signed-off-by: Akber Raza <akberr@nvidia.com> * ci(windows): pin x64 workflow actions Signed-off-by: Akber Raza <akberr@nvidia.com> * ci(windows): align x64 Rust toolchain Signed-off-by: Akber Raza <akberr@nvidia.com> * ci(windows): align ARM64 workflow setup Signed-off-by: Akber Raza <akberr@nvidia.com> * refactor(windows): exclude unsupported runtime crates Signed-off-by: Akber Raza <akberr@nvidia.com> * refactor(windows): exclude unsupported crates at workspace boundary Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com> * refactor(server): gate builtin driver config by platform Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com> * fix(sandbox): restore crate documentation Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com> * ci(windows): make build workflow manual Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com> * ci(windows): temporarily enable pull request builds Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com> * ci(windows): cache Rust dependencies Signed-off-by: Akber Raza <akberr@nvidia.com> * refactor(windows): remove unnecessary platform changes Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com> * ci(windows): make build workflow manual Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com> * fix(ci): synchronize mise lockfile Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com> * fix(ci): normalize mise provenance metadata Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com> * refactor(python): isolate Windows atomic replace retry Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com> * fix(python): type Windows permission test errors Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com> --------- Signed-off-by: Shailendra Singh <shailendras@nvidia.com> Signed-off-by: Akber Raza <akberr@nvidia.com> Signed-off-by: Giedrius Burachas <gburachas@nvidia.com> Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com> Co-authored-by: Shailendra Singh <shailendras@nvidia.com> Co-authored-by: Giedrius Burachas <gburachas@nvidia.com> Co-authored-by: Jamie King <jamiek@nvidia.com> Co-authored-by: Piotr Mlocek <pmlocek@nvidia.com> Co-authored-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com> |
||
|
|
8eacb4779f |
feat(kubernetes): add sidecar supervisor topology (#2076)
* feat(kubernetes): add sidecar supervisor topology Add the Kubernetes sidecar supervisor topology, its Helm/Skaffold configuration, topology documentation, and sidecar e2e matrix coverage. Skip root-only sandbox identity rewriting when process enforcement is network-only so the low-permission sidecar process container can start successfully. Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(supervisor): avoid similar process id names Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(supervisor): avoid similar process id names Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(sandbox): avoid similar proxy id names Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * docs(kubernetes): clarify sidecar topology limits Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(kubernetes): keep sidecar process leaf capless Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(kubernetes): refresh sidecar provider env snapshots Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * test(supervisor): align hot-swap identity regression Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(kubernetes): stage sidecar mtls files before proxy chown Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(kubernetes): simplify sidecar supervisor topology Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * chore(helm): reuse sidecar skaffold values Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(supervisor): avoid similar iptables helper names Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(e2e): harden kube gateway wrapper setup Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(supervisor): avoid nft batch rollback on OCP Run nftables setup as individual commands so optional conntrack and log expressions can fail without rolling back required table, chain, and reject rules. Signed-off-by: Seth Jennings <sjenning@redhat.com> Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(kubernetes): preserve process identity in sidecar topology Render sidecar pods with a shared process namespace, keep binary-aware network policy enabled, and move Kubernetes sidecar settings under the nested sidecar config table. Also apply unprivileged Landlock/seccomp setup in NetworkOnly supervisor mode so sidecar topology keeps sandbox child hardening without privileged process setup. Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * refactor(kubernetes): replace sidecar snapshots with control socket Coordinate sidecar policy and provider bootstrap over a local Unix socket so the process leaf no longer reads policy/provider snapshot files. Report entrypoint startup through the control channel and keep gateway credentials confined to the network sidecar. Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * feat(kubernetes): support relaxed sidecar network identity Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(sandbox): satisfy sidecar clippy lint Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * refactor(kubernetes): standardize topology naming Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(sandbox): satisfy linux clippy timeout import Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(kubernetes): support kata sidecar on ipv4 pods Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(kubernetes): satisfy linux clippy for sidecar fallback Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * chore(kubernetes): remove stale supervisor topology references Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(kubernetes): enable sidecar binary policy inspection Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(kubernetes): harden sidecar control boundary Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(kubernetes): couple sidecar supervisor lifecycles Signed-off-by: Taylor Mutch <taylormutch@gmail.com> --------- Signed-off-by: Taylor Mutch <taylormutch@gmail.com> Signed-off-by: Seth Jennings <sjenning@redhat.com> Co-authored-by: Seth Jennings <sjenning@redhat.com> |
||
|
|
70fed042cb |
fix(helm): build chart dependencies before lint (#1947)
Signed-off-by: Evan Lezar <elezar@nvidia.com> |
||
|
|
c4ca283c1a | refactor(helm): require external postgres for ha (#1844) | ||
|
|
5007042e79 |
feat(helm): add optional PostgreSQL backing store (#1579)
* feat(helm): add optional PostgreSQL backing store with Secret-based credentials - Add postgres.enabled and postgres.deploy values to control database backend (SQLite vs PostgreSQL) and subchart deployment independently. - Introduce db-secret.yaml template for Opaque Secret with assembled postgresql:// connection string injected via OPENSHELL_DB_URL env var. - Add Bitnami PostgreSQL as optional subchart dependency keyed on postgres.deploy to prevent subchart deployment in external mode. - Externalize JWT signing key file mode via sandboxJwt.secretDefaultMode with 0400 default matching upstream. - Add validation guard for postgres.deploy=true without postgres.enabled. - Add helm unit tests covering internal, external, URL-override, special character encoding, and misconfiguration error paths. - Update README with Kubernetes and OpenShift install examples for bundled and external PostgreSQL configurations. - Add helm dependency build to lint and unittest tasks. * fix(helm): add database backend docs to README.md.gotmpl and regenerate The helm-docs CI check failed because the Database backend section was added directly to README.md instead of README.md.gotmpl. Move the content to the template and regenerate so the check passes. * fix(helm): use Secret-based DB credentials and support existingSecret Replace the inline db-url stringData pattern with a proper Secret containing individual fields plus a uri key. When postgres.deploy=true the Bitnami service-binding secret is referenced directly; when deploy=false users can supply postgres.external.existingSecret to bring their own Secret, or let the chart generate one from the external field values. Also restructures the README database section for clarity, adds helm-unittest coverage for the new secret resolution paths, and fixes a markdown lint issue in the root README. * refactor(helm): move OpenShift e2e script to e2e/rust/ and add mise task Move test-openshift-scenarios.sh from deploy/helm/openshell/ci/ to e2e/rust/e2e-openshift.sh, matching the existing e2e script naming convention. Register it as `e2e:openshift` in tasks/test.toml — not wired into the `test` or `e2e` aggregates so it only runs on explicit invocation against a live OpenShift cluster. * feat(e2e): add database backend scenarios to Kubernetes e2e Extend with-kube-gateway.sh with an optional multi-scenario loop gated by OPENSHELL_E2E_KUBE_DB_SCENARIOS=1. When enabled, the script installs the Helm chart three times — SQLite (default), bundled PostgreSQL, and external PostgreSQL with existingSecret — running the full test suite against each backend. When unset, existing single-install behavior is unchanged. Also adds helm dependency build before helm install, fixing CI failures caused by the missing PostgreSQL subchart dependency. * refactor(helm): simplify PostgreSQL config to two orthogonal controls Replace postgres.deploy and postgres.external.* with two simple controls: - postgres.enabled: deploy the bundled Bitnami PostgreSQL subchart - server.externalDbSecret: name of a pre-existing Secret with a uri key Delete db-secret.yaml — the chart no longer generates Secrets from individual credential fields. Users either get the Bitnami service-binding secret (bundled) or bring their own via server.externalDbSecret. Add validation that postgres.serviceBindings.enabled must stay true when using bundled PostgreSQL, preventing a confusing runtime failure. |
||
|
|
a7cd1608f3 |
docs(helm): add chart readme generation (#1437)
Signed-off-by: Taylor Mutch <taylormutch@gmail.com> |
||
|
|
dbba580e84 |
fix(security): refresh CI and gateway image dependencies (#1432)
Refresh the CI image tool pins so Go-built tools are rebuilt with patched Go releases and move the sandbox Python runtime to 3.14.5. Rebase the gateway runtime to a pinned distroless Debian 13 image with glibc 2.41-12+deb13u3 while preserving the existing UID/GID 1000 runtime identity for upgrade compatibility. Update rustls-webpki to 0.103.13 and clarify Linux k3d guidance now that k3d is not installed through mise on Linux. Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> |
||
|
|
ea2fddbe2d |
feat(policy): agent-driven policy management — the agent half (#1323)
* feat(policy): plumb chunk_ids and rejection_reason through proposal pipeline Prereq plumbing for the agent revise-and-resubmit loop. Two narrow additive proto changes unblock the upcoming /wait endpoint (#1092), prover validation badge (#1097), and reject --guidance surfaces (#1098). - SubmitPolicyAnalysisResponse: add accepted_chunk_ids so the in-sandbox agent gets handles to watch its proposals. Surfaced through the typed grpc_client wrapper and policy.local's POST /v1/proposals 202 body. Closes #1094. - PolicyChunk + StoredDraftChunk + DraftChunkPayload: add validation_result (gateway prover verdict, populated by #1097) and rejection_reason (operator free-form text). Both plain strings; no enums, no parsing on the read path. Closes #1096. - RejectDraftChunk now persists the existing reason field into the chunk's rejection_reason so it round-trips back to the agent via GetDraftPolicy. UndoDraftChunk clears it on the way back to pending so consumers cannot read a stale guidance string from a prior reject -> re-approve -> undo cycle. Whole surface stays gated behind agent_policy_proposals_enabled. Two focused tests cover the round-trip and the undo-clears guarantee. Signed-off-by: Alexander Watson <zredlined@gmail.com> * feat(sandbox): add /v1/proposals/{id} and /wait long-poll to policy.local The agent feedback channel back from policy.local. Two new routes let the in-sandbox agent learn its proposal's outcome on a single blocking HTTP call — zero LLM tokens during the wait. - GET /v1/proposals/{chunk_id} returns the chunk's current state in one gateway call. - GET /v1/proposals/{chunk_id}/wait?timeout=<s> blocks until the chunk transitions out of pending. Default 60s, clamped [1, 300]. Agent re-issues on timeout to extend. Response carries the chunk's status plus the two feedback fields shipped in the prereq commit: rejection_reason (free-form reviewer text) and validation_result (gateway prover verdict, empty until #1097). On timeout: same shape with timed_out: true so the agent can disambiguate without parsing. Wait handler short-polls GetDraftPolicy every 1s inside the request with a tokio::time::Instant deadline. One gateway connection is opened per request and reused across all polls, so a 60s wait does one TLS handshake instead of sixty. A future commit can swap the loop body for a tokio::sync::broadcast driven by a watcher task — the agent-visible contract (URL, query, response shape) is independent of the polling implementation. All routes stay behind agent_policy_proposals_enabled. Closes #1092. Signed-off-by: Alexander Watson <zredlined@gmail.com> * docs(sandbox): teach policy_advisor skill the wait + redraft loop The agent-facing instructions for the feedback loop. The endpoints exist; this is the doc that makes them usable. policy_advisor.md gains: - API entries for GET /v1/proposals/{chunk_id} and /wait?timeout=<s>, including the field semantics (status, rejection_reason, validation_result, timed_out). - A note on the submit response's accepted_chunk_ids / rejection_reasons split so the agent handles partial acceptance. - Step 6 saves the chunk_ids and addresses any submit-time rejections before waiting. - Step 7 walks the four wait outcomes: approved (retry, with the honest "may still fail" caveat), rejected (read rejection_reason AND validation_result; address whichever has content), still-pending with timed_out (re-call), non-2xx (surface, do not retry). skills.rs gains two assertions on the skill content so a future edit cannot drop the wait endpoint or the rejection_reason directive silently. Closes #1095. Signed-off-by: Alexander Watson <zredlined@gmail.com> * test(policy-advisor): add end-to-end smoke for the agent feedback loop A focused smoke that exercises the new policy.local /wait endpoint on a live gateway + sandbox, separate from the existing no-LLM regression harness (which still drives the OLD retry-with-bash-loop recovery pattern). Two flows: - Flow A — approve-and-retry: agent submits, /wait blocks, host runs `openshell rule approve`, /wait returns status=approved. Confirms the happy path round-trip latency. - Flow B — reject-with-guidance: agent submits, /wait blocks, host runs `openshell rule reject --reason "..."`, /wait returns status=rejected with the exact reviewer text in rejection_reason. Confirms the free-form guidance contract round-trips through the agent feedback channel. No GitHub credentials needed — proposals are synthetic and never trigger outbound traffic. Both flows expect agent_policy_proposals_enabled=true and a running gateway. Adds three cases to sandbox-runner.sh: submit-test-proposal (no GH deps), proposal-status, proposal-wait. The existing put-file and submit-proposal cases used by test.sh are untouched. Signed-off-by: Alexander Watson <zredlined@gmail.com> * fix(policy-advisor): surface real CLI errors from wait-smoke preflight The preflight piped openshell's stderr to /dev/null and relied on jq to default the missing setting key to "<unset>", but under `set -euo pipefail` a non-zero exit from openshell makes the whole pipeline fail and the command substitution exits the script silently before the intended fail() message can print. Capture stderr explicitly, check the CLI exit code, and surface the real error plus the expected fix (port-forward + gateway add + select) when the CLI cannot reach the gateway. Signed-off-by: Alexander Watson <zredlined@gmail.com> * fix(policy-advisor): pass --json to settings get in wait-smoke preflight `openshell settings get --global` defaults to a human-readable table; jq cannot parse it and the preflight died with a numeric-literal error. Pass --json so jq gets actual JSON. Also touched up the suggested recovery commands in the preflight error to match the real CLI shape (`gateway add <endpoint> --name <name>` and the env-var override warning). Signed-off-by: Alexander Watson <zredlined@gmail.com> * fix(policy): dedup draft chunks only in mechanistic mode; return effective id The smoke harness for the agent feedback loop caught a real bug in the gateway: SubmitPolicyAnalysis's response carried a chunk_id that was never persisted whenever the SQL ON CONFLICT path fired. Two failure modes, both load-bearing: - Agent-authored proposals targeting the same host/port/binary (e.g. the redraft-after-rejection loop) silently folded into one row and any RejectDraftChunk by the new chunk_id failed with "chunk not found." Latent since #1151, surfaced by #1094 returning chunk_ids. - Mechanistic mode had the same class of bug — the dedup fold-in is the intended behavior there, but the response still advertised the newly-generated UUID instead of the existing row's id. Less visible because no current caller reads mechanistic chunk_ids back, but the proto contract was violated either way. Fix in three parts: - put_draft_chunk now takes Option<&str> dedup_key explicitly and returns the effective row id (via RETURNING). None binds NULL to the dedup_key column, which bypasses the partial-index ON CONFLICT path entirely. Caller-decides semantics replace store-side magic. - handle_submit_policy_analysis picks dedup_key per chunk using an allowlist (only "mechanistic" dedups) and pushes the returned effective_id to accepted_chunk_ids. New modes default to no-dedup so a misconfigured caller cannot silently lose proposals. - The two-copy draft_chunk_dedup_key helper consolidated to one observation_dedup_key in policy_store.rs with a doc comment. Tests: - agent_authored_submits_for_same_endpoint_do_not_dedup pins the redraft-loop contract: two intentional submissions with the same host/port/binary get distinct chunk_ids, both findable via GetDraftPolicy, both rejectable by id. - mechanistic_submits_for_same_endpoint_dedup_into_one_chunk locks in the observation-mode dedup AND asserts both submits return the same effective_id — would have caught the deeper bug. Proto: SubmitPolicyAnalysisRequest.analysis_mode doc updated to describe the actual semantics (mechanistic dedups, agent_authored and unknown modes do not). Signed-off-by: Alexander Watson <zredlined@gmail.com> * docs(examples): retarget policy-management demo at the /wait endpoint The narrated demo (examples/agent-driven-policy-management) has been the public face of this feature since #1151. Its agent prompt told Codex to retry the original PUT every few seconds for up to 120 seconds — a polling workaround for the missing /wait endpoint that this branch shipped. Update the demo to exercise /wait so the canonical reading of the feature reflects the actual UX win. - agent-task.md: step 4 is now "call /wait, branch on status" with the three outcomes spelled out (approved → retry once; rejected → read rejection_reason and revise or stop; pending+timed_out → re-issue /wait once, do NOT busy-loop or shorten the timeout). Also makes explicit that the demo submits one rule per proposal so accepted_chunk_ids[0] is the safe single id to wait on. - demo.sh: header docstring rewritten as a six-step loop that mirrors the README. narrate_sandbox_workflow drops its parallel numbering and uses bullets (the runtime narration is the agent's sub-actions, not a separate decomposition of the loop). Approve step header and success message now reference /wait waking the agent, not "policy hot-reload retry." - README.md: top-of-file flow expanded from 5 to 6 steps to include the /wait call and chunk_id capture; "Going further" section now describes both regression scripts and the boundary between them (real-GitHub retry vs. synthetic /wait wire test). Slow-path qualifier corrected from "image pull on first run" to "sandbox cold-start (SSH bring-up plus Codex install)". - wait-smoke.sh header rewritten to make it unambiguous this is a regression, NOT a tutorial, with explicit prereq commands instead of prereq descriptions, and a pointer at demo.sh for the narrated story. No code paths change; this is the readability pass. Signed-off-by: Alexander Watson <zredlined@gmail.com> * fix(examples): pass --yes on demo.sh's global setting writes Global setting updates require explicit confirmation in non-interactive mode; demo.sh's enable_agent_proposals and the cleanup restore path were missing --yes and hard-failed the preflight. Pre-existing issue that surfaced now that more of the demo runs through this path. No other behavior change. Signed-off-by: Alexander Watson <zredlined@gmail.com> * feat(sandbox): emit OCSF audit events for policy proposal lifecycle The demo's policy decision trace previously showed only the proxy enforcement story (HTTP:PUT DENIED, CONFIG:LOADED, HTTP:PUT ALLOWED). It was silent about who proposed what or who decided what — the audit-trail receipts for the agent feedback loop were missing. policy.local now emits sandbox-side OCSF events at the observation moments, into the same stream as the existing CONFIG:LOADED: - CONFIG:PROPOSED on submit_proposal acceptance. Per accepted chunk: the message names the chunk_id, target endpoint, L7 method/path, and binary so the trace correlates against the inbox card via chunk_id. - CONFIG:APPROVED on /wait observation of approved status. - CONFIG:REJECTED on /wait observation of rejected status. Carries the reviewer's free-form rejection_reason in the message AND as an unmapped field, both sanitized (control chars stripped, capped at 200 chars with an ellipsis marker). The agent still reads the raw text via GET /v1/proposals/{id}; sanitization is audit-side only, per AGENTS.md's no-secrets-in-OCSF rule. The submit path defends the audit_summaries / accepted_chunk_ids index pairing against a future gateway change that compresses past rejected chunks (the proto doesn't promise 1:1 ordering with the request). Today client-side validation makes the lengths always match; if they don't, the pairing falls back to a generic per-id event rather than mis-attribute. The wait handler's emit site fires once per terminal-status observation. Multiple concurrent waiters on the same chunk would each emit one event; acceptable for single-waiter-per-chunk demos and the right place to dedup is the SIEM. demo.sh's trace filter now surfaces the four CONFIG: events alongside HTTP:PUT, so the trace at the end of every run tells the full story from deny to allow via propose -> approve. wait-smoke.sh's prereq notes recommend redirecting kubectl port-forward output so its "Handling connection for 8090" lines don't bleed into demo narration. Three new unit tests on the sandbox-side helpers — summary builder happy path, fallback, and the rejection_reason sanitizer. Signed-off-by: Alexander Watson <zredlined@gmail.com> * feat(policy): /wait awaits local policy reload; demo auto-approves redrafts Three things in one commit, all surfaced by running the demo end-to-end against a real gateway and finding the agent had to draft a broader second proposal. 1. /wait race fix. Previously /wait returned `approved` the moment it observed the gateway's chunk status flip, but the local supervisor reloads policy on its own poll cycle (~10s in practice). The agent's retry would race the reload and hit the still-old policy, getting denied. Codex then drafted a broader rule and re-submitted — sound agent behavior, but not what /wait should provoke. Now /wait captures the local policy version at start, and after observed-approved waits for the supervisor to load a strictly-newer version before returning. Bounded by the caller's deadline; best-effort return if the deadline elapses without the version bumping. Two new unit tests pin the happy path and the deadline-clamped fallback. 2. demo.sh auto-approve loop. Replaces approve_when_pending + wait_for_agent with one approve_pending_until_agent_exits function that keeps watching for pending chunks and approving them until the agent process exits (or the configured timeout). Defense in depth against future redraft scenarios for any reason; today (post-fix #1) the agent should only submit one proposal per task, but we don't want to hang silently if it does submit more. 3. UX. Step headers now carry "[t+1.2s]" relative timestamps so reading the run output makes latency visible (the demo's whole point is the wait is cheap — surface that). A spin_wait helper renders an ASCII spinner during the watch loop so the demo never looks frozen on a TTY. Falls back to plain sleep on non-TTY contexts. Closes the race condition diagnosed from the trace timing where the gateway approved at t+0, sandbox observed at t+0.3s, but the supervisor didn't load v2 until t+9.4s — well after the agent had already retried and been denied. Signed-off-by: Alexander Watson <zredlined@gmail.com> * fix(sandbox): /wait detects policy reload by content, not the schema version The previous attempt at the /wait-after-approve race fix compared `SandboxPolicy.version` between /wait start and the policy reload — but that field is the *schema* version (constant 1), not a revision counter. Every comparison was `current(1) > baseline(1) == false`, so the wait blocked until the agent's 300s timeout regardless of whether the supervisor had actually reloaded. The demo SSH connection then timed out around the 240s mark. Diagnosed from a live run's OCSF trace: supervisor pulled v2 at +8.5s after approval (CONFIG:LOADED), but the sandbox-side CONFIG:APPROVED that my /wait emits didn't fire until +304s — exactly at the 300s deadline. Fix: compare the whole policy via prost's derived PartialEq. Any field change (network_policies map being the only one that actually mutates today) flips equality. A clone-per-200ms-tick on a few-KB proto is cheap inside the bounded wait window. Tests rewritten to match the new contract: the supervisor-reload fixture now keeps `version: 1` constant and changes `network_policies` contents, mirroring the exact failure mode from the live run. Signed-off-by: Alexander Watson <zredlined@gmail.com> * fix(examples): redact tokens with python literal-string replace, not sed The sed-based redact_log in demo.sh broke when one of the auth tokens contained a character that conflicted with sed's pattern parser ("unterminated substitute pattern" on the Codex JWT). The whole log tail then blanks on failure, hiding the very failure context we're trying to surface. Switch to a python subprocess that takes the tokens via argv and does literal str.replace. No regex, no delimiter games, no truncation. Signed-off-by: Alexander Watson <zredlined@gmail.com> * fix(sandbox): scope /wait reload check to the approved rule Reviewer (John Myers) flagged two failure modes in the prior whole-policy fingerprint approach used by policy.local /wait: - False sleep: when the supervisor reloads between two /wait calls (the skill tells the agent to re-issue on timed_out), the new call snapshots the already-updated policy as baseline and burns the full timeout waiting for a change that never comes. - False wakeup: any unrelated reload (other agent's approval, settings change) flips the diff, but the chunk's actual rule may not be loaded yet — the agent retries and hits policy_denied for no real signal. Replace the diff with rule-coverage. New public helper openshell_policy::policy_covers_rule reuses endpoints_overlap (so it matches add_rule's merge semantics, including the fold-into-existing-key case) plus an L7 allow check on method/path (so an existing endpoint that doesn't yet contain the proposed method doesn't signal coverage). Add policy_reloaded: true|false to the /wait response on approve, with a 500ms floor on the reload-wait phase so approvals arriving near the deadline still get a fair shot at reloaded=true. Update the policy_advisor skill to branch on it: reloaded=true → retry; reloaded=false → re-issue /wait once with timeout=30, then surface to user. Don't loop tightly. Tests: - 9 new unit tests in openshell-policy pinning coverage semantics (L4-only, L7 method gap, fold-into-existing-key, empty binaries). - 4 new tokio tests in policy_local mirroring John's exact scenarios. - wait-smoke.sh asserts policy_reloaded=true on Flow A. Signed-off-by: Alexander Watson <zredlined@gmail.com> * fix(server): make GetDraftPolicy dual-auth so /wait works under OIDC policy.local calls GetDraftPolicy from inside the sandbox supervisor via the sandbox gRPC client, which authenticates with the shared x-sandbox-secret. GetDraftPolicy was listed only in the Bearer-auth scope table (config:read) and was not in SANDBOX_SECRET_METHODS or DUAL_AUTH_METHODS, so OIDC-enabled gateways rejected those calls and the /wait long-poll surfaced gateway_lookup_failed. Local/no-OIDC setups happened to work because the auth check is short-circuited. Add GetDraftPolicy to DUAL_AUTH_METHODS, matching the existing GetSandboxConfig pattern (called by both CLI reviewer surfaces with Bearer and the sandbox supervisor with x-sandbox-secret). Dual-auth short-circuits the scope check for sandbox-secret callers, so the config:read entry in authz.rs continues to gate Bearer-only flows. Mirror the openshell_get_sandbox_config_is_dual_auth assertion for GetDraftPolicy. Note: ssh_handshake_secret is server-wide, not per-sandbox, so a sandbox-secret caller can today name any sandbox in a SubmitPolicyAnalysis request — and now in a GetDraftPolicy request. The exposure is symmetric with the existing SANDBOX_SECRET_METHODS pattern. Filed as a follow-up: per-sandbox secret binding, tracked separately. Signed-off-by: Alexander Watson <zredlined@gmail.com> * fix(ci): address rebased check failures Signed-off-by: Alexander Watson <zredlined@gmail.com> --------- Signed-off-by: Alexander Watson <zredlined@gmail.com> |
||
|
|
96d909d9a1 |
feat(ci): add helm-unittest mise task and CI step (#1367)
Adds a helm:test mise task that installs the helm-unittest plugin if not present and runs chart unit tests under deploy/helm/openshell. Installs the plugin into Dockerfile.ci so CI runs do not need to download it each time. Closes #1281 Signed-off-by: Mesut Oezdil <versusfinem@gmail.com> |
||
|
|
909e9034aa |
ci(helm): add helm lint workflow and reorganize chart values under ci/ (#1223)
* test(helm): Reorganize values under /ci, update mise helm:lint task loop WIP * ci: Add a helm lint job to validate helm linting passes on changes |
||
|
|
5116cc27b7 | feat(helm): add kubernetes local-dev environment (#1158) | ||
|
|
d6c6e97679 | chore: remove navigator references from codebase (#208) | ||
|
|
984d1a6e5c | chore: rename project from NemoClaw to OpenShell (#198) | ||
|
|
90da02a7ed | chore: simplify contributing workflow and documentation (#92) |