21 Commits
Author SHA1 Message Date
Drew Newberry 1374672967 fix(helm): restore Kubernetes e2e chart rendering (#3692)
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-24 18:55:45 -07:00
Drew Newberry 52cb8ecee7 fix(kubernetes): remove NetworkPolicy acknowledgement (#3677)
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-24 22:41:45 +00:00
Drew Newberry c1f2e7189f feat(isolation): implement the RFC 0012 sandbox architecture (#2942)
* feat(isolation): add RFC 0012 backend contract

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>

* refactor(isolation): name the interface crate explicitly

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): expose trusted host gateway

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(agents): inventory the MXC driver

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): add mediated DNS transport

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): tighten interface error and digest contracts

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(isolation): remove unrelated driver inventory

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): define capability-free launch contract

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): seal confirmed boundary state

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): validate confirmation for external backend implementations

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): clarify mediated DNS identity

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(isolation): generalize loopback connector

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(isolation): unify typed network mediation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): bind launches to sandbox sessions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(mxc): initialize extended sandbox status

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): add boundary protocol and Linux primitives

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): harden signals and separate process status from transport

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): validate remote confirmation through public contract

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): validate wire state and propagate snapshot failures

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(isolation): import owned agent specification explicitly

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(isolation): describe mediated DNS channel

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): bound mediation attach without nested retries

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(isolation): generalize loopback protocol

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): add transport-neutral session authentication

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(isolation): separate sandbox backend protocol

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): harden runtime boundary controls

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): add terminal boundary operation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): split supervisor and sandbox runtimes

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): harden boundary isolation and lifecycle ownership

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): reject private root redirects and adopt typed errors

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): preserve accept thread ownership on musl

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(sandbox): isolate credential probes from filtered threads

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): return retained exec exit status to independent waiters

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): bound network mediation and preserve socket authorization

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): bound control admission and retire stale mediation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci(e2e): select migrated drivers per stack layer

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(sandbox): implement loopback connector

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): authenticate the Sandbox Protocol

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(supervisor): rotate launch-scoped authentication

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(sandbox): consume dedicated backend crate

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(sandbox): align topology session fixture

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): align projected bootstrap bundle

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(auth): validate refreshed credentials before rotation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): fail closed across supervisor disconnects

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): repair rebased sandbox CI

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* build(runtime): publish separate sandbox and supervisor images

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(config): configure the sandbox runtime image

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(ci): validate sandbox binary linkage

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(isolation): use backend and runtime terminology

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(sandbox): use a scratch runtime image

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(ci): refresh schema and dependency policy

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): bind reconnects to supervisor process

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs: align runtime split operational guidance

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* chore(security): document Kubernetes runtime RBAC

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): enforce runtime lifecycle invariants

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(compute): identify sandbox start generations

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(server): restore sandbox launch sessions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): support authenticated runtime replacement

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(auth): bind sandbox session successors

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(auth): retry pending sandbox successors

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(vm): run the supervisor outside the guest workload

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(vm): use sandbox backend protocol

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): repair rebase integration

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): use unified build toolchain

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(vm): use sandbox runtime terminology

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(vm): own guest network bootstrap

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): expose guest init version

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): select native supervisor artifacts

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): guard guest init Linux symbols

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): scope Linux test imports

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): avoid guest interface casts

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): reconcile admitted sandbox identity

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): share resolved sandbox identity

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): surface host supervisor failures

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): include guest logs on supervisor exit

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): rotate and clean runtime generations

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): make sandbox starts generation-aware

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): rotate restored sandbox sessions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(vm): keep shared paths in the base layer

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): bind sandbox session lineage

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(docker): isolate workloads behind the host supervisor

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(docker): rotate launch-scoped authentication

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(docker): use sandbox backend protocol

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(docker): use host networking for supervisor

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): preserve host gateway alias resolution

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(docker): use separate sandbox and supervisor images

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): restore startup validation after rebase

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(docker): name the sandbox runtime directly

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): narrow supervisor CA runtime storage

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): close companion isolation gaps

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(docker): align mediated network expectations

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(docker): exercise mediated network paths

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): attach supervisor to managed network

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): defer supervisor recovery until gateway is ready

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): make sandbox starts generation-aware

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): rotate restored sandbox sessions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): preserve workloads during session rotation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(docker): remove unrelated configuration RFC changes

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): bind sandbox session lineage

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(kubernetes): add proxy-pod isolation topology

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(kubernetes): use sandbox backend protocol

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): use stable sandbox service authority

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(kubernetes): split sandbox and supervisor images

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): adapt proxy pods to current runtime APIs

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(kubernetes): describe the single runtime placement

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(kubernetes): simplify sandbox orchestration

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): validate deployment prerequisites

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): update Trivy Helm profile inventory

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(kubernetes): update Trivy scan inventory count

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): reuse preloaded runtime images in e2e

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): type and clean runtime resources

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): make sandbox restarts recoverable

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): rotate restored sandbox sessions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): preserve supervisor egress

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): bind sandbox session lineage

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(podman): adopt isolated sandbox and supervisor containers

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): stage bootstrap archives at named volume destinations

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(podman): rotate launch-scoped authentication

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(podman): use sandbox backend protocol

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(podman): use host networking for supervisor

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(podman): split sandbox and supervisor images

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): repair rebase integration

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(podman): name the sandbox runtime directly

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): provision supervisor CA runtime storage

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): address isolation review findings

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): inspect Debian supervisor provenance

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): use libpod-compatible tmpfs options

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): bind verified sandbox runtime binary

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): provide external driver data directory

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): start sandbox before joining user namespace

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): separate supervisor user namespace

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): make sandbox starts generation-aware

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): rotate restored sandbox sessions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): bind sandbox session lineage

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* perf(isolation): add TCP and DNS benchmark harnesses

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(perf): align benchmark timing and supported protocols

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(perf): report TCP benchmark metrics accurately

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(perf): cancel failed worker startup

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): build matching local supervisor image

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): make local sandbox smoke test runnable

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): wire local sandbox runtime image

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): narrow sandbox service RBAC

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(ci): validate split runtime artifacts

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): harden runtime session handling

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(supervisor): add standalone network proxy role

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(rfc): remove implementation companion notes

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(vm): standardize runtime release name

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): pin renamed runtime artifacts

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(auth): persist sandbox runtime identity

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(runtime): restore branch validation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): reconcile main after rebase

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(network): close unframed HTTP 1.0 responses

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* chore(isolation): preserve upstream OCSF updates

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(security): close credential and TLS replay paths

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(auth): make sandbox refresh retries idempotent

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-16 00:49:14 +00:00
Dhiraj Bokde cc4ded2088 feat(helm): split gateway and workspace charts (#2643)
* feat(helm): split gateway and workspace charts

Signed-off-by: Dhiraj Bokde <dbokde@nvidia.com>

* fix(helm): preserve split chart upgrade compatibility

Keep workspace manifests valid after value validation and default legacy reused values to the combined resource topology.

* fix(ci): preserve VM runtime for E2E

The Rust cache restores target/ after VM runtime artifacts are staged,
overwriting target/vm-runtime-compressed before openshell-driver-vm is built.
Stage the compressed runtime outside target and pass that location through
OPENSHELL_VM_RUNTIME_COMPRESSED_DIR so build.rs can embed the supervisor.

Also locate the Helm split-ownership test repository root from the script
path rather than git rev-parse. The test runs in a container where the
GitHub checkout can be owned by a different UID and rejected as dubious
ownership.

Signed-off-by: Dhiraj Bokde <dbokde@nvidia.com>

* fix(ci): install yq for Helm ownership test

The split-chart ownership regression uses yq to inspect rendered YAML,
but the Helm CI container installs only tools declared in mise.
Declare and lock yq so mise install --locked provides the test dependency.

Signed-off-by: Dhiraj Bokde <dbokde@nvidia.com>

---------

Signed-off-by: Dhiraj Bokde <dbokde@nvidia.com>
2026-09-02 00:23:39 +00:00
krishicks d0dfb22baf feat(kubernetes): export driver traces over OTLP (#2958)
Mirror the VM, Podman, and Docker driver tracing setup for Kubernetes.
Export standalone driver spans through OTLP/gRPC as the distinct
openshell-driver-kubernetes service, preserve gateway trace context, record
lifecycle operations and gRPC failures, and flush spans on shutdown.

Kubernetes currently runs in-process when selected as a built-in gateway
driver. Use the temporary server-boundary shim shared with Podman and Docker
so traces retain the shape they will have when Kubernetes moves to a
separate process. Move the common ComputeDriver RPC tracing layer into
openshell-otel to keep all drivers aligned.

Propagate the active W3C context through the controller-reserved Sandbox
annotation and enable Agent Sandbox OTLP export in the local k3s workflow.
This connects asynchronous controller reconciliation spans to the originating
OpenShell create trace.

Expose gateway OTLP configuration through Helm and add an Aspire collector
to the local k3s workflow. Extend helm:k3s:forward with OTLP ingest and trace
UI forwarding for Kubernetes and local container gateway development.

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-08-26 21:22:51 +00:00
krishicks c399342649 feat(dev): unify local Kubernetes gateway workflow (#2914)
Make the local k3s gateway workflow match the Docker and Podman flows by
registering and selecting successful plaintext Skaffold deployments with the
OpenShell CLI. Derive the registration name from the worktree-specific k3d
cluster name so parallel worktrees retain independent gateway metadata.

Add helm:k3s:forward as the standard way to expose the Kubernetes gateway on
localhost:8090, and update the development and debugging guidance to use the
active registered gateway instead of one-off endpoint flags.

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-08-26 14:44:47 +00:00
Evan Lezar 3be2cd8a29 fix(helm): preflight Agent Sandbox APIs (#2867)
* fix(helm): preflight Agent Sandbox APIs

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(kubernetes): share Agent Sandbox setup

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(e2e): wait for Agent Sandbox CRD status

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* ci(canary): sparse-checkout sandbox helper

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-08-21 14:32:19 +00:00
2f96c53b8c feat(gateway,cli): windows compilation support (#2496)
* chore(windows): gate Unix-only workspace code for MSVC

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* feat(windows): stub unsupported compute drivers

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* ci(windows): add MSVC mise build lane

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* docs(windows): document MSVC build-only design

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* docs(agent): add Windows MSVC build skill

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* feat(windows): add Windows build support

Signed-off-by: Akber Raza <akberr@nvidia.com>

* refactor(windows): consolidate Windows-specific dependencies and improve build logic

Signed-off-by: Akber Raza <akberr@nvidia.com>

* feat(windows): add libclang path resolution and update cargo commands with bundled Z3 features

Signed-off-by: Akber Raza <akberr@nvidia.com>

* chore(tooling): lock Windows tool artifacts

Signed-off-by: Giedrius Burachas <gburachas@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* feat(windows): enhance libclang path resolution to support architecture-specific subdirectories

Signed-off-by: Akber Raza <akberr@nvidia.com>

* Fix Windows dependency gating after sync merge

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(z3): update Z3 header path requirements in Windows build documentation and scripts

Signed-off-by: Akber Raza <akberr@nvidia.com>

* docs(windows): relocate Windows MSVC build design to architecture/

Why: windows-msvc-build-design.mdx is a design document ("design decisions for
the native Windows MSVC build lane"), but it lived in the published, user-facing
docs/reference/ tree. Per AGENTS.md (Documentation) and architecture/README.md
("rfc/ vs architecture/"), design content belongs in architecture/ (or rfc/),
not in published reference. It also shared Fern sidebar "position: 6" with the
MXC compute-driver design page, colliding in the Reference nav ordering.

What:
- Move docs/reference/windows-msvc-build-design.mdx ->
  architecture/windows-msvc-build.md.
- Strip the Fern publish frontmatter and add a plain H1, matching the other
  architecture docs.
- Register it in the architecture doc index in architecture/README.md.
- Repoint the inbound references (build-openshell-mxc-windows skill + reference,
  implement-openshell-mxc-driver skill) to the new path.

With both design pages moved out of docs/reference/, the duplicate position-6
sidebar collision is resolved.

Signed-off-by: Akber Raza <akberr@nvidia.com>

* remove openshell-supervisor-network from unsupported driver package test exclusion list

Signed-off-by: Akber Raza <akberr@nvidia.com>

# Conflicts:
#	tasks/scripts/windows-msvc.ps1

* fix(interceptors): gate unix-only imports so the crate builds on Windows

openshell-gateway-interceptors failed to compile on Windows (E0432: no UnixStream in tokio::net), breaking any Windows build of openshell-server (which depends on it unconditionally). The connect_unix_endpoint fn was already #[cfg(unix)]-gated, but the imports it uses (UnixStream, TokioIo, Uri, service_fn) were left ungated. Gate those four imports with #[cfg(unix)] too. No behavior change on unix; Windows now compiles (no errors, no unused-import warnings).

Signed-off-by: Akber Raza <akberr@nvidia.com>

* feat(windows): add native ARM64 test support

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mise): skip Skaffold on Windows

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(windows): harden ARM64 toolchain discovery

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(windows): scope ARM64 toolchain preflight

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(windows): restore compatibility after GitHub sync

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(windows): avoid rate-limited Z3 source lookup

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mise): skip Helm checks on Windows

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(windows): support repository pre-commit checks

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(windows): stabilize native MSVC validation

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(windows): harden shared Z3 source cache

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>

* fix(windows): avoid leaking MSVC flags into clang-cl

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(windows): complete ARM64 migration audit

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(windows): restore ARM64 Ninja discovery

Signed-off-by: Akber Raza <akberr@nvidia.com>

* refactor(windows): separate platform crate roots

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(windows): restore proto include cfg gating

Signed-off-by: Akber Raza <akberr@nvidia.com>

* refactor: address lint errors

* fix(windows): add preflight check for proxy auth file path

* docs(windows): update GitHub checkout guidance

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(windows): restore CI after dependency updates

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mise): repair Windows sccache lock entry

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(windows): reconcile validation after rebase

Signed-off-by: Akber Raza <akberr@nvidia.com>

* refactor(server): exclude unsupported drivers on Windows

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* refactor(server): isolate platform driver config

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(windows): repair unsupported driver contract test

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(sandbox): remove stale dependencies

Signed-off-by: Akber Raza <akberr@nvidia.com>

* ci(windows): pin x64 workflow actions

Signed-off-by: Akber Raza <akberr@nvidia.com>

* ci(windows): align x64 Rust toolchain

Signed-off-by: Akber Raza <akberr@nvidia.com>

* ci(windows): align ARM64 workflow setup

Signed-off-by: Akber Raza <akberr@nvidia.com>

* refactor(windows): exclude unsupported runtime crates

Signed-off-by: Akber Raza <akberr@nvidia.com>

* refactor(windows): exclude unsupported crates at workspace boundary

Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com>

* refactor(server): gate builtin driver config by platform

Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com>

* fix(sandbox): restore crate documentation

Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com>

* ci(windows): make build workflow manual

Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com>

* ci(windows): temporarily enable pull request builds

Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com>

* ci(windows): cache Rust dependencies

Signed-off-by: Akber Raza <akberr@nvidia.com>

* refactor(windows): remove unnecessary platform changes

Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com>

* ci(windows): make build workflow manual

Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com>

* fix(ci): synchronize mise lockfile

Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com>

* fix(ci): normalize mise provenance metadata

Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com>

* refactor(python): isolate Windows atomic replace retry

Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com>

* fix(python): type Windows permission test errors

Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com>

---------

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>
Signed-off-by: Giedrius Burachas <gburachas@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com>
Co-authored-by: Shailendra Singh <shailendras@nvidia.com>
Co-authored-by: Giedrius Burachas <gburachas@nvidia.com>
Co-authored-by: Jamie King <jamiek@nvidia.com>
Co-authored-by: Piotr Mlocek <pmlocek@nvidia.com>
Co-authored-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com>
2026-08-11 21:00:36 +00:00
Taylor MutchandSeth Jennings 8eacb4779f feat(kubernetes): add sidecar supervisor topology (#2076)
* feat(kubernetes): add sidecar supervisor topology

Add the Kubernetes sidecar supervisor topology, its Helm/Skaffold configuration, topology documentation, and sidecar e2e matrix coverage. Skip root-only sandbox identity rewriting when process enforcement is network-only so the low-permission sidecar process container can start successfully.

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(supervisor): avoid similar process id names

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(supervisor): avoid similar process id names

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(sandbox): avoid similar proxy id names

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* docs(kubernetes): clarify sidecar topology limits

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(kubernetes): keep sidecar process leaf capless

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(kubernetes): refresh sidecar provider env snapshots

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* test(supervisor): align hot-swap identity regression

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(kubernetes): stage sidecar mtls files before proxy chown

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(kubernetes): simplify sidecar supervisor topology

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* chore(helm): reuse sidecar skaffold values

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(supervisor): avoid similar iptables helper names

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(e2e): harden kube gateway wrapper setup

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(supervisor): avoid nft batch rollback on OCP

Run nftables setup as individual commands so optional conntrack and log expressions can fail without rolling back required table, chain, and reject rules.

Signed-off-by: Seth Jennings <sjenning@redhat.com>
Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(kubernetes): preserve process identity in sidecar topology

Render sidecar pods with a shared process namespace, keep binary-aware network policy enabled, and move Kubernetes sidecar settings under the nested sidecar config table.

Also apply unprivileged Landlock/seccomp setup in NetworkOnly supervisor mode so sidecar topology keeps sandbox child hardening without privileged process setup.

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* refactor(kubernetes): replace sidecar snapshots with control socket

Coordinate sidecar policy and provider bootstrap over a local Unix socket so the process leaf no longer reads policy/provider snapshot files.

Report entrypoint startup through the control channel and keep gateway credentials confined to the network sidecar.

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* feat(kubernetes): support relaxed sidecar network identity

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(sandbox): satisfy sidecar clippy lint

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* refactor(kubernetes): standardize topology naming

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(sandbox): satisfy linux clippy timeout import

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(kubernetes): support kata sidecar on ipv4 pods

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(kubernetes): satisfy linux clippy for sidecar fallback

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* chore(kubernetes): remove stale supervisor topology references

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(kubernetes): enable sidecar binary policy inspection

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(kubernetes): harden sidecar control boundary

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(kubernetes): couple sidecar supervisor lifecycles

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

---------

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
Signed-off-by: Seth Jennings <sjenning@redhat.com>
Co-authored-by: Seth Jennings <sjenning@redhat.com>
2026-07-10 13:01:39 -07:00
Evan Lezar 70fed042cb fix(helm): build chart dependencies before lint (#1947)
Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-06-17 17:20:54 -05:00
Taylor Mutch c4ca283c1a refactor(helm): require external postgres for ha (#1844) 2026-06-09 16:39:24 -07:00
Saurabh Agarwal 5007042e79 feat(helm): add optional PostgreSQL backing store (#1579)
* feat(helm): add optional PostgreSQL backing store with Secret-based credentials

- Add postgres.enabled and postgres.deploy values to control database
  backend (SQLite vs PostgreSQL) and subchart deployment independently.
- Introduce db-secret.yaml template for Opaque Secret with assembled
  postgresql:// connection string injected via OPENSHELL_DB_URL env var.
- Add Bitnami PostgreSQL as optional subchart dependency keyed on
  postgres.deploy to prevent subchart deployment in external mode.
- Externalize JWT signing key file mode via sandboxJwt.secretDefaultMode
  with 0400 default matching upstream.
- Add validation guard for postgres.deploy=true without postgres.enabled.
- Add helm unit tests covering internal, external, URL-override, special
  character encoding, and misconfiguration error paths.
- Update README with Kubernetes and OpenShift install examples for
  bundled and external PostgreSQL configurations.
- Add helm dependency build to lint and unittest tasks.

* fix(helm): add database backend docs to README.md.gotmpl and regenerate

The helm-docs CI check failed because the Database backend section was
added directly to README.md instead of README.md.gotmpl. Move the
content to the template and regenerate so the check passes.

* fix(helm): use Secret-based DB credentials and support existingSecret

Replace the inline db-url stringData pattern with a proper Secret
containing individual fields plus a uri key.  When postgres.deploy=true
the Bitnami service-binding secret is referenced directly; when
deploy=false users can supply postgres.external.existingSecret to
bring their own Secret, or let the chart generate one from the external
field values.

Also restructures the README database section for clarity, adds
helm-unittest coverage for the new secret resolution paths, and
fixes a markdown lint issue in the root README.

* refactor(helm): move OpenShift e2e script to e2e/rust/ and add mise task

Move test-openshift-scenarios.sh from deploy/helm/openshell/ci/ to
e2e/rust/e2e-openshift.sh, matching the existing e2e script naming
convention. Register it as `e2e:openshift` in tasks/test.toml — not
wired into the `test` or `e2e` aggregates so it only runs on explicit
invocation against a live OpenShift cluster.

* feat(e2e): add database backend scenarios to Kubernetes e2e

Extend with-kube-gateway.sh with an optional multi-scenario loop gated
by OPENSHELL_E2E_KUBE_DB_SCENARIOS=1. When enabled, the script installs
the Helm chart three times — SQLite (default), bundled PostgreSQL, and
external PostgreSQL with existingSecret — running the full test suite
against each backend. When unset, existing single-install behavior is
unchanged.

Also adds helm dependency build before helm install, fixing CI failures
caused by the missing PostgreSQL subchart dependency.

* refactor(helm): simplify PostgreSQL config to two orthogonal controls

Replace postgres.deploy and postgres.external.* with two simple controls:
- postgres.enabled: deploy the bundled Bitnami PostgreSQL subchart
- server.externalDbSecret: name of a pre-existing Secret with a uri key

Delete db-secret.yaml — the chart no longer generates Secrets from
individual credential fields. Users either get the Bitnami service-binding
secret (bundled) or bring their own via server.externalDbSecret.

Add validation that postgres.serviceBindings.enabled must stay true
when using bundled PostgreSQL, preventing a confusing runtime failure.
2026-05-28 18:31:41 -07:00
Taylor Mutch a7cd1608f3 docs(helm): add chart readme generation (#1437)
Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
2026-05-18 15:26:17 -07:00
John T. Myers dbba580e84 fix(security): refresh CI and gateway image dependencies (#1432)
Refresh the CI image tool pins so Go-built tools are rebuilt with patched Go releases and move the sandbox Python runtime to 3.14.5.

Rebase the gateway runtime to a pinned distroless Debian 13 image with glibc 2.41-12+deb13u3 while preserving the existing UID/GID 1000 runtime identity for upgrade compatibility. Update rustls-webpki to 0.103.13 and clarify Linux k3d guidance now that k3d is not installed through mise on Linux.

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-05-18 14:25:08 -07:00
Alexander Watson ea2fddbe2d feat(policy): agent-driven policy management — the agent half (#1323)
* feat(policy): plumb chunk_ids and rejection_reason through proposal pipeline

Prereq plumbing for the agent revise-and-resubmit loop. Two narrow
additive proto changes unblock the upcoming /wait endpoint (#1092),
prover validation badge (#1097), and reject --guidance surfaces (#1098).

- SubmitPolicyAnalysisResponse: add accepted_chunk_ids so the in-sandbox
  agent gets handles to watch its proposals. Surfaced through the typed
  grpc_client wrapper and policy.local's POST /v1/proposals 202 body.
  Closes #1094.
- PolicyChunk + StoredDraftChunk + DraftChunkPayload: add
  validation_result (gateway prover verdict, populated by #1097) and
  rejection_reason (operator free-form text). Both plain strings; no
  enums, no parsing on the read path. Closes #1096.
- RejectDraftChunk now persists the existing reason field into the
  chunk's rejection_reason so it round-trips back to the agent via
  GetDraftPolicy. UndoDraftChunk clears it on the way back to pending
  so consumers cannot read a stale guidance string from a prior reject
  -> re-approve -> undo cycle.

Whole surface stays gated behind agent_policy_proposals_enabled. Two
focused tests cover the round-trip and the undo-clears guarantee.

Signed-off-by: Alexander Watson <zredlined@gmail.com>

* feat(sandbox): add /v1/proposals/{id} and /wait long-poll to policy.local

The agent feedback channel back from policy.local. Two new routes let
the in-sandbox agent learn its proposal's outcome on a single blocking
HTTP call — zero LLM tokens during the wait.

- GET /v1/proposals/{chunk_id} returns the chunk's current state in one
  gateway call.
- GET /v1/proposals/{chunk_id}/wait?timeout=<s> blocks until the chunk
  transitions out of pending. Default 60s, clamped [1, 300]. Agent
  re-issues on timeout to extend.

Response carries the chunk's status plus the two feedback fields shipped
in the prereq commit: rejection_reason (free-form reviewer text) and
validation_result (gateway prover verdict, empty until #1097). On
timeout: same shape with timed_out: true so the agent can disambiguate
without parsing.

Wait handler short-polls GetDraftPolicy every 1s inside the request with
a tokio::time::Instant deadline. One gateway connection is opened per
request and reused across all polls, so a 60s wait does one TLS
handshake instead of sixty. A future commit can swap the loop body for
a tokio::sync::broadcast driven by a watcher task — the agent-visible
contract (URL, query, response shape) is independent of the polling
implementation.

All routes stay behind agent_policy_proposals_enabled. Closes #1092.

Signed-off-by: Alexander Watson <zredlined@gmail.com>

* docs(sandbox): teach policy_advisor skill the wait + redraft loop

The agent-facing instructions for the feedback loop. The endpoints
exist; this is the doc that makes them usable.

policy_advisor.md gains:

- API entries for GET /v1/proposals/{chunk_id} and /wait?timeout=<s>,
  including the field semantics (status, rejection_reason,
  validation_result, timed_out).
- A note on the submit response's accepted_chunk_ids /
  rejection_reasons split so the agent handles partial acceptance.
- Step 6 saves the chunk_ids and addresses any submit-time rejections
  before waiting.
- Step 7 walks the four wait outcomes: approved (retry, with the
  honest "may still fail" caveat), rejected (read rejection_reason
  AND validation_result; address whichever has content), still-pending
  with timed_out (re-call), non-2xx (surface, do not retry).

skills.rs gains two assertions on the skill content so a future edit
cannot drop the wait endpoint or the rejection_reason directive
silently.

Closes #1095.

Signed-off-by: Alexander Watson <zredlined@gmail.com>

* test(policy-advisor): add end-to-end smoke for the agent feedback loop

A focused smoke that exercises the new policy.local /wait endpoint
on a live gateway + sandbox, separate from the existing no-LLM
regression harness (which still drives the OLD retry-with-bash-loop
recovery pattern).

Two flows:

- Flow A — approve-and-retry: agent submits, /wait blocks, host runs
  `openshell rule approve`, /wait returns status=approved. Confirms
  the happy path round-trip latency.
- Flow B — reject-with-guidance: agent submits, /wait blocks, host
  runs `openshell rule reject --reason "..."`, /wait returns
  status=rejected with the exact reviewer text in rejection_reason.
  Confirms the free-form guidance contract round-trips through the
  agent feedback channel.

No GitHub credentials needed — proposals are synthetic and never
trigger outbound traffic. Both flows expect agent_policy_proposals_enabled=true
and a running gateway.

Adds three cases to sandbox-runner.sh: submit-test-proposal (no GH
deps), proposal-status, proposal-wait. The existing put-file and
submit-proposal cases used by test.sh are untouched.

Signed-off-by: Alexander Watson <zredlined@gmail.com>

* fix(policy-advisor): surface real CLI errors from wait-smoke preflight

The preflight piped openshell's stderr to /dev/null and relied on jq to
default the missing setting key to "<unset>", but under `set -euo
pipefail` a non-zero exit from openshell makes the whole pipeline fail
and the command substitution exits the script silently before the
intended fail() message can print.

Capture stderr explicitly, check the CLI exit code, and surface the
real error plus the expected fix (port-forward + gateway add + select)
when the CLI cannot reach the gateway.

Signed-off-by: Alexander Watson <zredlined@gmail.com>

* fix(policy-advisor): pass --json to settings get in wait-smoke preflight

`openshell settings get --global` defaults to a human-readable table;
jq cannot parse it and the preflight died with a numeric-literal error.
Pass --json so jq gets actual JSON. Also touched up the suggested
recovery commands in the preflight error to match the real CLI shape
(`gateway add <endpoint> --name <name>` and the env-var override
warning).

Signed-off-by: Alexander Watson <zredlined@gmail.com>

* fix(policy): dedup draft chunks only in mechanistic mode; return effective id

The smoke harness for the agent feedback loop caught a real bug in the
gateway: SubmitPolicyAnalysis's response carried a chunk_id that was
never persisted whenever the SQL ON CONFLICT path fired. Two failure
modes, both load-bearing:

- Agent-authored proposals targeting the same host/port/binary
  (e.g. the redraft-after-rejection loop) silently folded into one row
  and any RejectDraftChunk by the new chunk_id failed with "chunk not
  found." Latent since #1151, surfaced by #1094 returning chunk_ids.
- Mechanistic mode had the same class of bug — the dedup fold-in is
  the intended behavior there, but the response still advertised the
  newly-generated UUID instead of the existing row's id. Less visible
  because no current caller reads mechanistic chunk_ids back, but the
  proto contract was violated either way.

Fix in three parts:

- put_draft_chunk now takes Option<&str> dedup_key explicitly and
  returns the effective row id (via RETURNING). None binds NULL to the
  dedup_key column, which bypasses the partial-index ON CONFLICT path
  entirely. Caller-decides semantics replace store-side magic.
- handle_submit_policy_analysis picks dedup_key per chunk using an
  allowlist (only "mechanistic" dedups) and pushes the returned
  effective_id to accepted_chunk_ids. New modes default to no-dedup so
  a misconfigured caller cannot silently lose proposals.
- The two-copy draft_chunk_dedup_key helper consolidated to one
  observation_dedup_key in policy_store.rs with a doc comment.

Tests:

- agent_authored_submits_for_same_endpoint_do_not_dedup pins the
  redraft-loop contract: two intentional submissions with the same
  host/port/binary get distinct chunk_ids, both findable via
  GetDraftPolicy, both rejectable by id.
- mechanistic_submits_for_same_endpoint_dedup_into_one_chunk locks in
  the observation-mode dedup AND asserts both submits return the same
  effective_id — would have caught the deeper bug.

Proto: SubmitPolicyAnalysisRequest.analysis_mode doc updated to
describe the actual semantics (mechanistic dedups, agent_authored and
unknown modes do not).

Signed-off-by: Alexander Watson <zredlined@gmail.com>

* docs(examples): retarget policy-management demo at the /wait endpoint

The narrated demo (examples/agent-driven-policy-management) has been the
public face of this feature since #1151. Its agent prompt told Codex to
retry the original PUT every few seconds for up to 120 seconds — a
polling workaround for the missing /wait endpoint that this branch
shipped. Update the demo to exercise /wait so the canonical reading of
the feature reflects the actual UX win.

- agent-task.md: step 4 is now "call /wait, branch on status" with the
  three outcomes spelled out (approved → retry once; rejected → read
  rejection_reason and revise or stop; pending+timed_out → re-issue
  /wait once, do NOT busy-loop or shorten the timeout). Also makes
  explicit that the demo submits one rule per proposal so
  accepted_chunk_ids[0] is the safe single id to wait on.

- demo.sh: header docstring rewritten as a six-step loop that mirrors
  the README. narrate_sandbox_workflow drops its parallel numbering and
  uses bullets (the runtime narration is the agent's sub-actions, not a
  separate decomposition of the loop). Approve step header and success
  message now reference /wait waking the agent, not "policy hot-reload
  retry."

- README.md: top-of-file flow expanded from 5 to 6 steps to include the
  /wait call and chunk_id capture; "Going further" section now describes
  both regression scripts and the boundary between them (real-GitHub
  retry vs. synthetic /wait wire test). Slow-path qualifier corrected
  from "image pull on first run" to "sandbox cold-start (SSH bring-up
  plus Codex install)".

- wait-smoke.sh header rewritten to make it unambiguous this is a
  regression, NOT a tutorial, with explicit prereq commands instead of
  prereq descriptions, and a pointer at demo.sh for the narrated story.

No code paths change; this is the readability pass.

Signed-off-by: Alexander Watson <zredlined@gmail.com>

* fix(examples): pass --yes on demo.sh's global setting writes

Global setting updates require explicit confirmation in non-interactive
mode; demo.sh's enable_agent_proposals and the cleanup restore path
were missing --yes and hard-failed the preflight. Pre-existing issue
that surfaced now that more of the demo runs through this path.

No other behavior change.

Signed-off-by: Alexander Watson <zredlined@gmail.com>

* feat(sandbox): emit OCSF audit events for policy proposal lifecycle

The demo's policy decision trace previously showed only the proxy
enforcement story (HTTP:PUT DENIED, CONFIG:LOADED, HTTP:PUT ALLOWED).
It was silent about who proposed what or who decided what — the
audit-trail receipts for the agent feedback loop were missing.

policy.local now emits sandbox-side OCSF events at the observation
moments, into the same stream as the existing CONFIG:LOADED:

- CONFIG:PROPOSED on submit_proposal acceptance. Per accepted chunk:
  the message names the chunk_id, target endpoint, L7 method/path, and
  binary so the trace correlates against the inbox card via chunk_id.
- CONFIG:APPROVED on /wait observation of approved status.
- CONFIG:REJECTED on /wait observation of rejected status. Carries the
  reviewer's free-form rejection_reason in the message AND as an
  unmapped field, both sanitized (control chars stripped, capped at
  200 chars with an ellipsis marker). The agent still reads the raw
  text via GET /v1/proposals/{id}; sanitization is audit-side only,
  per AGENTS.md's no-secrets-in-OCSF rule.

The submit path defends the audit_summaries / accepted_chunk_ids
index pairing against a future gateway change that compresses past
rejected chunks (the proto doesn't promise 1:1 ordering with the
request). Today client-side validation makes the lengths always
match; if they don't, the pairing falls back to a generic per-id
event rather than mis-attribute.

The wait handler's emit site fires once per terminal-status
observation. Multiple concurrent waiters on the same chunk would
each emit one event; acceptable for single-waiter-per-chunk demos
and the right place to dedup is the SIEM.

demo.sh's trace filter now surfaces the four CONFIG: events
alongside HTTP:PUT, so the trace at the end of every run tells the
full story from deny to allow via propose -> approve.

wait-smoke.sh's prereq notes recommend redirecting kubectl
port-forward output so its "Handling connection for 8090" lines
don't bleed into demo narration.

Three new unit tests on the sandbox-side helpers — summary builder
happy path, fallback, and the rejection_reason sanitizer.

Signed-off-by: Alexander Watson <zredlined@gmail.com>

* feat(policy): /wait awaits local policy reload; demo auto-approves redrafts

Three things in one commit, all surfaced by running the demo end-to-end
against a real gateway and finding the agent had to draft a broader
second proposal.

1. /wait race fix. Previously /wait returned `approved` the moment it
   observed the gateway's chunk status flip, but the local supervisor
   reloads policy on its own poll cycle (~10s in practice). The agent's
   retry would race the reload and hit the still-old policy, getting
   denied. Codex then drafted a broader rule and re-submitted — sound
   agent behavior, but not what /wait should provoke. Now /wait captures
   the local policy version at start, and after observed-approved waits
   for the supervisor to load a strictly-newer version before returning.
   Bounded by the caller's deadline; best-effort return if the deadline
   elapses without the version bumping. Two new unit tests pin the
   happy path and the deadline-clamped fallback.

2. demo.sh auto-approve loop. Replaces approve_when_pending +
   wait_for_agent with one approve_pending_until_agent_exits function
   that keeps watching for pending chunks and approving them until the
   agent process exits (or the configured timeout). Defense in depth
   against future redraft scenarios for any reason; today (post-fix #1)
   the agent should only submit one proposal per task, but we don't
   want to hang silently if it does submit more.

3. UX. Step headers now carry "[t+1.2s]" relative timestamps so reading
   the run output makes latency visible (the demo's whole point is the
   wait is cheap — surface that). A spin_wait helper renders an ASCII
   spinner during the watch loop so the demo never looks frozen on a
   TTY. Falls back to plain sleep on non-TTY contexts.

Closes the race condition diagnosed from the trace timing where the
gateway approved at t+0, sandbox observed at t+0.3s, but the supervisor
didn't load v2 until t+9.4s — well after the agent had already retried
and been denied.

Signed-off-by: Alexander Watson <zredlined@gmail.com>

* fix(sandbox): /wait detects policy reload by content, not the schema version

The previous attempt at the /wait-after-approve race fix compared
`SandboxPolicy.version` between /wait start and the policy reload —
but that field is the *schema* version (constant 1), not a revision
counter. Every comparison was `current(1) > baseline(1) == false`, so
the wait blocked until the agent's 300s timeout regardless of whether
the supervisor had actually reloaded. The demo SSH connection then
timed out around the 240s mark.

Diagnosed from a live run's OCSF trace: supervisor pulled v2 at
+8.5s after approval (CONFIG:LOADED), but the sandbox-side
CONFIG:APPROVED that my /wait emits didn't fire until +304s — exactly
at the 300s deadline.

Fix: compare the whole policy via prost's derived PartialEq. Any
field change (network_policies map being the only one that actually
mutates today) flips equality. A clone-per-200ms-tick on a few-KB
proto is cheap inside the bounded wait window.

Tests rewritten to match the new contract: the supervisor-reload
fixture now keeps `version: 1` constant and changes `network_policies`
contents, mirroring the exact failure mode from the live run.

Signed-off-by: Alexander Watson <zredlined@gmail.com>

* fix(examples): redact tokens with python literal-string replace, not sed

The sed-based redact_log in demo.sh broke when one of the auth tokens
contained a character that conflicted with sed's pattern parser
("unterminated substitute pattern" on the Codex JWT). The whole log
tail then blanks on failure, hiding the very failure context we're
trying to surface.

Switch to a python subprocess that takes the tokens via argv and does
literal str.replace. No regex, no delimiter games, no truncation.

Signed-off-by: Alexander Watson <zredlined@gmail.com>

* fix(sandbox): scope /wait reload check to the approved rule

Reviewer (John Myers) flagged two failure modes in the prior whole-policy
fingerprint approach used by policy.local /wait:

- False sleep: when the supervisor reloads between two /wait calls
  (the skill tells the agent to re-issue on timed_out), the new call
  snapshots the already-updated policy as baseline and burns the full
  timeout waiting for a change that never comes.
- False wakeup: any unrelated reload (other agent's approval, settings
  change) flips the diff, but the chunk's actual rule may not be loaded
  yet — the agent retries and hits policy_denied for no real signal.

Replace the diff with rule-coverage. New public helper
openshell_policy::policy_covers_rule reuses endpoints_overlap (so it
matches add_rule's merge semantics, including the fold-into-existing-key
case) plus an L7 allow check on method/path (so an existing endpoint
that doesn't yet contain the proposed method doesn't signal coverage).

Add policy_reloaded: true|false to the /wait response on approve, with
a 500ms floor on the reload-wait phase so approvals arriving near the
deadline still get a fair shot at reloaded=true. Update the
policy_advisor skill to branch on it: reloaded=true → retry;
reloaded=false → re-issue /wait once with timeout=30, then surface to
user. Don't loop tightly.

Tests:
- 9 new unit tests in openshell-policy pinning coverage semantics
  (L4-only, L7 method gap, fold-into-existing-key, empty binaries).
- 4 new tokio tests in policy_local mirroring John's exact scenarios.
- wait-smoke.sh asserts policy_reloaded=true on Flow A.

Signed-off-by: Alexander Watson <zredlined@gmail.com>

* fix(server): make GetDraftPolicy dual-auth so /wait works under OIDC

policy.local calls GetDraftPolicy from inside the sandbox supervisor
via the sandbox gRPC client, which authenticates with the shared
x-sandbox-secret. GetDraftPolicy was listed only in the Bearer-auth
scope table (config:read) and was not in SANDBOX_SECRET_METHODS or
DUAL_AUTH_METHODS, so OIDC-enabled gateways rejected those calls and
the /wait long-poll surfaced gateway_lookup_failed. Local/no-OIDC
setups happened to work because the auth check is short-circuited.

Add GetDraftPolicy to DUAL_AUTH_METHODS, matching the existing
GetSandboxConfig pattern (called by both CLI reviewer surfaces with
Bearer and the sandbox supervisor with x-sandbox-secret). Dual-auth
short-circuits the scope check for sandbox-secret callers, so the
config:read entry in authz.rs continues to gate Bearer-only flows.

Mirror the openshell_get_sandbox_config_is_dual_auth assertion for
GetDraftPolicy.

Note: ssh_handshake_secret is server-wide, not per-sandbox, so a
sandbox-secret caller can today name any sandbox in a SubmitPolicyAnalysis
request — and now in a GetDraftPolicy request. The exposure is
symmetric with the existing SANDBOX_SECRET_METHODS pattern. Filed as a
follow-up: per-sandbox secret binding, tracked separately.

Signed-off-by: Alexander Watson <zredlined@gmail.com>

* fix(ci): address rebased check failures

Signed-off-by: Alexander Watson <zredlined@gmail.com>

---------

Signed-off-by: Alexander Watson <zredlined@gmail.com>
2026-05-13 15:24:18 -07:00
Mesut Oezdil 96d909d9a1 feat(ci): add helm-unittest mise task and CI step (#1367)
Adds a helm:test mise task that installs the helm-unittest plugin if
not present and runs chart unit tests under deploy/helm/openshell.
Installs the plugin into Dockerfile.ci so CI runs do not need to
download it each time.

Closes #1281

Signed-off-by: Mesut Oezdil <versusfinem@gmail.com>
2026-05-13 14:06:44 -07:00
Taylor Mutch 909e9034aa ci(helm): add helm lint workflow and reorganize chart values under ci/ (#1223)
* test(helm): Reorganize values under /ci, update mise helm:lint task loop

WIP

* ci: Add a helm lint job to validate helm linting passes on changes
2026-05-07 10:36:27 -07:00
Taylor Mutch 5116cc27b7 feat(helm): add kubernetes local-dev environment (#1158) 2026-05-05 13:42:21 -07:00
Drew Newberry d6c6e97679 chore: remove navigator references from codebase (#208) 2026-03-10 14:46:11 -07:00
Drew Newberry 984d1a6e5c chore: rename project from NemoClaw to OpenShell (#198) 2026-03-10 11:49:09 -07:00
Drew Newberry 90da02a7ed chore: simplify contributing workflow and documentation (#92) 2026-03-04 13:06:32 -08:00