Refresh the CI image tool pins so Go-built tools are rebuilt with patched Go releases and move the sandbox Python runtime to 3.14.5.
Rebase the gateway runtime to a pinned distroless Debian 13 image with glibc 2.41-12+deb13u3 while preserving the existing UID/GID 1000 runtime identity for upgrade compatibility. Update rustls-webpki to 0.103.13 and clarify Linux k3d guidance now that k3d is not installed through mise on Linux.
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
Adds a helm:test mise task that installs the helm-unittest plugin if
not present and runs chart unit tests under deploy/helm/openshell.
Installs the plugin into Dockerfile.ci so CI runs do not need to
download it each time.
Closes#1281
Signed-off-by: Mesut Oezdil <versusfinem@gmail.com>
* wip
* refactor(docker): use native rust builds for split gateway/supervisor images
Drop the in-Docker BUILD_FROM_SOURCE path so both images consume only
prebuilt binaries staged natively via tasks/scripts/stage-prebuilt-binaries.sh.
This mirrors what CI does and reuses the host's cargo target cache and
sccache across rebuilds.
- Dockerfile.gateway: nvcr.io/nvidia/distroless/cc:v4.0.4 base (the 4.0.0
tag does not exist on nvcr.io; the registry uses a v prefix). GNU-linked
binary copied to /usr/local/bin.
- Dockerfile.supervisor: scratch base, static musl binary. Static linkage
lets the image stay scratch while still being executable as a Kubernetes
init container.
- skaffold.yaml: each artifact invokes tasks/scripts/docker-build-image.sh,
which stages the binary natively (cargo / cargo-zigbuild) and then builds
the image. Drops the cross-build.sh dependency from the supervisor build.
- seccomp.rs: add a local SYS_kexec_file_load constant for musl/aarch64.
libc 0.2.185 omits the symbol from its musl/aarch64 bindings, so the
supervisor's seccomp filter previously failed to compile for that target.
- architecture/build.md: describe the native-first pipeline and per-image
runtime choices.
Local validation: gateway image 101MB (was 194MB), supervisor image 21.7MB.
helm:skaffold:run deploys cleanly; the static musl supervisor binary runs
correctly in a non-glibc agent container.
* refactor(docker): tighten binary perms via --chown + 0550
Replace `COPY --chmod=755` with `COPY --chown=<user> --chmod=0550` in
the gateway and supervisor Dockerfiles. The binary is no longer
world-readable or world-executable; ownership is pinned to the runtime
user.
- Gateway uses `--chown=nvs:nvs` + `USER nvs:nvs`, matching the only
non-root user defined in `nvcr.io/nvidia/distroless/cc` (UID 1000) and
the Helm chart's `securityContext.runAsUser: 1000`, which overrides
the Dockerfile USER at runtime.
- Supervisor uses numeric `--chown=65534:65534` because the scratch base
has no `/etc/passwd` for name resolution. The supervisor image is
only consumed by the init-container copy-self path; the destination
pod's runAsUser governs execute access.
Validated by deploying to a local k3d cluster via `helm:skaffold:run`
and confirming the gateway StatefulSet reaches 1/1 Running.
* ci(gpu): repoint GPU probe image lookup at Dockerfile.gateway
The previous awk parsed `FROM <image> AS gateway` from the now-deleted
`Dockerfile.images`. The new `Dockerfile.gateway` uses an ARG with a
default (`ARG GATEWAY_BASE_IMAGE=nvcr.io/nvidia/distroless/cc:v4.0.4`)
and `FROM ${GATEWAY_BASE_IMAGE} AS gateway`, so the old script returns
nothing.
Parse the ARG default value directly so the GPU prerequisites check
keeps using the gateway base image as a `nvidia-smi` probe target.
* ci(gpu): pin GPU probe to nvcr.io/nvidia/base/ubuntu:noble
The previous probe parsed the gateway base image out of the Dockerfile,
relying on the fact that the gateway ran on `nvcr.io/nvidia/base/ubuntu`
and that NVIDIA Container Toolkit CDI injection would populate
`nvidia-smi` and the supporting libs at runtime. The new gateway base
(`nvcr.io/nvidia/distroless/cc`) lacks `ldconfig`, a populated
`/usr/bin`, and the broader filesystem layout CDI injection assumes,
so it cannot serve as a GPU probe.
Pin the probe image explicitly to the NVIDIA-managed Ubuntu base. The
probe is independent of the gateway runtime and survives future base
swaps.
* fix(e2e-gpu): pass GPU probe image via env, drop Dockerfile.images parse
The Rust e2e test `gpu_request_for_each_discovered_device_matches_plain_container`
was still parsing the deleted `Dockerfile.images` to derive its GPU probe
image, panicking with `No such file or directory` after this PR's
Dockerfile split.
Move the probe image to a single source of truth in the workflow
(`OPENSHELL_E2E_GPU_PROBE_IMAGE` env at the job level) and require the
e2e test to read it from there. No silent codebase default — the test
panics with a pointer at the workflow if the env is missing, so it
fails loudly rather than drifting from CI.
The prereq probe step in the workflow now consumes the same env, so
the probe image is declared exactly once.
* fix(docker): keep supervisor binary root-owned for rootless Podman
The Podman driver mounts the supervisor image read-only into the sandbox
container at /opt/openshell/bin and runs that container as UID 0, but
deliberately drops DAC_OVERRIDE for hardening (container.rs:419). With
--chown=65534:65534 --chmod=0550 the binary was r-xr-x--- owned by UID
65534, so the container's UID 0 fell into "other" with no read or exec
access and the supervisor crashed on start (ContainerExited code 1).
Docker and Kubernetes both retain DAC_OVERRIDE, so root could still
exec the file — which is why this regression only surfaced in the
Podman e2e job.
Drop --chown from the supervisor COPY so the binary stays root-owned.
Keep --chmod=0550: the security win was dropping world-execute, not
changing the owner. The chown bought nothing here because the container
is always UID 0 regardless of driver, but it actively broke the only
driver that drops DAC_OVERRIDE.
---------
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
* feat(prover): add native Rust policy prover with Z3 solver
Add openshell-prover crate implementing formal policy verification
using Z3 SMT solving. Answers two questions about any sandbox policy:
"Can data leave?" and "Can the agent write despite read-only intent?"
Native Rust — no Python subprocess, no PYTHONPATH, no uv dependency.
Z3 bundled via z3-sys for self-contained builds.
Replaces the Python prototype from #703.
Closes#699
Signed-off-by: Alexander Watson <zredlined@gmail.com>
* fix(prover): skip L7-write in exfil, write bypass only on read-only intent
Port two fixes from the Python branch:
- Exfil query skips endpoints where L7 is enforced and working
- Write bypass only fires on explicit read-only intent, not L4-only
Signed-off-by: Alexander Watson <zredlined@gmail.com>
* chore(prover): add missing SPDX license headers to registry and testdata YAML files
* fix(prover): revert serde_yaml to serde_yml to match workspace dependency on main
* fix(prover): apply cargo fmt formatting to prover and cli source files
* ci: add clang and libclang-dev to CI image for z3-sys bindgen
z3-sys requires libclang at build time for bindgen to generate FFI
bindings. Without it, the Rust CI jobs fail on the prover crate.
* fix(prover): use system libz3 instead of compiling from source
Add libz3-dev to the CI image and drop the z3 `bundled` feature from
the workspace dependency. This eliminates the ~30 min z3 C++ build
that ran on every CI cache miss.
sccache only wraps rustc — it does not intercept the cmake C++ build
that z3-sys runs when `bundled` is enabled. The cargo target cache
helped on warm runs but evicts on toolchain bumps and fork PRs.
With the system library pre-installed in the CI image, z3 link time
is always zero.
A `bundled-z3` opt-in feature is added to the prover crate for local
development without system z3:
cargo build -p openshell-prover --features bundled-z3
Regular local dev: brew install z3 (macOS) or apt install libz3-dev
(Linux), then cargo build just works.
Signed-off-by: Alexander Watson <zredlined@gmail.com>
Made-with: Cursor
* fix(prover): use u16 for ports, align include_workdir default with runtime
Port fields changed from u32 to u16 across all prover types (policy,
model, finding, queries). Prevents the prover from silently accepting
port values >65535 that the runtime rejects, which would produce
misleading PASS results on invalid policies.
Change include_workdir serde default from true to false to match the
runtime (openshell-policy uses #[serde(default)] which gives false).
The previous mismatch caused the prover to model a /sandbox path that
does not exist at runtime, producing false analysis.
Signed-off-by: Alexander Watson <zredlined@gmail.com>
Made-with: Cursor
* refactor(prover): embed registry at compile time, replace hand-rolled glob
Embed binary and API capability registry YAML files into the binary at
compile time using include_dir!. The previous approach used
env!("CARGO_MANIFEST_DIR") which bakes in the build machine's source
path — works in tests but breaks for installed binaries. The CWD
fallback was equally fragile.
The --registry CLI flag still works as a filesystem override for custom
registries. Credentials remain filesystem-loaded (user-supplied data).
Replace the hand-rolled glob matching (~60 lines) with the glob crate's
Pattern::matches(), which is already a transitive dependency.
Remove unused dependencies: openshell-policy, openshell-core, thiserror,
tracing. The prover parses policy YAML directly into its own types and
does not use the policy or core crate APIs.
Signed-off-by: Alexander Watson <zredlined@gmail.com>
Made-with: Cursor
* docs: add Z3 system library to prerequisites
The prover crate now links against system libz3 instead of compiling
from source. Document the install steps for macOS, Ubuntu, and Fedora,
and note the bundled-z3 feature flag as a fallback.
Signed-off-by: Alexander Watson <zredlined@gmail.com>
Made-with: Cursor
* fix(prover): apply cargo fmt formatting
Made-with: Cursor
---------
Signed-off-by: Alexander Watson <zredlined@gmail.com>
Migrate all production container runtime stages from Debian/Ubuntu
upstream to NVIDIA's hardened Ubuntu base image for supply chain
consistency.
- Dockerfile.ci: ubuntu:24.04 -> nvidia base
- Dockerfile.server: debian:bookworm-slim -> nvidia base (runtime stage)
- Dockerfile.base (sandbox): python:3.12-slim-bookworm -> nvidia base
with system Python 3.12 from Ubuntu Noble repos
- Dockerfile.cluster: convert from single-stage rancher/k3s to
multistage build extracting k3s artifacts onto nvidia base
Closes#244
## Summary
- Add `Provider` entity for managing 3p deps from a sandbox
- Add provider CRUD API/server persistence and new CLI workflows (`nav provider create/get/list/update/delete`), including `--from-existing` laptop discovery.
- Integrate providers into sandbox create flow: infer from command (`-- claude`), support repeatable `--provider <type>`, prompt before auto-create, and allow manual in-sandbox setup.
- Add a dedicated `navigator-providers` crate with per-provider modules and mockable discovery test helpers.
## Key UX Changes
- `nav sandbox create --provider gitlab -- claude`
- Missing provider prompt now asks before creating from local state.
- `nav provider list --names` for scripting/cleanup.
## Test Plan
- `mise run cluster:deploy`
- `mise run test:e2e:sandbox`
- `mise run pre-commit`
Closes#19Closes#22Closes#11