* feat(sandbox): default to official Alpine sandbox image default_sandbox_image() now returns docker.io/library/alpine:3.22, a generic version-qualified official image, so a fresh install no longer depends on the community sandbox image catalog. All compute drivers (docker, podman, kubernetes, vm) inherit this fallback. Part of #3116. Signed-off-by: Akram Signed-off-by: Akram <akram.benaissi@gmail.com> * feat(deploy): default deployment configs to the official Alpine sandbox image Update the shared gateway default_image, Helm chart values, the standalone Kubernetes manifest, and the dev gateway task scripts to use docker.io/library/alpine:3.22 instead of the community base image, consistent with default_sandbox_image(). GPU e2e image-build base is left unchanged (CUDA needs a glibc base). Part of #3116. Signed-off-by: Akram Signed-off-by: Akram <akram.benaissi@gmail.com> * feat(driver): default to numeric non-root identity for USER-less images With the default sandbox image now Alpine, images that declare no OCI USER must start instead of being rejected. When the image declares no USER and the policy requests none, the Podman and Docker drivers now supply a numeric non-root identity (DEFAULT_SANDBOX_UID/GID = 1000) instead of rejecting, matching the numeric-identity behavior of the Kubernetes and VM drivers. The supervisor's resolved-identity path runs the sandbox as a synthesized non-root account without the account existing in the image. Images that declare a USER keep the OCI resolution path unchanged. Part of #3116. Signed-off-by: Akram <akram.benaissi@gmail.com> Signed-off-by: Evan Lezar <elezar@nvidia.com> * test(conformance): use Alpine workload image Signed-off-by: Evan Lezar <elezar@nvidia.com> * refactor(policy): drop community image /app path from default policy The restrictive default policy granted read-only access to /app, a directory that only existed in the community base image. A generic Alpine default has no /app, so remove it. Landlock best-effort already ignores absent paths; this just stops advertising a community-specific layout in the default. Part of #3116. Signed-off-by: Akram Signed-off-by: Akram <akram.benaissi@gmail.com> * docs(config): document Alpine default images Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(podman): report early sandbox termination Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(podman): initialize rootless workspace ownership Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(sandbox): qualify NVIDIA Ubuntu default Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): initialize rootful default workspace Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(sftp): add native sandbox adapter Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sftp): gate runtime helper support to Linux Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sftp): support standard OpenSSH file operations Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sftp): harden rename and special file handling Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(runtime): remove community image dependencies Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(e2e): build provider readiness tool fixture Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(e2e): use a dedicated Noble fixture for Docker tests Signed-off-by: Evan Lezar <elezar@nvidia.com> --------- Signed-off-by: Akram Signed-off-by: Akram <akram.benaissi@gmail.com> Signed-off-by: Evan Lezar <elezar@nvidia.com> Signed-off-by: Drew Newberry <anewberry@nvidia.com> Co-authored-by: Evan Lezar <elezar@nvidia.com> Co-authored-by: Drew Newberry <anewberry@nvidia.com>
Agent-Driven Policy Management Demo
Run the full agent-driven policy loop end-to-end:
- A Codex agent inside an OpenShell sandbox tries to write a markdown file to GitHub via the Contents API.
- OpenShell denies the request with a structured
policy_denied403 because the initial policy only allows read-only access toapi.github.com. - The agent reads
/etc/openshell/skills/policy_advisor.md, drafts the narrowest rule needed, and submits it tohttp://policy.local/v1/proposals. It saves the returnedchunk_id. - The gateway merges the proposed rule with the current sandbox policy, runs
the policy prover, and stores a concise
validation_resulton the pending chunk. This is deterministic control-plane evidence, not agent prose. - The agent calls
GET /v1/proposals/{chunk_id}/wait?timeout=300— a single HTTP request that the supervisor holds open until the developer decides. This is the load-bearing UX point: the agent burns zero LLM tokens while it waits; it's literally sleeping on a socket. - You approve the proposal from the host with one keystroke after seeing the
exact rule and the prover verdict in
openshell rule get. - The agent's
/waitreturns within ~1 second of the approval. The sandbox has hot-reloaded the merged policy; the agent retries the original PUT once and exits.
The whole loop usually finishes in under two minutes; most of that time is sandbox cold-start (SSH bring-up + Codex install inside the sandbox), not the policy round-trip itself.
Prerequisites
-
An active OpenShell gateway (
openshell gateway start). -
gh auth login(or aGITHUB_TOKENenv var with contents-write on a scratch repo). -
codex loginon the host. -
A scratch GitHub repository with at least one commit on the default branch. If you don't have one yet:
gh repo create "$(gh api user --jq .login)/openshell-policy-demo" \ --private --add-readme \ --description "OpenShell policy advisor demo scratch repo"
Run it
bash examples/agent-driven-policy-management/demo.sh
That's the whole thing. The demo resolves your GitHub handle from gh, picks
openshell-policy-demo as the repo, and writes one timestamped markdown file
under openshell-policy-advisor-demo/ per run.
Driving it manually (real-session UX)
DEMO_MANUAL_APPROVE=1 bash examples/agent-driven-policy-management/demo.sh
Same flow, but the script no longer auto-approves. When the agent submits a
proposal, the demo prints the exact approve and reject --reason commands
and pauses until you run one from another terminal. This is how you'd review
a coding agent's privilege ask in a real session — read the structured grant,
decide, type one command, watch the agent's /wait unblock within ~1s.
Try a rejection-with-guidance to see the full revise-and-resubmit loop:
reject with --reason "scope to docs/ paths only" and the agent reads
rejection_reason, drafts a tighter proposal, and pauses again.
Overrides (all optional)
| Env var | Default |
|---|---|
DEMO_GITHUB_OWNER |
gh api user --jq .login |
DEMO_GITHUB_REPO |
openshell-policy-demo |
DEMO_BRANCH |
main |
DEMO_RUN_ID |
timestamp |
DEMO_GITHUB_TOKEN |
falls back to GITHUB_TOKEN, GH_TOKEN, or gh auth token |
DEMO_KEEP_SANDBOX |
0 (set 1 to inspect the sandbox after the demo) |
DEMO_MANUAL_APPROVE |
0 (set 1 to pause for host-side rule approve / rule reject --reason) |
DEMO_APPROVAL_TIMEOUT_SECS |
240 (auto), 1800 (manual mode) |
DEMO_CODEX_MODEL |
gpt-5.4-mini (pinned for ChatGPT-account compatibility; override if your account supports a different model) |
DEMO_CODEX_REASONING |
low (the demo task is mechanical; medium/high slow it down without changing outcomes) |
OPENSHELL_BIN |
target/debug/openshell if present, else openshell on PATH |
What the agent sees
policy.template.yaml is the initial restrictive policy: a read-only L7 REST
rule for api.github.com plus the binary set Codex needs. The agent has to
ask for the additional PUT /repos/.../contents/... write itself — that's the
proposal you approve.
What gets approved (trust model)
Every proposal lands in the gateway as a PolicyChunk — a structured object
with three parts, each with a different trust level:
| Field | Source | Trust |
|---|---|---|
proposed_rule (host, port, method, path, binary) |
agent, schema-validated by the gateway | structured contract — this is what you're approving |
rationale (free-form prose) |
agent | hint only — a compromised agent can lie here |
validation_result (prover output) |
gateway-side prover | trust signal — but this surface is in progress (see RFC 0002) |
The MVP today shows the structured rule plus the agent's rationale in
openshell rule get and the TUI inbox panel. With prover validation wired
into the gateway, openshell rule get also shows a Validation: line for
agent-authored chunks. The value is the prover's verdict in OCSF-shorthand
style — one short, scannable string per chunk:
Validation: prover: no new findings
Validation: prover: 1 new finding
capability_expansion: PUT on api.github.com:443 via /usr/bin/curl
Other possible verdicts: validation unavailable (gateway-side prover infra
issue — surfaces in the gateway log, not as proposal failure), merge failed: … (proposal won't merge into the current policy), and policy invalid: …
(merged policy fails the structural safety check).
Read the structured rule (Endpoints + Binary). Read the Validation line.
Approve if both look right. The demo's openshell rule approve-all
auto-approves to keep the loop short; in a real session a developer makes
that judgment per chunk before pressing a.
Going further
Two LLM-less regression scripts cover adjacent slices of the same surface when you're iterating on the sandbox or gateway code:
e2e/policy-advisor/test.sh— drives the original deny-observe-approve loop end-to-end against a real GitHub repo, usingcurlfrom inside the sandbox in a retry loop until policy hot-reloads. Exercises the L7 proxy enforcement, the proposal-submit path, and the merged-policy reload.e2e/policy-advisor/wait-smoke.sh— pure wire-contract regression for theGET /v1/proposals/{id}/waitendpoint shipped here. No LLM, no GitHub, no real network traffic; just submits a synthetic proposal, blocks on/wait, and asserts the developer's approve orreject --reasontext round-trips back into the response body. Faster (~10s) and the right thing to add to when changingpolicy.localor the gateway draft-chunk persistence.