The gateway chart always rendered the ClusterRole and ClusterRoleBinding, so
every install and upgrade required cluster-admin even when only namespaced
objects were needed. Installers that are namespace-admin GitOps or platform
controllers could not run the release at all, and clusters where cluster-scoped
RBAC is owned by a separate team had no supported way to split the install.
Add an rbac values block so a cluster-admin can apply the cluster-scoped objects
once and a namespace-admin can install and upgrade the release without
cluster-scoped permissions:
rbac:
create: true
clusterScoped:
create: true
clusterRoleName: ""
clusterRoleBindingName: ""
rbac.clusterScoped.create gates the ClusterRole and ClusterRoleBinding, and is
independent of the workspace mode. rbac.create additionally gates the namespaced
sandbox Role and RoleBinding, which matters because Kubernetes escalation
prevention stops an installer holding only the built-in admin role from creating
a Role that grants agents.x-k8s.io verbs it does not itself hold. The certgen
hook and credential driver RBAC keep their existing flags.
Both flags default to true, so current installs are unchanged. The helpers treat
a missing rbac block as enabled so upgrades with --reuse-values do not drop RBAC,
matching the existing workspaceResources pattern. The ClusterRoleBinding roleRef
follows clusterRoleName so a separately applied ClusterRole can carry a name the
cluster-admin chooses.
Document the migration for a release that already owns the cluster-scoped
objects: Helm deletes objects that leave the manifest, so annotate them with
helm.sh/resource-policy=keep before setting the flag, otherwise the gateway
loses TokenReview until a cluster-admin re-applies them.
Signed-off-by: ansjindal <ansjindal@nvidia.com>
Closes#3756
Replace the idle 50 ms procfs scan with SIGCHLD notifications and a 30 second recovery sweep. Preserve managed-child wait ownership and cover idle and exit behavior in isolated tests.
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* docs: align page file names and nav labels with published URLs
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs: redirect moved dev pages and fix agent guide redirects
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* ci(docs): check that page URLs match their file paths
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* fix(docs): align navigation checks with repo conventions
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs: omit historical redirect aliases
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* fix(docs): preserve published overview and TypeScript URLs
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* fix(docs): redirect unversioned overview and TypeScript URLs
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
---------
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
Closes#3746
Use Homebrew atomic_write for exact legacy config migrations and keep the first-install write distinct in the formula test.
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
The proxy read the synthesized CONNECT header of a mediated open with a
full-width read into an 8192-byte buffer through a BufReader of the same
size, so the workload's request could arrive in the same read. The
CONNECT path never looked past the header, and the request was lost
until the workload timed out. Take the header out of the BufReader with
fill_buf and consume only through the terminator, so the bytes behind it
stay buffered for the relay.
Signed-off-by: Davanum Srinivas <dsrinivas@nvidia.com>
* docs(readme): show Rust SDK installation command
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* docs(architecture): lead with policy enforcement
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* docs(architecture): reuse README how it works text
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* docs(architecture): refresh architecture page and diagrams
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* docs(extensibility): simplify extension points diagram and overview
Redraw the extension points diagram as two left-to-right rows for the
data plane and control plane, describe which layers each extension point
extends, and move authentication after Building Extensions as
Authenticating Extensions with updated cross-page links.
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
---------
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* fix(snap): require mTLS for the snap gateway
Replace the installer opt-in with an authenticated snap gateway. The wrapper
no longer forces plaintext, so the gateway serves TLS from the bundle it
already generates in $SNAP_COMMON/tls. The install hook writes a config that
enables mTLS user auth instead of unauthenticated access, and a new
post-refresh hook migrates the exact legacy default on existing installs.
install.sh waits for the gateway, detects whether it serves TLS, copies the
client bundle into the target user's snap state directory, and registers the
gateway over HTTPS. Older plaintext snap revisions still register over HTTP
with a warning. The release canary asserts mTLS auth and HTTPS registration.
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* fix(snap): pass config preflight and detect the mTLS gateway reliably
An explicit [openshell.gateway.mtls_auth] table fails config preflight,
which validates mTLS auth before the local TLS bundle supplies the client
CA. Write a default that pins the Docker driver instead; with the wrapper's
TLS bundle the gateway requires client certificates and enables mTLS user
auth automatically, as the native packages do.
The mTLS gateway rejects TLS handshakes without a client certificate, and
it still answers plaintext loopback HTTP for sandbox service routing, so the
installer could misdetect it as a legacy plaintext gateway. Probe HTTPS with
the root-owned client bundle, and treat a gateway as legacy only when a
plaintext gRPC Health call succeeds.
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* chore(snap): simplify install hook comment
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* fix(snap): migrate insecure gateway configs on refresh
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* refactor(snap): simplify mTLS detection and config migration
Detect the mTLS snap from the installed revision's post-refresh hook instead
of probing plaintext gRPC, and drop the scheme global. Remove the installer's
pre-hook config fallback, which is dead now that every channel ships the
install hook and which wrote the insecure default. Give the install hook a
single write path with a simple backup name, and shorten the manual client
certificate steps.
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* fix(snap): stop keeping a copy of replaced insecure configs
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* feat(install): make the snap an opt-in install method
Stop selecting the OpenShell snap just because the snap command exists. Linux
installs default to the Debian or RPM package; OPENSHELL_INSTALL_METHOD=snap
(or deb, rpm) selects the package explicitly. Hosts that already have the
OpenShell snap keep refreshing it rather than gaining a second gateway on the
same port. The release canary and snap repro script opt in explicitly.
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* fix(snap): let the gateway auto-detect its compute driver
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* fix(snap): restart the gateway after refresh
Published revisions use refresh-mode: endure, and snapd honors the old
revision's setting during a refresh, so the plaintext gateway kept running
with the migrated config unused until a manual restart. Restart the gateway
from the post-refresh hook so the mTLS config takes effect immediately.
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* docs(snap): drop refresh notes from the snap description
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* docs(snap): trim snap refresh notes from installation docs
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
---------
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* docs(tutorials): run Pi with OpenRouter
Switch the Pi tutorial from Anthropic to OpenRouter, add fd to the Pi image
so Pi does not try to download it from GitHub inside the sandbox, and rename
the page to Run Pi with OpenRouter with a dev redirect from the old URL.
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(tutorials): drop redirect for dev-only Pi page
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(tutorials): make the Pi tutorial easier to follow
Add prerequisites and steps, explain providers and the profile fields in
plain terms, inspect the sandbox while Pi is still running, and add clean-up
and troubleshooting sections.
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(tutorials): clarify Pi's providers and sandbox additions
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(tutorials): make the Pi tutorial more conversational
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
---------
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
The /_ws_tunnel endpoint pipes a WebSocket into the full gRPC service. On a
plaintext loopback gateway any web page could open it, since browsers do not
apply CORS to WebSocket upgrades. Mount the tunnel only when
enable_websocket_tunnel is set (config file, --enable-websocket-tunnel, or
OPENSHELL_ENABLE_WEBSOCKET_TUNNEL; server.enableWebsocketTunnel in Helm).
BREAKING CHANGE: gateways behind an authenticating edge proxy must set
enable_websocket_tunnel = true for CLI edge-tunnel connections.
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* docs(readme): streamline README and move reference detail to docs
Restructure the README as a short path from overview to quickstart to
further reading. Move prerelease install steps into the installation
guide and telemetry build flags into a new observability page. Fix
broken docs links and outdated runtime and credential descriptions.
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* docs(readme): describe 0.1.0 as adding new isolation primitives
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* docs(architecture): add policy prover as a gateway component
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* docs: describe OpenShell as a runtime for fleets of autonomous AI agents
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* docs(architecture): describe policy prover as formal verification
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* docs(architecture): name OpenShell Sandbox in component table
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* docs(architecture): fold isolation backend into supervisor row
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* docs(architecture): mention formal verification in overview
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* docs(run-agent): use the published OpenCode image in the first-agent guide
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* docs(inference): correct provider examples and readiness
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
* docs(providers): correct Google binding and provider selection
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
* docs(readme): sharpen value prop, how it works, and explore further
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* docs(run-agent): use example Anthropic profile and add policy advisor step
Import the example Anthropic profile, which now allows OpenCode, instead
of editing it with sed. Add a step that shows how to review and approve
mechanistic policy proposals as the agent needs more access.
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* feat(providers): add OpenRouter example for OpenCode
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* docs(run-agent): run OpenCode against OpenRouter with a free model
Add an example OpenRouter provider profile scoped to OpenCode and switch
the first-agent guide to it, using a free Nemotron model so readers do
not need OpenRouter credits. Revert the OpenCode binary added to the
example Anthropic profile.
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* docs(readme): link first-agent guide and add agent skills section
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
---------
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Co-authored-by: Piotr Mlocek <pmlocek@nvidia.com>
* docs(fern): publish and order versioned release docs
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
* docs(fern): remove version availability badges
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
* docs(fern): keep version badges optional
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
---------
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
* docs(policy): correct schema and default policy guidance
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): add network recipes and update command reference
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): organize lifecycle guidance and troubleshooting
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): split policy overview into concepts and management tasks
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): reorganize network recipes as a cookbook
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): restructure schema reference by field group and protocol
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): align troubleshooting, advisor, and reference pages
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): fix first policy tutorial and security guidance
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): keep overview high level and move network rules to their own page
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): focus policy management on CLI workflows and remove command reference
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): clarify policy views and sandbox deletion in management guide
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): streamline network rule concepts and examples
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): correct request path wildcard semantics
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): rewrite policy advisor guide for clarity
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): clarify policy advisor scope, setup, and review
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): rewrite policy prover guide for clarity
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): explain the two uses of the policy prover
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): describe policy prover uses, boundaries, and coverage
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): place prover before advisor and troubleshooting last
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): remove unsupported CI guidance from prover page
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): tighten policy prover introduction
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): move policy change behavior into management guide
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): name prover check types and note expanding coverage
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): prefix prover and advisor sidebar labels
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): streamline policy schema reference
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): place default policy before schema reference
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): fold troubleshooting into policy management guide
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): correct tutorial log samples and GitHub push policy steps
The first policy tutorial said the 403 body begins with error, policy, and
rule, but the proxy serializes the body with sorted keys. Its log samples also
showed the wrong CONNECT deny reason for a sandbox without network rules, and
the L7 deny sample omitted the :443 authority, the `l7` engine, and the reason
tag that the shorthand formatter emits.
The GitHub tutorial filtered denials with `--level warn`, which hides the INFO
level OCSF policy events, and showed the retired key=value log format. Its
hand-written policy also omitted /bin from the restrictive default, so
`policy set` would reject the file for removing a filesystem path on a live
sandbox. Start from `policy get --base` and add only the network rules.
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): improve flow and terminology across policy pages
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): correct network rule matching and protocol details
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): align policy management steps with CLI behavior
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): correct policy advisor proposal and approval details
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): correct policy section, default, and schema details
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): correct prover installation and coverage limits
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): recommend tls skip for server-first protocols
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): fix stale baseline path and interpreter examples
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): move policy pages under how-it-works and fix links
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): align native TCP guidance in security best practices
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): restore policy.local and policy DNS details from main
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(policy): state exact glob matching rules
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
---------
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* fix(server): retry internal store updates on version conflict
- Re-read and reapply internal CAS updates (expected version 0) up to 5 times on conflict
- Client-supplied versions still fail on conflict
- Add concurrent-writer test
Signed-off-by: divesh <dgude@nvidia.com>
* fix(kubernetes): keep sandboxes running while the supervisor reconnects
- Treat a running but not-Ready supervisor Pod as degraded, not unavailable
- Report degraded sandboxes as not ready without suspending them
- Suspend only when the supervisor Pod is missing, terminated, or deleting
- Add availability test
Signed-off-by: divesh <dgude@nvidia.com>
* fix(kubernetes): skip fence generation check for suspended sandboxes
- Check the fence generation only for running or bootstrapping sandboxes
- Stop re-suspending stopped sandboxes and logging a warning every reconcile
Signed-off-by: divesh <dgude@nvidia.com>
* fix(kubernetes): rank degraded supervisor above unknown dependencies
- Report a degraded supervisor as unavailable even when another dependency read is unknown
- Extract dependency aggregation and readiness mapping into pure helpers
- Add mixed degraded and unknown regression test
Signed-off-by: divesh <dgude@nvidia.com>
---------
Signed-off-by: divesh <dgude@nvidia.com>
* fix(examples): add quickstart rule with policy update and show OCSF logs
The quickstart applied policy.yaml with `openshell policy set`, which replaces
the whole policy. The file omitted /bin from the restrictive default, so the
live filesystem additivity check could reject it. Add the rule with
`openshell policy update` instead, and keep policy.yaml as a complete policy
for `sandbox create --policy` that covers the default read-only paths.
The demo and README filtered logs with `--level warn`, but the server ranks
OCSF events as INFO, which hid the policy decisions the demo shows. Query
`--source sandbox` without a level filter and match the OCSF shorthand
(DENIED/ALLOWED) instead of the retired key=value format.
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(skills): align generate-sandbox-policy with proxy behavior
Remove the `protocol: sql` validation check and the SQL `command` matcher,
which the published policy docs no longer describe.
Correct the private IP guidance: exact user-declared hostnames may reach
private addresses without allowed_ips. Wildcard, hostless, and
advisor-proposed endpoints still need allowed_ips, and loopback, link-local,
unspecified, and cloud metadata addresses stay blocked.
Stop describing an omitted protocol as pure L4 or uninspected. The proxy
still terminates TLS, parses HTTP strictly, and enforces request authority;
it only skips method and path rules.
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* fix(providers): drop pip script paths from pypi profile binaries
OpenShell identifies a process by /proc/<pid>/exe, so a pip script runs as
its Python interpreter and the .venv/bin/pip entries could never match. The
venv interpreters that run those scripts are already listed, so remove the
script paths and explain in the header that users must list interpreters.
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(skills): fix openshell-cli policy iteration steps
Monitoring denials with `--level warn` hides them, because the server ranks
OCSF policy events as INFO. Drop the level filter and describe the OCSF
shorthand DENIED lines instead of the retired `action: deny` format.
`policy get --full > file` produced input that `policy set` cannot parse: the
output starts with revision details before the `---` separator, and --full
adds provider-composed rules. Export `--base` and keep only the YAML after the
separator. Also stop recommending full replacement for filesystem, Landlock,
or process changes, which require recreating the sandbox, and drop the SQL
mention that generate-sandbox-policy no longer covers.
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* docs(skills): qualify authority checks for omitted-protocol endpoints
An endpoint without `protocol` only receives authority checks on HTTP requests
the proxy parses after default TLS handling. Without an L7 route or required
middleware, other CONNECT payloads such as HTTP/2 prior knowledge can use the
raw relay, and `tls: skip` bypasses termination and parsing. Stop describing
omitted-protocol endpoints as always authority-checked.
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
---------
Signed-off-by: Johnny Greco <jogreco@nvidia.com>
* fix(policy): propose rules for unknown DNS hosts
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
* docs(policy): clarify synthetic DNS use across protocols
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
* fix(policy): harden unknown-host DNS observations
- Emit the policy_dns_ineligible denial for every unknown name and
report observation staging failures as DNS failure events.
- Refuse unknown names during fail-closed quarantine and after the
observation budget, now a quarter of each address family's pool.
- Pin transparent TCP to the mapping of the deciding policy generation
so a reload between DNS and authorization fails closed.
- Stop Docker workloads from inheriting host DNS search domains, which
let the first expanded short name claim an observation address.
- Share mechanistic draft polling in conformance, register
new-hostname-proposal in the installed suite, and update docs.
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
* test(policy): build policy DNS proxy tests on every target
The proxy tests name PolicyEndpointId, which proxy.rs imported only on
Linux, so the macOS test build failed. Import it for test builds too.
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
* fix(policy): name DNS queries and mapped hosts in OCSF denials
DNS denial and failure events attached port 53 to the queried name,
which read as a connection to that host. They now carry only the name.
Transparent TCP denials for a policy DNS address show the mapped
hostname and keep the synthetic address in dst_endpoint.ip.
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
---------
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
* test(podman): run driver-podman userns suite against rootful Podman too
The driver-specific-integration job only ran the driver-podman testsuite
(default/auto/keep-id/private userns reference checks) against
fedora-podman-rootless, leaving rootful behavior for this scenario
unverified even though the compute driver auto-detects and explicitly
supports rootful Podman.
The default-userns-baseline and userns-profile playbooks hard-asserted a
rootless tmachine gateway user, so pointing them at a rootful environment
would have failed that assertion immediately rather than exercising
anything. They now detect rootful vs. rootless via the existing
tmachine_container_runtime role and branch the reference-capture user
accordingly, while keeping the captured reference file itself owned by
tmachine, since the archived test binary that reads it back always runs
unprivileged as tmachine regardless of daemon mode.
Signed-off-by: politerealism <burdcat17@gmail.com>
* test(podman): add real-daemon coverage for resource limits and daemon failure
Neither the Podman driver's resource-limit enforcement nor its behavior
when the Podman daemon is unreachable had any test coverage against a
real daemon; both were only exercised through unit tests against a
mocked Podman client.
podman_resource_limits.rs creates a sandbox with --cpu/--memory flags and
reads /sys/fs/cgroup/memory.max and cpu.max from inside the sandbox
itself, verifying the limit is actually enforced rather than just echoed
back by the template API. Expected values are cross-checked against the
driver's own parse_cpu_to_microseconds/parse_memory_to_bytes and against
a real local `podman run --cpus/--memory` container.
podman_preflight.rs spawns the standalone openshell-driver-podman binary
against a guaranteed-nonexistent Podman socket and asserts it exits
non-zero within its bounded retry window with an actionable error naming
the socket path, rather than hanging or failing silently.
Signed-off-by: politerealism <burdcat17@gmail.com>
* test(podman): make rootful userns and cgroup checks pass
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* test(podman): match lifecycle containers by isolation role label
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* test(podman): accept non-expiring bootstrap tokens
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
---------
Signed-off-by: politerealism <burdcat17@gmail.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
* docs(extensibility): reorganize extensibility and middleware guides
Add an extensibility overview that introduces the extension protocol and
links every extension point, and move extension caller authentication to
a shared page used by middleware and gateway interceptors.
Split the supervisor middleware guide into an overview, a Configure and
Operate guide, and a Middleware Operations reference for service authors.
The overview explains when to use middleware, shows where it runs, lists
current limitations, and defines the service contract. Configure and
Operate covers policy attachment, service registration, failure behavior,
and observability. Middleware Operations describes HTTP request, HTTP
response, and WebSocket message operations with shared inputs and results,
per-operation diagrams, and detail accordions.
Pin page slugs so links resolve, redirect the replaced dev middleware
URL, and update reference and architecture links.
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
* docs(middleware): clarify navigation and service contracts
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
* docs(middleware): drop obsolete dev URL redirect
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
---------
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
The snap installer tests added in #3656 checked the gateway config mode
with GNU stat -c, which BSD stat on macOS rejects, so mise run ci failed
locally on macOS. Check the mode with find -perm instead, which matches
the exact mode on both GNU and BSD systems.
Signed-off-by: Kris Hicks <khicks@nvidia.com>
* fix(install): avoid installing incompatible docker snap
The work to land RFC-0012 added new restrictions when interacting with
Docker by setting `NoNewPrivs`. This prevents the `docker` snap from
transitioning its AppArmor profile from `snap.docker.dockerd` to
`docker-default` when it tries to launch a container. Thus, the `docker`
snap is currently incompatible with OpenShell.
This commit prevents `install.sh` from installing the `docker` snap
before installing the `openshell` snap, and instead requires the user to
install a non-snap Docker daemon before proceeding with installing the
snap. Systems without the `snap` command are unaffected, since they
install native packages without checking for the presence of Docker or
other compute providers.
Signed-off-by: Oliver Calder <oliver.calder@canonical.com>
* fix(snap): only snapd 2.76 for openshell snap since store installs work
Signed-off-by: Oliver Calder <oliver.calder@canonical.com>
* fixup! fix(install): avoid installing incompatible docker snap
Signed-off-by: Oliver Calder <oliver.calder@canonical.com>
* fixup! fix(snap): only snapd 2.76 for openshell snap since store installs work
Signed-off-by: Oliver Calder <oliver.calder@canonical.com>
---------
Signed-off-by: Oliver Calder <oliver.calder@canonical.com>
* perf(server): enable WAL and NORMAL sync for the SQLite store
On-disk SQLite stores ran with sqlx defaults: rollback journal
(`journal_mode=delete`) and `synchronous=FULL`. Every autocommit write paid
several fsyncs and blocked readers while it held the lock, so gateway hot
paths made of many small writes serialized on disk latency. The clearest
case is `openshell forward service`, which mints and revokes an SSH session
token around every forwarded TCP connection: two commits per connection,
tens of milliseconds each on a virtual disk, wall clock linear in the
number of concurrent connections, and enough queueing that bursts hit the
per-sandbox connection cap and get refused.
Switch on-disk databases to WAL with `synchronous=NORMAL`. The mode change
runs once on a single connection before the pool opens: entering WAL needs
exclusive access to the file, so doing it up front means pool connections
only ever re-apply the pragma to a file already in WAL mode, and a failure
surfaces as one clear connect error. The first start after upgrading an
existing database therefore needs the file to be otherwise unopened.
`synchronous` is applied through the connect options on every pooled
connection. In-memory databases keep their defaults. A crash can now roll
back the most recent transactions without corrupting the database, which
is the standard WAL trade-off and fits the single-node scope of the SQLite
backend.
Tests cover a fresh store, an existing rollback-journal file that must be
switched on connect, sidecar permissions, and concurrent readers under a
burst of insert-then-update writes. Architecture, configuration and Helm
docs describe the durability trade-off, the sidecar files, and the local
filesystem requirement.
Signed-off-by: Jason T. Greene <jason.greene@redhat.com>
* fix(server): keep SQLite commits durable except SSH session issuance
WAL with synchronous=NORMAL can roll back acknowledged commits after a
power loss or kernel crash, including SSH session revocations and other
authorization-tightening writes. Run the main pool with synchronous=FULL
so every acknowledged write is durable; in WAL mode that is a single
fsync of the WAL per commit.
Add Store::create_relaxed for inserts that are safe to lose, and use it
only for SSH session issuance: a dropped token just fails validation.
On file-backed SQLite it runs on a dedicated single-connection pool with
synchronous=NORMAL. Both pools share one WAL, so the next FULL commit
also makes earlier relaxed commits durable. Postgres treats it as an
ordinary durable MustCreate insert.
Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
---------
Signed-off-by: Jason T. Greene <jason.greene@redhat.com>
Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
Co-authored-by: Mrunal Patel <mrunalp@gmail.com>
* fix(podman): restore host gateway alias mediation
Signed-off-by: Gordon Sim <gsim@redhat.com>
* fix(podman-e2e-tests): enable broader test podman e2e coverage
Signed-off-by: Gordon Sim <gsim@redhat.com>
* fix(tests): make test more reliable
Signed-off-by: Gordon Sim <gsim@redhat.com>
* fix(podman): fix macos linting error
Signed-off-by: Gordon Sim <gsim@redhat.com>
---------
Signed-off-by: Gordon Sim <gsim@redhat.com>
* fix(snap): update stale snap docs and tests
Previously, the `openshell` snap required the `docker` snap. Now, it
works with any Docker daemon running on the system. Furthermore, the
`snap-declaration` assertion on the `openshell` snap when installed from
the Snap Store causes the `openshell` snap to always connect to the
system `:docker` slot, rather than a slot provided by the `docker` snap.
This commit updates the documentation, including the `description` field in
`snapcraft.yaml`, to ensure that all information is correct and
up-to-date.
Additionally, some tests connected the `openshell:docker` plug to the
`docker` snap's `docker:docker-daemon` slot, which is inconsistent with
how the `openshell` snap operates when installed from the store. Update
those tests to connect to the system `:docker` slot as well.
Signed-off-by: Oliver Calder <oliver.calder@canonical.com>
* fix(snap): require snapd 2.76 for openshell snap
Signed-off-by: Oliver Calder <oliver.calder@canonical.com>
* fix(nix): align snap gateway reproducer timeout with release-canary
Signed-off-by: Oliver Calder <oliver.calder@canonical.com>
* fix(snap): require snapd 2.77 for openshell snap
Signed-off-by: Oliver Calder <oliver.calder@canonical.com>
* fixup! fix(snap): require snapd 2.77 for openshell snap
Signed-off-by: Oliver Calder <oliver.calder@canonical.com>
* feat(snap): install openshell snap via install.sh when snap available
Change `install.sh` to install the `openshell` snap by default when
snapd is installed on the host. This installs the snap from the
`latest/stable` channel, which should match the most up-to-date release
tag on github.
If `OPENSHELL_VERSION=dev` is set for `install.sh`, then it will install
the `openshell` snap from the `latest/edge` channel, which matches the
latest dev release available on github.
The `openshell` snap currently requires Docker in order to function. If
a Docker daemon is already installed on the system, it will be used by
the `openshell` snap. Otherwise, `install.sh` will install the `docker`
snap first, wait for the Docker daemon to be ready, and then install the
`openshell` snap.
Also, update the `release-canary.yml` to split the `ubuntu-snap` job
into `ubuntu-snap-system-docker` and `ubuntu-snap-provisions-docker`,
which test the two aforementioned scenarios. Previously, `ubuntu-snap`
manually installed a given snap artifact as built from CI, but with
these new jobs, it instead uses `install.sh` to install the published
`openshell` snap from the `latest/edge` track, thus matching the
behavior of the other release canary jobs.
Make corresponding changes to the `nix` guest reproducer, documentation,
and tests.
Signed-off-by: Oliver Calder <oliver.calder@canonical.com>
* fix(snap): configure local gateway authentication
The `openshell` snap runs the gateway as a systemd system service, which
runs as root. Thus, the mTLS certs are generated by root and stored in a
root-owned directory to which non-root users do not have access. For
this reason, the snap's `openshell-gateway-wrapper` script sets
`OPENSHELL_DISABLE_TLS=true`.
This commit ensures that the `openshell` snap's gateway allows
unauthenticated local access by writing a default `gateway.toml`
configuration file during the install hook, which runs after the snap is
first installed but before services are started. The config file
contains sets `allow_unauthenticated_users = true`.
Signed-off-by: Oliver Calder <oliver.calder@canonical.com>
* fix(install): configure snap gateway authentication
Recently, a new install hook was added which writes a default config
file for the `openshell` snap to allow unauthenticated local access to
the gateway. This is because the gateway service runs as root and the
mTLS certificates are not accessible to non-root users.
However, the `install.sh` script installs the `openshell` snap from the
snap store, and the published version may not yet have that new install
hook. Or, the user may already have the snap installed, in which case
the install hook does not run. In either case, we need `install.sh` to
ensure that the config file is written to set the gateway auth to
`allow_unauthenticated_users = true`, and then restart the openshell
gateway service.
Signed-off-by: Oliver Calder <oliver.calder@canonical.com>
* fixup! feat(snap): install openshell snap via install.sh when snap available
Signed-off-by: Oliver Calder <oliver.calder@canonical.com>
---------
Signed-off-by: Oliver Calder <oliver.calder@canonical.com>