Commit Graph
1530 Commits
Author SHA1 Message Date
Gordon Sim 6588c7601a fix(auth): bound OIDC CA bundle reads
Signed-off-by: Gordon Sim <gsim@redhat.com>
2026-09-25 22:20:37 +01:00
Gordon Sim 2af0cb53d2 fix(helm): scope OIDC CA trust to the OIDC client
Signed-off-by: Gordon Sim <gsim@redhat.com>
2026-09-25 22:20:37 +01:00
Drew NewberryandPiotr Mlocek 73a181d32f docs: streamline README, add policy prover to architecture docs (#3718)
* docs(readme): streamline README and move reference detail to docs

Restructure the README as a short path from overview to quickstart to
further reading. Move prerelease install steps into the installation
guide and telemetry build flags into a new observability page. Fix
broken docs links and outdated runtime and credential descriptions.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(readme): describe 0.1.0 as adding new isolation primitives

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(architecture): add policy prover as a gateway component

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs: describe OpenShell as a runtime for fleets of autonomous AI agents

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(architecture): describe policy prover as formal verification

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(architecture): name OpenShell Sandbox in component table

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(architecture): fold isolation backend into supervisor row

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(architecture): mention formal verification in overview

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(run-agent): use the published OpenCode image in the first-agent guide

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(inference): correct provider examples and readiness

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(providers): correct Google binding and provider selection

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(readme): sharpen value prop, how it works, and explore further

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(run-agent): use example Anthropic profile and add policy advisor step

Import the example Anthropic profile, which now allows OpenCode, instead
of editing it with sed. Add a step that shows how to review and approve
mechanistic policy proposals as the agent needs more access.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(providers): add OpenRouter example for OpenCode

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(run-agent): run OpenCode against OpenRouter with a free model

Add an example OpenRouter provider profile scoped to OpenCode and switch
the first-agent guide to it, using a free Nemotron model so readers do
not need OpenRouter credits. Revert the OpenCode binary added to the
example Anthropic profile.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(readme): link first-agent guide and add agent skills section

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Co-authored-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-09-25 13:53:45 -07:00
Piotr Mlocek 6f596828ad docs(fern): publish ordered version snapshots (#3721)
* docs(fern): publish and order versioned release docs

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(fern): remove version availability badges

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(fern): keep version badges optional

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

---------

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-09-25 19:56:05 +00:00
Johnny Greco d7f921190b docs(policy): refresh policy documentation and references (#3563)
* docs(policy): correct schema and default policy guidance

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): add network recipes and update command reference

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): organize lifecycle guidance and troubleshooting

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): split policy overview into concepts and management tasks

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): reorganize network recipes as a cookbook

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): restructure schema reference by field group and protocol

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): align troubleshooting, advisor, and reference pages

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): fix first policy tutorial and security guidance

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): keep overview high level and move network rules to their own page

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): focus policy management on CLI workflows and remove command reference

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): clarify policy views and sandbox deletion in management guide

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): streamline network rule concepts and examples

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): correct request path wildcard semantics

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): rewrite policy advisor guide for clarity

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): clarify policy advisor scope, setup, and review

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): rewrite policy prover guide for clarity

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): explain the two uses of the policy prover

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): describe policy prover uses, boundaries, and coverage

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): place prover before advisor and troubleshooting last

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): remove unsupported CI guidance from prover page

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): tighten policy prover introduction

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): move policy change behavior into management guide

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): name prover check types and note expanding coverage

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): prefix prover and advisor sidebar labels

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): streamline policy schema reference

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): place default policy before schema reference

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): fold troubleshooting into policy management guide

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): correct tutorial log samples and GitHub push policy steps

The first policy tutorial said the 403 body begins with error, policy, and
rule, but the proxy serializes the body with sorted keys. Its log samples also
showed the wrong CONNECT deny reason for a sandbox without network rules, and
the L7 deny sample omitted the :443 authority, the `l7` engine, and the reason
tag that the shorthand formatter emits.

The GitHub tutorial filtered denials with `--level warn`, which hides the INFO
level OCSF policy events, and showed the retired key=value log format. Its
hand-written policy also omitted /bin from the restrictive default, so
`policy set` would reject the file for removing a filesystem path on a live
sandbox. Start from `policy get --base` and add only the network rules.

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): improve flow and terminology across policy pages

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): correct network rule matching and protocol details

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): align policy management steps with CLI behavior

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): correct policy advisor proposal and approval details

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): correct policy section, default, and schema details

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): correct prover installation and coverage limits

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): recommend tls skip for server-first protocols

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): fix stale baseline path and interpreter examples

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): move policy pages under how-it-works and fix links

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): align native TCP guidance in security best practices

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): restore policy.local and policy DNS details from main

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): state exact glob matching rules

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

---------

Signed-off-by: Johnny Greco <jogreco@nvidia.com>
2026-09-25 18:19:16 +00:00
Divesh Chowdary 585bcf0b82 (Fix) ha sandbox resilience with k8s (#3644)
* fix(server): retry internal store updates on version conflict

- Re-read and reapply internal CAS updates (expected version 0) up to 5 times on conflict
- Client-supplied versions still fail on conflict
- Add concurrent-writer test

Signed-off-by: divesh <dgude@nvidia.com>

* fix(kubernetes): keep sandboxes running while the supervisor reconnects

- Treat a running but not-Ready supervisor Pod as degraded, not unavailable
- Report degraded sandboxes as not ready without suspending them
- Suspend only when the supervisor Pod is missing, terminated, or deleting
- Add availability test

Signed-off-by: divesh <dgude@nvidia.com>

* fix(kubernetes): skip fence generation check for suspended sandboxes

- Check the fence generation only for running or bootstrapping sandboxes
- Stop re-suspending stopped sandboxes and logging a warning every reconcile

Signed-off-by: divesh <dgude@nvidia.com>

* fix(kubernetes): rank degraded supervisor above unknown dependencies

- Report a degraded supervisor as unavailable even when another dependency read is unknown
- Extract dependency aggregation and readiness mapping into pure helpers
- Add mixed degraded and unknown regression test

Signed-off-by: divesh <dgude@nvidia.com>

---------

Signed-off-by: divesh <dgude@nvidia.com>
2026-09-25 18:10:55 +00:00
Johnny Greco 854b2370b8 fix(policy): align quickstart, policy skills, and pypi profile with current behavior (#3695)
* fix(examples): add quickstart rule with policy update and show OCSF logs

The quickstart applied policy.yaml with `openshell policy set`, which replaces
the whole policy. The file omitted /bin from the restrictive default, so the
live filesystem additivity check could reject it. Add the rule with
`openshell policy update` instead, and keep policy.yaml as a complete policy
for `sandbox create --policy` that covers the default read-only paths.

The demo and README filtered logs with `--level warn`, but the server ranks
OCSF events as INFO, which hid the policy decisions the demo shows. Query
`--source sandbox` without a level filter and match the OCSF shorthand
(DENIED/ALLOWED) instead of the retired key=value format.

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(skills): align generate-sandbox-policy with proxy behavior

Remove the `protocol: sql` validation check and the SQL `command` matcher,
which the published policy docs no longer describe.

Correct the private IP guidance: exact user-declared hostnames may reach
private addresses without allowed_ips. Wildcard, hostless, and
advisor-proposed endpoints still need allowed_ips, and loopback, link-local,
unspecified, and cloud metadata addresses stay blocked.

Stop describing an omitted protocol as pure L4 or uninspected. The proxy
still terminates TLS, parses HTTP strictly, and enforces request authority;
it only skips method and path rules.

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(providers): drop pip script paths from pypi profile binaries

OpenShell identifies a process by /proc/<pid>/exe, so a pip script runs as
its Python interpreter and the .venv/bin/pip entries could never match. The
venv interpreters that run those scripts are already listed, so remove the
script paths and explain in the header that users must list interpreters.

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(skills): fix openshell-cli policy iteration steps

Monitoring denials with `--level warn` hides them, because the server ranks
OCSF policy events as INFO. Drop the level filter and describe the OCSF
shorthand DENIED lines instead of the retired `action: deny` format.

`policy get --full > file` produced input that `policy set` cannot parse: the
output starts with revision details before the `---` separator, and --full
adds provider-composed rules. Export `--base` and keep only the YAML after the
separator. Also stop recommending full replacement for filesystem, Landlock,
or process changes, which require recreating the sandbox, and drop the SQL
mention that generate-sandbox-policy no longer covers.

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(skills): qualify authority checks for omitted-protocol endpoints

An endpoint without `protocol` only receives authority checks on HTTP requests
the proxy parses after default TLS handling. Without an L7 route or required
middleware, other CONNECT payloads such as HTTP/2 prior knowledge can use the
raw relay, and `tls: skip` bypasses termination and parsing. Stop describing
omitted-protocol endpoints as always authority-checked.

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

---------

Signed-off-by: Johnny Greco <jogreco@nvidia.com>
2026-09-25 16:34:43 +00:00
Drew Newberry 9244868056 docs: refresh architecture and agent guides (#3705)
* docs: refresh architecture and agent guides

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs: describe updated security architecture neutrally

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs: highlight new isolation primitives

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(sandboxes): clarify how to disconnect

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs: align architecture and guides with current navigation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(extensibility): streamline extension authentication guidance

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
v0.1.0-pre.12
2026-09-25 02:38:00 -07:00
Piotr Mlocek aead95b7ab fix(policy): propose rules for unknown DNS hosts (#3707)
* fix(policy): propose rules for unknown DNS hosts

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(policy): clarify synthetic DNS use across protocols

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(policy): harden unknown-host DNS observations

- Emit the policy_dns_ineligible denial for every unknown name and
  report observation staging failures as DNS failure events.
- Refuse unknown names during fail-closed quarantine and after the
  observation budget, now a quarter of each address family's pool.
- Pin transparent TCP to the mapping of the deciding policy generation
  so a reload between DNS and authorization fails closed.
- Stop Docker workloads from inheriting host DNS search domains, which
  let the first expanded short name claim an observation address.
- Share mechanistic draft polling in conformance, register
  new-hostname-proposal in the installed suite, and update docs.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* test(policy): build policy DNS proxy tests on every target

The proxy tests name PolicyEndpointId, which proxy.rs imported only on
Linux, so the macOS test build failed. Import it for test builds too.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(policy): name DNS queries and mapped hosts in OCSF denials

DNS denial and failure events attached port 53 to the queried name,
which read as a connection to that host. They now carry only the name.
Transparent TCP denials for a policy DNS address show the mapped
hostname and keep the synthetic address in dst_endpoint.ip.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

---------

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-09-25 08:28:41 +00:00
Drew Newberry c93fd94a5f feat(cli): import provider profiles from HTTP URLs (#3706)
* feat(cli): import provider profiles from HTTP URLs

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(cli): initialize crypto provider for remote profiles

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-25 06:12:18 +00:00
Drew Newberry 7a50c0899f fix(cli): stream piped exec stdin beyond gRPC request limit (#3687)
* fix(cli): stream piped exec stdin across gRPC messages

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(cli): preserve small exec requests and surface stdin errors

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(cli): preserve exec stdin limit across gRPC streaming

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-25 04:45:57 +00:00
Drew Newberry 4688061882 fix(sandbox): deliver complete exec output before success (#3688)
* fix(sandbox): preserve exec output through channel close

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(exec): propagate output delivery failures before exit

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(cli): clarify exec output delivery failures

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-25 04:03:49 +00:00
Piotr Mlocek c9257c8447 fix(policy): restore policy.local and proposal conformance (#3689)
* fix(policy): restore policy.local and proposal conformance

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* test(conformance): select policy scenarios by name

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* test(conformance): reduce policy scenario timing flakes

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* test(conformance): assert proposals target Bash

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

---------

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-09-24 21:03:26 -07:00
Piotr Mlocek 5448fb4f98 fix(auth): skip renewal for non-expiring sandbox JWTs (#3686)
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-09-25 03:40:09 +00:00
Polite_realismandDrew Newberry d376c90755 test(podman): close rootful userns, resource-limit, and daemon-failure CI gaps (#3690)
* test(podman): run driver-podman userns suite against rootful Podman too

The driver-specific-integration job only ran the driver-podman testsuite
(default/auto/keep-id/private userns reference checks) against
fedora-podman-rootless, leaving rootful behavior for this scenario
unverified even though the compute driver auto-detects and explicitly
supports rootful Podman.

The default-userns-baseline and userns-profile playbooks hard-asserted a
rootless tmachine gateway user, so pointing them at a rootful environment
would have failed that assertion immediately rather than exercising
anything. They now detect rootful vs. rootless via the existing
tmachine_container_runtime role and branch the reference-capture user
accordingly, while keeping the captured reference file itself owned by
tmachine, since the archived test binary that reads it back always runs
unprivileged as tmachine regardless of daemon mode.

Signed-off-by: politerealism <burdcat17@gmail.com>

* test(podman): add real-daemon coverage for resource limits and daemon failure

Neither the Podman driver's resource-limit enforcement nor its behavior
when the Podman daemon is unreachable had any test coverage against a
real daemon; both were only exercised through unit tests against a
mocked Podman client.

podman_resource_limits.rs creates a sandbox with --cpu/--memory flags and
reads /sys/fs/cgroup/memory.max and cpu.max from inside the sandbox
itself, verifying the limit is actually enforced rather than just echoed
back by the template API. Expected values are cross-checked against the
driver's own parse_cpu_to_microseconds/parse_memory_to_bytes and against
a real local `podman run --cpus/--memory` container.

podman_preflight.rs spawns the standalone openshell-driver-podman binary
against a guaranteed-nonexistent Podman socket and asserts it exits
non-zero within its bounded retry window with an actionable error naming
the socket path, rather than hanging or failing silently.

Signed-off-by: politerealism <burdcat17@gmail.com>

* test(podman): make rootful userns and cgroup checks pass

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(podman): match lifecycle containers by isolation role label

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(podman): accept non-expiring bootstrap tokens

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: politerealism <burdcat17@gmail.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
2026-09-24 20:16:29 -07:00
Drew Newberry 1374672967 fix(helm): restore Kubernetes e2e chart rendering (#3692)
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-24 18:55:45 -07:00
Jim Meyer 48725c5fcc ci(release): move CodeQL, Trivy, and Zizmor to advisory (#3693)
* ci(release): Move CodeQL, Trivy, and Zizmor to advisory

* docs(ci): describe advisory static findings for tagged releases

Signed-off-by: Jim Meyer <jimeyer@nvidia.com>

---------

Signed-off-by: Jim Meyer <jimeyer@nvidia.com>
2026-09-25 01:26:59 +00:00
Drew Newberry 08548713c5 fix(install): honor pinned releases and speed up prerelease discovery (#3681)
* fix(install): use native packages for prereleases and speed up discovery

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(install): honor pinned versions and guard Snap migration

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(install): omit Snap migration guard

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-24 17:57:42 -07:00
Piotr Mlocek f213e9a49b docs(middleware): reorganize and expand middleware guides (#3636)
* docs(extensibility): reorganize extensibility and middleware guides

Add an extensibility overview that introduces the extension protocol and
links every extension point, and move extension caller authentication to
a shared page used by middleware and gateway interceptors.

Split the supervisor middleware guide into an overview, a Configure and
Operate guide, and a Middleware Operations reference for service authors.
The overview explains when to use middleware, shows where it runs, lists
current limitations, and defines the service contract. Configure and
Operate covers policy attachment, service registration, failure behavior,
and observability. Middleware Operations describes HTTP request, HTTP
response, and WebSocket message operations with shared inputs and results,
per-operation diagrams, and detail accordions.

Pin page slugs so links resolve, redirect the replaced dev middleware
URL, and update reference and architecture links.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(middleware): clarify navigation and service contracts

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(middleware): drop obsolete dev URL redirect

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

---------

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-09-24 23:43:31 +00:00
Jim Meyer a00ea31c66 ci: restrict copy-pr-bot manual vetters (#3678)
Signed-off-by: Jim Meyer <jimeyer@nvidia.com>
2026-09-24 22:42:51 +00:00
Drew Newberry 52cb8ecee7 fix(kubernetes): remove NetworkPolicy acknowledgement (#3677)
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-24 22:41:45 +00:00
krishicks 4cd1351e18 fix(test): use portable file mode checks in snap installer tests (#3670)
The snap installer tests added in #3656 checked the gateway config mode
with GNU stat -c, which BSD stat on macOS rejects, so mise run ci failed
locally on macOS. Check the mode with find -perm instead, which matches
the exact mode on both GNU and BSD systems.

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-09-24 21:17:12 +00:00
Oliver Calder 6c864ec9ab fix(install): avoid installing incompatible docker snap (#3666)
* fix(install): avoid installing incompatible docker snap

The work to land RFC-0012 added new restrictions when interacting with
Docker by setting `NoNewPrivs`. This prevents the `docker` snap from
transitioning its AppArmor profile from `snap.docker.dockerd` to
`docker-default` when it tries to launch a container. Thus, the `docker`
snap is currently incompatible with OpenShell.

This commit prevents `install.sh` from installing the `docker` snap
before installing the `openshell` snap, and instead requires the user to
install a non-snap Docker daemon before proceeding with installing the
snap. Systems without the `snap` command are unaffected, since they
install native packages without checking for the presence of Docker or
other compute providers.

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* fix(snap): only snapd 2.76 for openshell snap since store installs work

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* fixup! fix(install): avoid installing incompatible docker snap

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* fixup! fix(snap): only snapd 2.76 for openshell snap since store installs work

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

---------

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>
2026-09-24 19:46:00 +00:00
Jason T. GreeneandMrunal Patel 7691f88e0e perf(server): enable WAL for the SQLite store; relax sync only for SSH session issuance (#3543)
* perf(server): enable WAL and NORMAL sync for the SQLite store

On-disk SQLite stores ran with sqlx defaults: rollback journal
(`journal_mode=delete`) and `synchronous=FULL`. Every autocommit write paid
several fsyncs and blocked readers while it held the lock, so gateway hot
paths made of many small writes serialized on disk latency. The clearest
case is `openshell forward service`, which mints and revokes an SSH session
token around every forwarded TCP connection: two commits per connection,
tens of milliseconds each on a virtual disk, wall clock linear in the
number of concurrent connections, and enough queueing that bursts hit the
per-sandbox connection cap and get refused.

Switch on-disk databases to WAL with `synchronous=NORMAL`. The mode change
runs once on a single connection before the pool opens: entering WAL needs
exclusive access to the file, so doing it up front means pool connections
only ever re-apply the pragma to a file already in WAL mode, and a failure
surfaces as one clear connect error. The first start after upgrading an
existing database therefore needs the file to be otherwise unopened.
`synchronous` is applied through the connect options on every pooled
connection. In-memory databases keep their defaults. A crash can now roll
back the most recent transactions without corrupting the database, which
is the standard WAL trade-off and fits the single-node scope of the SQLite
backend.

Tests cover a fresh store, an existing rollback-journal file that must be
switched on connect, sidecar permissions, and concurrent readers under a
burst of insert-then-update writes. Architecture, configuration and Helm
docs describe the durability trade-off, the sidecar files, and the local
filesystem requirement.

Signed-off-by: Jason T. Greene <jason.greene@redhat.com>

* fix(server): keep SQLite commits durable except SSH session issuance

WAL with synchronous=NORMAL can roll back acknowledged commits after a
power loss or kernel crash, including SSH session revocations and other
authorization-tightening writes. Run the main pool with synchronous=FULL
so every acknowledged write is durable; in WAL mode that is a single
fsync of the WAL per commit.

Add Store::create_relaxed for inserts that are safe to lose, and use it
only for SSH session issuance: a dropped token just fails validation.
On file-backed SQLite it runs on a dedicated single-connection pool with
synchronous=NORMAL. Both pools share one WAL, so the next FULL commit
also makes earlier relaxed commits durable. Postgres treats it as an
ordinary durable MustCreate insert.

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

---------

Signed-off-by: Jason T. Greene <jason.greene@redhat.com>
Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
Co-authored-by: Mrunal Patel <mrunalp@gmail.com>
2026-09-24 18:42:43 +00:00
grs 8369bc11a5 fix(podman): restore host gateway alias mediation (#3606)
* fix(podman): restore host gateway alias mediation

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(podman-e2e-tests): enable broader test podman e2e coverage

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(tests): make test more reliable

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(podman): fix macos linting error

Signed-off-by: Gordon Sim <gsim@redhat.com>

---------

Signed-off-by: Gordon Sim <gsim@redhat.com>
2026-09-24 18:41:59 +00:00
Evan Lezar 6af5520f37 fix(vm): confine OCI layer application to the rootfs (#3550)
* refactor(vm): isolate OCI layer application

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(vm): confine OCI layer application

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(vm): require trusted bootstrap image

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(vm): validate bootstrap image in gateway

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(vm): configure bootstrap image in tracing test

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-09-24 17:22:46 +00:00
Drew Newberry 0518bd4c83 fix(sandbox): preserve local sessions across host sleep (#3573)
* fix(cli): recover sandbox connect transport

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(auth): propagate non-expiring local sandbox sessions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(cli): distinguish main exit from transport loss

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(cli): bound sandbox connect recovery

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-24 17:22:30 +00:00
Oliver Calder e60098d748 fix(snap): install openshell snap via install.sh when snap available (#3656)
* fix(snap): update stale snap docs and tests

Previously, the `openshell` snap required the `docker` snap. Now, it
works with any Docker daemon running on the system. Furthermore, the
`snap-declaration` assertion on the `openshell` snap when installed from
the Snap Store causes the `openshell` snap to always connect to the
system `:docker` slot, rather than a slot provided by the `docker` snap.

This commit updates the documentation, including the `description` field in
`snapcraft.yaml`, to ensure that all information is correct and
up-to-date.

Additionally, some tests connected the `openshell:docker` plug to the
`docker` snap's `docker:docker-daemon` slot, which is inconsistent with
how the `openshell` snap operates when installed from the store. Update
those tests to connect to the system `:docker` slot as well.

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* fix(snap): require snapd 2.76 for openshell snap

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* fix(nix): align snap gateway reproducer timeout with release-canary

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* fix(snap): require snapd 2.77 for openshell snap

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* fixup! fix(snap): require snapd 2.77 for openshell snap

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* feat(snap): install openshell snap via install.sh when snap available

Change `install.sh` to install the `openshell` snap by default when
snapd is installed on the host. This installs the snap from the
`latest/stable` channel, which should match the most up-to-date release
tag on github.

If `OPENSHELL_VERSION=dev` is set for `install.sh`, then it will install
the `openshell` snap from the `latest/edge` channel, which matches the
latest dev release available on github.

The `openshell` snap currently requires Docker in order to function. If
a Docker daemon is already installed on the system, it will be used by
the `openshell` snap. Otherwise, `install.sh` will install the `docker`
snap first, wait for the Docker daemon to be ready, and then install the
`openshell` snap.

Also, update the `release-canary.yml` to split the `ubuntu-snap` job
into `ubuntu-snap-system-docker` and `ubuntu-snap-provisions-docker`,
which test the two aforementioned scenarios. Previously, `ubuntu-snap`
manually installed a given snap artifact as built from CI, but with
these new jobs, it instead uses `install.sh` to install the published
`openshell` snap from the `latest/edge` track, thus matching the
behavior of the other release canary jobs.

Make corresponding changes to the `nix` guest reproducer, documentation,
and tests.

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* fix(snap): configure local gateway authentication

The `openshell` snap runs the gateway as a systemd system service, which
runs as root. Thus, the mTLS certs are generated by root and stored in a
root-owned directory to which non-root users do not have access. For
this reason, the snap's `openshell-gateway-wrapper` script sets
`OPENSHELL_DISABLE_TLS=true`.

This commit ensures that the `openshell` snap's gateway allows
unauthenticated local access by writing a default `gateway.toml`
configuration file during the install hook, which runs after the snap is
first installed but before services are started. The config file
contains sets `allow_unauthenticated_users = true`.

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* fix(install): configure snap gateway authentication

Recently, a new install hook was added which writes a default config
file for the `openshell` snap to allow unauthenticated local access to
the gateway. This is because the gateway service runs as root and the
mTLS certificates are not accessible to non-root users.

However, the `install.sh` script installs the `openshell` snap from the
snap store, and the published version may not yet have that new install
hook. Or, the user may already have the snap installed, in which case
the install hook does not run. In either case, we need `install.sh` to
ensure that the config file is written to set the gateway auth to
`allow_unauthenticated_users = true`, and then restart the openshell
gateway service.

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* fixup! feat(snap): install openshell snap via install.sh when snap available

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

---------

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>
2026-09-24 15:10:08 +00:00
Piotr Mlocek a8f98ec09d fix(auth): remove legacy sandbox JWT admission (#3562)
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
v0.1.0-pre.11
2026-09-23 23:57:21 +00:00
Mark CampbellandKris Hicks 679b190677 feat(testing): support independent gateway and supervisor image overrides (#3341)
* feat(testing): normalize configurable test images

Signed-off-by: Bobbins228 <mcampbel@redhat.com>
Signed-off-by: Kris Hicks <khicks@nvidia.com>

* feat(helm): add global image overrides

Signed-off-by: Bobbins228 <mcampbel@redhat.com>

* feat(helm): support image registry overrides

Signed-off-by: Bobbins228 <mcampbel@redhat.com>
Signed-off-by: Kris Hicks <khicks@nvidia.com>

* refactor(helm): simplify image configuration

Signed-off-by: Bobbins228 <mcampbel@redhat.com>

* fix(e2e): avoid reloading reused kind sandbox image

Signed-off-by: Bobbins228 <mcampbel@redhat.com>

* fix(helm): default sandbox image to nvcr.io/nvidia/base/ubuntu:24.04

Signed-off-by: Kris Hicks <khicks@nvidia.com>

---------

Signed-off-by: Bobbins228 <mcampbel@redhat.com>
Signed-off-by: Kris Hicks <khicks@nvidia.com>
Co-authored-by: Kris Hicks <khicks@nvidia.com>
2026-09-23 23:35:30 +00:00
Drew Newberry 490055b426 fix(sandbox): keep boundary connection live under stalled relays (#3642)
* fix(sandbox): keep boundary connection live under stalled relays

Stalled relay streams could exhaust the shared HTTP/2 connection window and
freeze DNS, exec, and control traffic. Size flow control so the connection
window exceeds all streams' windows plus a reserve, and cap relay streams.

Recovery no longer replaces a live connection after a single stream failure,
force-closes a dead transport before reattaching, and retries while the
sandbox still reports the old same-epoch connection as active. The sandbox
now reports attach/confirm rejections instead of closing the stream.

Closes #3396

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): warn only on sustained relay stream waits

Package managers routinely hold more than the relay budget's worth of connections for short bursts; report only waits that last longer than five seconds.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): keep aborted channel cached until reattach completes

Clearing the channel cache during the reattach retry window let concurrent requests cache and use a connection that never attached or confirmed. Keep the aborted generation installed so they fail fast and wait on the in-progress recovery.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-23 23:26:26 +00:00
Florent BENOIT bfd126868c fix(e2e): use POSIX-compatible lowercase conversion in parity runner (#3465)
Replace Bash 4+ parameter expansion (${VAR,,}) with tr-based
lowercasing so the script works on systems with older shells.

Signed-off-by: Florent Benoit <fbenoit@redhat.com>
2026-09-23 22:31:50 +00:00
krishicks 0b351c4a9b fix(helm)!: reduce gateway Secret privileges (#3616)
* fix(driver-kubernetes-secrets)!: store provider credentials in one namespace

The Kubernetes Secrets credential driver now stores every credential in its
configured namespace in all workspace modes and rejects handles that reference
any other namespace before contacting the Kubernetes API. The gateway reaches
credential Secrets through the Role in that namespace; this allows removing the
Secret rules from the ClusterRole.

- Remove the workspace_mode, gateway_id, and allow_reference_namespace driver
  settings and stop rendering them from Helm. Configurations that set them fail
  at startup. Existing credential state is not migrated.
- Add server.credentialDrivers.kubernetesSecrets.createNamespace to provision a
  dedicated credential namespace. The namespace is kept on uninstall, adopted
  by a reinstall of the same release, and left untouched when something else owns
  it.
- Update the gateway config reference, Kubernetes setup docs, 0.1.0
  upgrade guide, compute-runtime architecture, and cluster debugging skill.

Signed-off-by: Kris Hicks <khicks@nvidia.com>

* fix(helm): reduce gateway Secret permissions

Remove the gateway's Secret list permission in every workspace mode and
grant source Secret reads through a Role in the sandbox namespace.
Bootstrap Secret cleanup deletes Secrets by exact name instead of listing
them.

- Grant get on the copied client TLS and image-pull Secrets through a Role in
  the sandbox namespace. The ClusterRole keeps get and patch on those names
  for the ownership check and server-side apply into workspace namespaces.
- Delete sandbox and supervisor bootstrap Secrets by exact name, derived from
  the runtime generation recorded on the Sandbox and, on restart, the target
  generation, tolerating 404. The generation annotation is cleared only after
  cleanup succeeds, and each bootstrap Secret has a Pod owner reference, so
  garbage collection removes any generation the driver does not name.
- Drop Secret list from the ClusterRole and the shared-mode sandbox Role.
- Extend the managed e2e RBAC checks to Secret list.
- Update the Kubernetes setup and sandbox runtime docs, compute-runtime
  architecture, and cluster debugging skill.

Signed-off-by: Kris Hicks <khicks@nvidia.com>

* fix(driver-kubernetes): stage workspace Secrets per runtime generation

Every Secret the Kubernetes driver writes into a workspace namespace is
now scoped to one sandbox runtime generation, immutable, and created
with create only, so the gateway never reads, patches, or adopts an
existing Secret there. This removes the gateway's cluster-wide get and
patch on the copied client TLS and image-pull Secret names.

- In managed mode, create an immutable copy of each configured
  image-pull Secret per generation, named os-pull-<id>-<generation>-<n>
  and owned by the generation's workload and supervisor Pods. Pods and
  the restarted Sandbox template reference those names, and generation
  cleanup deletes them by name. A Secret already holding a generation
  name fails the create.
- Outside shared mode, stage the gateway client TLS material into the
  supervisor bootstrap Secret instead of copying the client TLS Secret
  into the workspace namespace.
- Remove the fixed-name TLS and image-pull copies, the target ownership
  read, and the ClusterRole get and patch rule on the copied names.
  Source reads stay in the sandbox-namespace Role.
- Update the managed e2e to expect generation image-pull Secrets and
  client TLS material in the supervisor bootstrap Secret, and to check
  that the gateway cannot read the copied names in workspace namespaces.
- Update the gateway config and compute driver references, Kubernetes
  setup and sandbox runtime docs, compute-runtime architecture, driver
  README, and cluster debugging skill.

Signed-off-by: Kris Hicks <khicks@nvidia.com>

* fix(helm)!: grant operator-mode Secret permissions through the workspace chart

The operator-mode gateway ClusterRole grants no Secret permissions.
The openshell-workspace chart Role, installed in each operator-managed
namespace, grants the gateway create and delete on Secrets for sandbox
runtime generations. Operator-managed namespaces require the workspace
chart.

- Fail the chart tests on any ClusterRole rule that includes Secrets in
  operator and shared modes.
- Install the workspace chart when the operator e2e provisions a
  namespace, and assert that the gateway has no Secret permissions in a
  namespace without it.
- Update the Kubernetes setup docs, 0.1.0 upgrade guide, compute-runtime
  architecture, and cluster debugging skill.

Signed-off-by: Kris Hicks <khicks@nvidia.com>

---------

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-09-23 21:35:22 +00:00
Jim Meyer 62df64625b fix(identity): assess leaf and ancestor executable identities (#3633)
* fix(security): pin supplied executable identity chains

Signed-off-by: Jim Meyer <jimeyer@nvidia.com>

* docs(security): document executable identity chain pinning

Signed-off-by: Jim Meyer <jimeyer@nvidia.com>

* fix(binary-identity): compile Linux ancestry hashing

Signed-off-by: Jim Meyer <jimeyer@nvidia.com>

* test(binary-identity): avoid cross-label ancestry fixture

Signed-off-by: Jim Meyer <jimeyer@nvidia.com>

* test(e2e): choose distinct denied TCP port

Fix flaky test due to sequential port assignment on MacOS

Signed-off-by: Jim Meyer <jimeyer@nvidia.com>

* fix(identity): bound executable evidence cache

Reject identity chains atomically when the supervisor cache reaches its hard limit, and preserve existing pins without eviction. Classify malformed or conflicting evidence separately from policy denials at the staged TCP boundary.

Signed-off-by: Jim Meyer <jimeyer@nvidia.com>

---------

Signed-off-by: Jim Meyer <jimeyer@nvidia.com>
2026-09-23 21:07:24 +00:00
Drew Newberry a649aa42fc docs: add 0.1.0 upgrade guide outline (#3540)
* docs: add 0.1.0 upgrade guide outline

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs: add upgrade change provenance

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs: update 0.1.0 upgrade guidance

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs: move upgrade navigation below security

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs: remove release notes page

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs: announce OpenShell 0.1.0

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs: reorganize guides and require user-owned images

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs: reorganize navigation around core concepts

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs: simplify navigation and tutorial catalog

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs: refine 0.1.0 release highlights

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs: align navigation and 0.1.0 guidance

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-23 20:26:25 +00:00
Gaizka Menendez 123d95e2ed fix(pagination): document list contract and harden SDK pagers (#3279)
* fix(pagination): document list contract and harden SDK pagers

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* fix(pagination): bound pager token history

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* fix(pagination): preflight token history limits

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* fix(pagination): paginate sandbox providers

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* fix(pagination): harden TypeScript pager

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* fix(pagination): expose provider pagers in SDKs

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* docs(pagination): describe a uniform list contract

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* fix(pagination): use stable provider cursors

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* docs(pagination): clarify mutation semantics

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* fix(cli): expose sandbox provider pagination

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* fix(proto): refresh pagination schema inventory

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* fix(proto): refresh rebased schema inventory

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>
Signed-off-by: Kris Hicks <khicks@nvidia.com>

---------

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>
Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-09-23 18:51:59 +00:00
Drew Newberry 52aac37866 fix(network): honor HTTP response connection closure (#3581)
* fix(network): honor HTTP response connection closure

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(network): fix EOF fixture socket setup

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(network): keep pipeline probe response reusable

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-23 12:01:46 -07:00
krishicks 95632406cd fix(server): make HA sandbox create and HA e2e tests reliable (#3635)
* fix(server): retry runtime identity persistence on sandbox create

With multiple gateway replicas, another replica can update a new sandbox
record between the create path's read and its compare-and-swap write of the
runtime identity. The create path made one attempt, so the conflict failed
the request and deleted the backend sandbox. Persist the identity through
the same retrying helper that start and startup recovery use, which rereads
the record and retries while the sandbox stays in the same generation and a
Provisioning or Ready phase.

Signed-off-by: Kris Hicks <khicks@nvidia.com>

* ci(e2e): run Kubernetes HA tests one at a time

The two HA tests scale and roll the shared gateway Deployment. Their
in-process lock does not apply under nextest, which runs each test in its
own process, so one test could delete a gateway pod while the other was
executing through it. Put both tests in a nextest test group limited to one
thread in the e2e-kubernetes profile. The override matches test names
because a binary() filter fails in workspaces that lack the HA test binary.

Signed-off-by: Kris Hicks <khicks@nvidia.com>

---------

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-09-23 18:45:07 +00:00
Mrunal Patel a408f5dd08 chore(kubernetes): update Agent Sandbox to v1.0.3 (#3578)
Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
2026-09-23 18:30:24 +00:00
Piotr Mlocek bed9e5eafc fix(supervisor): use better error message when sandbox connect is not available (#3572)
* fix(supervisor): explain unavailable main terminal attachments

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(supervisor): reset cursor after terminal attachment errors

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(supervisor): format read-only warnings for terminal clients

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs: drop terminal attachment documentation additions

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(supervisor): exit read-only viewers on Ctrl-C

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs: simplify read-only viewer Ctrl-C guidance

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

---------

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-09-23 17:31:13 +00:00
dependabot[bot] f3e097d6c1 chore(deps): bump anyio from 4.13.0 to 4.14.2 (#3474)
Bumps [anyio](https://github.com/agronholm/anyio) from 4.13.0 to 4.14.2.
- [Release notes](https://github.com/agronholm/anyio/releases)
- [Commits](https://github.com/agronholm/anyio/compare/4.13.0...4.14.2)

---
updated-dependencies:
- dependency-name: anyio
  dependency-version: 4.14.2
  dependency-type: indirect
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-23 17:28:36 +00:00
krishicks c02683688f ci(e2e): run the Kubernetes HA and credential-driver suites on test:e2e-kubernetes (#3626)
PRs labeled test:e2e-kubernetes now run the Kubernetes HA suite and the
Kubernetes credential-driver suite (Kubernetes Secrets and Vault). Both stay
optional and off in merge groups.

- Read the test:e2e-kubernetes label for the HA and credential-driver
  lanes in Branch E2E Checks.
- Point gator at test:e2e for Helm and Kubernetes coverage and at
  test:e2e-kubernetes for gateway high availability and credential
  driver storage.

Refs #3481

Signed-off-by: Kris Hicks <khicks@nvidia.com>
v0.1.0-pre.10
2026-09-23 15:47:35 +00:00
Jim Meyer fd49df4f41 fix(ci): use approved setup-oras revision (#3625)
Signed-off-by: Jim Meyer <jimeyer@nvidia.com>
2026-09-23 15:35:08 +00:00
Florent BENOIT 11f1fe5806 fix(vm): relocate per-sandbox Unix sockets to /tmp to fit macOS sun_path (#3544)
The VM driver bound control.sock and ssh.sock inside the per-sandbox
state directory, which defaults to
$HOME/.local/state/openshell/vm-driver/sandboxes/<uuid>/. On macOS the
sun_path field of sockaddr_un holds 104 bytes, so a normal home
directory plus the sandbox UUID already pushes the socket path past the
limit and bind fails with "path too long".

Bind both sockets under a short driver-owned namespace instead:
/tmp/os-<uid>-<128-bit-hex>/<sandbox-id>/. The worst case with a UUID
sandbox id is 101 bytes.

Security properties of the new location:

- The root is created with mkdir mode 0700 and a random 128-bit name,
  retrying on EEXIST, so a local user cannot pre-create the path to
  block the driver or plant a symlink.
- The root's directory FD is held for the driver lifetime. Per-sandbox
  leaves are created and removed with mkdirat/unlinkat relative to that
  FD, so there is no path re-resolution between check and mutation.
- Leaf creation rejects any sandbox id whose socket path would still
  exceed sun_path, returning a clear error instead of a bind failure.

The driver removes the whole root on shutdown so restarts do not
accumulate orphaned roots in /tmp. VM children are kill_on_drop, so no
socket is live at that point. Restore-on-restart goes through the
normal launch path and recreates leaves under the new root.

Signed-off-by: Florent Benoit <fbenoit@redhat.com>
2026-09-23 14:36:57 +00:00
Evan Lezar 6cb1140c66 fix(drivers): normalize label namespace (#3609)
Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-09-23 14:30:27 +00:00
Evan Lezar 907f894ebc ci(release): publish prereleases with qualification summary (#3593)
Signed-off-by: Evan Lezar <elezar@nvidia.com>
v0.1.0-pre.9
2026-09-23 13:18:30 +00:00
Florent BENOIT d3480d2a7e fix(sandbox-backend): use String for CA cert/bundle in boundary protocol (#3456)
CA certificates are PEM-encoded text. Using String instead of Vec<u8>
avoids base64 overhead in the JSON wire format, keeping large CA bundles
within the 1 MiB control frame limit without needing to raise it.

Signed-off-by: Florent Benoit <fbenoit@redhat.com>
v0.1.0-pre.8
2026-09-23 09:15:34 +00:00
Drew Newberry 069ae6bd96 fix(vm): scope GPU filesystem enrichment to assigned workloads (#3580)
* fix(vm): scope GPU filesystem enrichment to assigned workloads

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(supervisor): clarify proxy baseline enrichment

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-23 07:14:19 +00:00
Divesh Chowdary d02ebe2c4b fix(kubernetes): prevent false sandbox suspension (#3567)
Signed-off-by: divesh <dgude@nvidia.com>
2026-09-22 21:57:33 -07:00
Drew Newberry c8b20bf0a2 ci(windows): seed caches on windows branch (#3576)
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-22 21:12:05 -07:00