Commit Graph
263 Commits
Author SHA1 Message Date
Simon Scatton e21b7fd8cf chore(build): remove bundled Z3 support (#3275)
* chore(build): remove bundled Z3 support

Signed-off-by: Simon Scatton <sscatton@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(build): preserve vendored Z3 for local gateway artifacts

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

---------

Signed-off-by: Simon Scatton <sscatton@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-10-01 12:08:33 +00:00
Drew Newberry 021400be8a refactor(auth): separate sandbox identity from TLS (#3110)
* refactor(auth): separate sandbox identity from TLS

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(auth): clarify gateway mTLS behavior

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(auth): include workspace scope in TLS authorization checks

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(e2e): bound service auth sandbox names for large PIDs

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-10-01 04:33:25 +00:00
krishicks 33a8eac196 feat(helm): configure gateway OCSF JSONL output (#3876)
Previously, Helm installations could not enable the gateway OCSF JSONL
destination through chart values because generated `gateway.toml` omitted the
`openshell.gateway.ocsf_log` table.

Now, setting `server.ocsfLog.enabled` renders the path, optional schema
version, rotation, retention, and queue limits into gateway configuration.
Output is disabled by default. The default path, `/tmp/gateway-ocsf.jsonl`,
is writable in the gateway container with either the StatefulSet or
Deployment workload, so enabling output does not require persistent storage.
Invalid schema versions, rotation values, non-positive limits, or an empty
path while enabled fail chart rendering.

Additionally, `server.extraVolumes` and `server.extraVolumeMounts` add
operator-supplied volumes to the gateway pod, so operators who want records
to survive restarts can place the OCSF path on persistent storage without
replacing chart-generated configuration.

The gateway pod's default termination grace period rises from 5 to 30
seconds. Gateway shutdown can spend up to 10 seconds on supervisor session
cleanup before allowing 5 seconds to drain queued OCSF records, so the
5-second default risked a SIGKILL before the final records were written.
The grace period is only an upper bound: the gateway exits as soon as its
shutdown completes.

Refs #2762

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-09-29 19:23:43 +00:00
ansjindal d009f30121 feat(helm): make cluster-scoped RBAC optional (#3459)
The gateway chart always rendered the ClusterRole and ClusterRoleBinding, so
every install and upgrade required cluster-admin even when only namespaced
objects were needed. Installers that are namespace-admin GitOps or platform
controllers could not run the release at all, and clusters where cluster-scoped
RBAC is owned by a separate team had no supported way to split the install.

Add an rbac values block so a cluster-admin can apply the cluster-scoped objects
once and a namespace-admin can install and upgrade the release without
cluster-scoped permissions:

  rbac:
    create: true
    clusterScoped:
      create: true
      clusterRoleName: ""
      clusterRoleBindingName: ""

rbac.clusterScoped.create gates the ClusterRole and ClusterRoleBinding, and is
independent of the workspace mode. rbac.create additionally gates the namespaced
sandbox Role and RoleBinding, which matters because Kubernetes escalation
prevention stops an installer holding only the built-in admin role from creating
a Role that grants agents.x-k8s.io verbs it does not itself hold. The certgen
hook and credential driver RBAC keep their existing flags.

Both flags default to true, so current installs are unchanged. The helpers treat
a missing rbac block as enabled so upgrades with --reuse-values do not drop RBAC,
matching the existing workspaceResources pattern. The ClusterRoleBinding roleRef
follows clusterRoleName so a separately applied ClusterRole can carry a name the
cluster-admin chooses.

Document the migration for a release that already owns the cluster-scoped
objects: Helm deletes objects that leave the manifest, so annotate them with
helm.sh/resource-policy=keep before setting the flag, otherwise the gateway
loses TokenReview until a cluster-admin re-applies them.

Signed-off-by: ansjindal <ansjindal@nvidia.com>
2026-09-28 05:50:07 +00:00
Drew Newberry d4f5034d7f fix(gateway)!: make the WebSocket tunnel opt-in (#3727)
The /_ws_tunnel endpoint pipes a WebSocket into the full gRPC service. On a
plaintext loopback gateway any web page could open it, since browsers do not
apply CORS to WebSocket upgrades. Mount the tunnel only when
enable_websocket_tunnel is set (config file, --enable-websocket-tunnel, or
OPENSHELL_ENABLE_WEBSOCKET_TUNNEL; server.enableWebsocketTunnel in Helm).

BREAKING CHANGE: gateways behind an authenticating edge proxy must set
enable_websocket_tunnel = true for CLI edge-tunnel connections.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-25 23:43:18 +00:00
Drew Newberry 9244868056 docs: refresh architecture and agent guides (#3705)
* docs: refresh architecture and agent guides

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs: describe updated security architecture neutrally

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs: highlight new isolation primitives

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(sandboxes): clarify how to disconnect

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs: align architecture and guides with current navigation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(extensibility): streamline extension authentication guidance

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-25 02:38:00 -07:00
Drew Newberry 1374672967 fix(helm): restore Kubernetes e2e chart rendering (#3692)
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-24 18:55:45 -07:00
Drew Newberry 52cb8ecee7 fix(kubernetes): remove NetworkPolicy acknowledgement (#3677)
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-24 22:41:45 +00:00
Jason T. GreeneandMrunal Patel 7691f88e0e perf(server): enable WAL for the SQLite store; relax sync only for SSH session issuance (#3543)
* perf(server): enable WAL and NORMAL sync for the SQLite store

On-disk SQLite stores ran with sqlx defaults: rollback journal
(`journal_mode=delete`) and `synchronous=FULL`. Every autocommit write paid
several fsyncs and blocked readers while it held the lock, so gateway hot
paths made of many small writes serialized on disk latency. The clearest
case is `openshell forward service`, which mints and revokes an SSH session
token around every forwarded TCP connection: two commits per connection,
tens of milliseconds each on a virtual disk, wall clock linear in the
number of concurrent connections, and enough queueing that bursts hit the
per-sandbox connection cap and get refused.

Switch on-disk databases to WAL with `synchronous=NORMAL`. The mode change
runs once on a single connection before the pool opens: entering WAL needs
exclusive access to the file, so doing it up front means pool connections
only ever re-apply the pragma to a file already in WAL mode, and a failure
surfaces as one clear connect error. The first start after upgrading an
existing database therefore needs the file to be otherwise unopened.
`synchronous` is applied through the connect options on every pooled
connection. In-memory databases keep their defaults. A crash can now roll
back the most recent transactions without corrupting the database, which
is the standard WAL trade-off and fits the single-node scope of the SQLite
backend.

Tests cover a fresh store, an existing rollback-journal file that must be
switched on connect, sidecar permissions, and concurrent readers under a
burst of insert-then-update writes. Architecture, configuration and Helm
docs describe the durability trade-off, the sidecar files, and the local
filesystem requirement.

Signed-off-by: Jason T. Greene <jason.greene@redhat.com>

* fix(server): keep SQLite commits durable except SSH session issuance

WAL with synchronous=NORMAL can roll back acknowledged commits after a
power loss or kernel crash, including SSH session revocations and other
authorization-tightening writes. Run the main pool with synchronous=FULL
so every acknowledged write is durable; in WAL mode that is a single
fsync of the WAL per commit.

Add Store::create_relaxed for inserts that are safe to lose, and use it
only for SSH session issuance: a dropped token just fails validation.
On file-backed SQLite it runs on a dedicated single-connection pool with
synchronous=NORMAL. Both pools share one WAL, so the next FULL commit
also makes earlier relaxed commits durable. Postgres treats it as an
ordinary durable MustCreate insert.

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

---------

Signed-off-by: Jason T. Greene <jason.greene@redhat.com>
Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
Co-authored-by: Mrunal Patel <mrunalp@gmail.com>
2026-09-24 18:42:43 +00:00
Mark CampbellandKris Hicks 679b190677 feat(testing): support independent gateway and supervisor image overrides (#3341)
* feat(testing): normalize configurable test images

Signed-off-by: Bobbins228 <mcampbel@redhat.com>
Signed-off-by: Kris Hicks <khicks@nvidia.com>

* feat(helm): add global image overrides

Signed-off-by: Bobbins228 <mcampbel@redhat.com>

* feat(helm): support image registry overrides

Signed-off-by: Bobbins228 <mcampbel@redhat.com>
Signed-off-by: Kris Hicks <khicks@nvidia.com>

* refactor(helm): simplify image configuration

Signed-off-by: Bobbins228 <mcampbel@redhat.com>

* fix(e2e): avoid reloading reused kind sandbox image

Signed-off-by: Bobbins228 <mcampbel@redhat.com>

* fix(helm): default sandbox image to nvcr.io/nvidia/base/ubuntu:24.04

Signed-off-by: Kris Hicks <khicks@nvidia.com>

---------

Signed-off-by: Bobbins228 <mcampbel@redhat.com>
Signed-off-by: Kris Hicks <khicks@nvidia.com>
Co-authored-by: Kris Hicks <khicks@nvidia.com>
2026-09-23 23:35:30 +00:00
krishicks 0b351c4a9b fix(helm)!: reduce gateway Secret privileges (#3616)
* fix(driver-kubernetes-secrets)!: store provider credentials in one namespace

The Kubernetes Secrets credential driver now stores every credential in its
configured namespace in all workspace modes and rejects handles that reference
any other namespace before contacting the Kubernetes API. The gateway reaches
credential Secrets through the Role in that namespace; this allows removing the
Secret rules from the ClusterRole.

- Remove the workspace_mode, gateway_id, and allow_reference_namespace driver
  settings and stop rendering them from Helm. Configurations that set them fail
  at startup. Existing credential state is not migrated.
- Add server.credentialDrivers.kubernetesSecrets.createNamespace to provision a
  dedicated credential namespace. The namespace is kept on uninstall, adopted
  by a reinstall of the same release, and left untouched when something else owns
  it.
- Update the gateway config reference, Kubernetes setup docs, 0.1.0
  upgrade guide, compute-runtime architecture, and cluster debugging skill.

Signed-off-by: Kris Hicks <khicks@nvidia.com>

* fix(helm): reduce gateway Secret permissions

Remove the gateway's Secret list permission in every workspace mode and
grant source Secret reads through a Role in the sandbox namespace.
Bootstrap Secret cleanup deletes Secrets by exact name instead of listing
them.

- Grant get on the copied client TLS and image-pull Secrets through a Role in
  the sandbox namespace. The ClusterRole keeps get and patch on those names
  for the ownership check and server-side apply into workspace namespaces.
- Delete sandbox and supervisor bootstrap Secrets by exact name, derived from
  the runtime generation recorded on the Sandbox and, on restart, the target
  generation, tolerating 404. The generation annotation is cleared only after
  cleanup succeeds, and each bootstrap Secret has a Pod owner reference, so
  garbage collection removes any generation the driver does not name.
- Drop Secret list from the ClusterRole and the shared-mode sandbox Role.
- Extend the managed e2e RBAC checks to Secret list.
- Update the Kubernetes setup and sandbox runtime docs, compute-runtime
  architecture, and cluster debugging skill.

Signed-off-by: Kris Hicks <khicks@nvidia.com>

* fix(driver-kubernetes): stage workspace Secrets per runtime generation

Every Secret the Kubernetes driver writes into a workspace namespace is
now scoped to one sandbox runtime generation, immutable, and created
with create only, so the gateway never reads, patches, or adopts an
existing Secret there. This removes the gateway's cluster-wide get and
patch on the copied client TLS and image-pull Secret names.

- In managed mode, create an immutable copy of each configured
  image-pull Secret per generation, named os-pull-<id>-<generation>-<n>
  and owned by the generation's workload and supervisor Pods. Pods and
  the restarted Sandbox template reference those names, and generation
  cleanup deletes them by name. A Secret already holding a generation
  name fails the create.
- Outside shared mode, stage the gateway client TLS material into the
  supervisor bootstrap Secret instead of copying the client TLS Secret
  into the workspace namespace.
- Remove the fixed-name TLS and image-pull copies, the target ownership
  read, and the ClusterRole get and patch rule on the copied names.
  Source reads stay in the sandbox-namespace Role.
- Update the managed e2e to expect generation image-pull Secrets and
  client TLS material in the supervisor bootstrap Secret, and to check
  that the gateway cannot read the copied names in workspace namespaces.
- Update the gateway config and compute driver references, Kubernetes
  setup and sandbox runtime docs, compute-runtime architecture, driver
  README, and cluster debugging skill.

Signed-off-by: Kris Hicks <khicks@nvidia.com>

* fix(helm)!: grant operator-mode Secret permissions through the workspace chart

The operator-mode gateway ClusterRole grants no Secret permissions.
The openshell-workspace chart Role, installed in each operator-managed
namespace, grants the gateway create and delete on Secrets for sandbox
runtime generations. Operator-managed namespaces require the workspace
chart.

- Fail the chart tests on any ClusterRole rule that includes Secrets in
  operator and shared modes.
- Install the workspace chart when the operator e2e provisions a
  namespace, and assert that the gateway has no Secret permissions in a
  namespace without it.
- Update the Kubernetes setup docs, 0.1.0 upgrade guide, compute-runtime
  architecture, and cluster debugging skill.

Signed-off-by: Kris Hicks <khicks@nvidia.com>

---------

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-09-23 21:35:22 +00:00
Mrunal Patel a408f5dd08 chore(kubernetes): update Agent Sandbox to v1.0.3 (#3578)
Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
2026-09-23 18:30:24 +00:00
Drew Newberry 84960e70a3 fix(kubernetes): scope resource admission RBAC (#3571)
* fix(kubernetes): scope resource admission RBAC

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(helm): gate PVC admission reads

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-23 00:02:10 +00:00
Drew Newberry 1e34e8c576 fix(drivers): require admission labels for external resources (#3538)
* fix(drivers): require admission labels for external resources

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(drivers): address resource admission review findings

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(core): reserve driver-owned admission labels

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(core): clarify workspace admission label

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(drivers): clarify resource admission failures

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(e2e): configure resource admission fixtures

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): retry forbidden admission lookups

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(e2e): preserve external driver admission defaults

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-22 21:22:30 +00:00
Philippe Martin 35e0a68e4a feat(kubernetes): support corporate proxy CA bundle (#3447)
* feat(kubernetes): support corporate proxy CA bundle

The Kubernetes driver had no way to supply a CA bundle for the corporate
egress proxy, so an `https://` proxy with a private CA, or a TLS-intercepting
proxy, could not be used. Podman and VM already expose `proxy_ca_bundle`.

Add `proxy_ca_bundle` to `[openshell.drivers.kubernetes]` as a path the
gateway Pod reads. The gateway stages the PEM into the existing
per-generation supervisor bootstrap Secret and passes
`--upstream-proxy-ca-bundle` on the supervisor argv. That Secret is already
immutable, owner-referenced and garbage-collected, and its volume mounts
every key at /.openshell/supervisor with no items filter, so this needs no
new object kind, volume, mount, or RBAC verb, and works in shared, managed
and operator workspace modes.

The bundle is deliberately read from the gateway's filesystem rather than
referenced as an object in the sandbox namespace. It becomes a trust anchor
for every upstream the sandbox reaches, so it must stay in the gateway's
trust domain; the immutable staging Secret also keeps the anchor from
changing underneath a running sandbox.

Bound the staged bundle at 256 KiB. The shared reader's limit is exactly the
apiserver's own Secret limit and the bootstrap Secret carries four other
keys, so a bundle between the two would pass gateway startup and then fail
every sandbox create with an opaque `data: Too long`.

Delegate the URL, no_proxy, connect_by_hostname and ca_bundle rules to the
shared validate_upstream_proxy_settings, keeping the Secret-specific
credential block local: this driver accepts an explicit
`proxy_auth_allow_insecure = false` without credentials, which the shared
rules reject. This also fixes the acknowledgement being demanded for an
`https://` proxy, where the credential travels inside the verified TLS
session. Add auth_setting_label so the inline-credential diagnostic names
the Secret keys instead of proxy_auth_file, which this driver rejects as an
unknown key.

Document that the bundle should carry only the CA that signs the proxy's
certificate, or that an intercepting proxy re-signs upstream certificates
with. Public roots already reach the sandbox through the supervisor image and
its TLS stack, and the bundle is concatenated with that system store into a
single boundary control frame, so a full merged trust bundle spends the frame
budget on duplicated roots. The frame, not the apiserver Secret limit, is the
tighter of the two ceilings in practice; raising the staging bound requires
checking it.

Closes #3443

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(helm): quote proxy CA ConfigMap references

Signed-off-by: Philippe Martin <phmartin@redhat.com>

---------

Signed-off-by: Philippe Martin <phmartin@redhat.com>
2026-09-22 17:57:40 +00:00
50230616d5 refactor(runtime): retire Community image dependencies (#3386)
* feat(sandbox): default to official Alpine sandbox image

default_sandbox_image() now returns docker.io/library/alpine:3.22, a generic
version-qualified official image, so a fresh install no longer depends on the
community sandbox image catalog. All compute drivers (docker, podman,
kubernetes, vm) inherit this fallback.

Part of #3116.

Signed-off-by: Akram
Signed-off-by: Akram <akram.benaissi@gmail.com>

* feat(deploy): default deployment configs to the official Alpine sandbox image

Update the shared gateway default_image, Helm chart values, the standalone
Kubernetes manifest, and the dev gateway task scripts to use
docker.io/library/alpine:3.22 instead of the community base image, consistent
with default_sandbox_image(). GPU e2e image-build base is left unchanged (CUDA
needs a glibc base).

Part of #3116.

Signed-off-by: Akram
Signed-off-by: Akram <akram.benaissi@gmail.com>

* feat(driver): default to numeric non-root identity for USER-less images

With the default sandbox image now Alpine, images that declare no OCI USER
must start instead of being rejected. When the image declares no USER and
the policy requests none, the Podman and Docker drivers now supply a numeric
non-root identity (DEFAULT_SANDBOX_UID/GID = 1000) instead of rejecting,
matching the numeric-identity behavior of the Kubernetes and VM drivers. The
supervisor's resolved-identity path runs the sandbox as a synthesized
non-root account without the account existing in the image. Images that
declare a USER keep the OCI resolution path unchanged.

Part of #3116.

Signed-off-by: Akram <akram.benaissi@gmail.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(conformance): use Alpine workload image

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* refactor(policy): drop community image /app path from default policy

The restrictive default policy granted read-only access to /app, a directory
that only existed in the community base image. A generic Alpine default has no
/app, so remove it. Landlock best-effort already ignores absent paths; this
just stops advertising a community-specific layout in the default.

Part of #3116.

Signed-off-by: Akram
Signed-off-by: Akram <akram.benaissi@gmail.com>

* docs(config): document Alpine default images

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(podman): report early sandbox termination

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(podman): initialize rootless workspace ownership

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(sandbox): qualify NVIDIA Ubuntu default

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): initialize rootful default workspace

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(sftp): add native sandbox adapter

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sftp): gate runtime helper support to Linux

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sftp): support standard OpenSSH file operations

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sftp): harden rename and special file handling

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(runtime): remove community image dependencies

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(e2e): build provider readiness tool fixture

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(e2e): use a dedicated Noble fixture for Docker tests

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Akram
Signed-off-by: Akram <akram.benaissi@gmail.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Evan Lezar <elezar@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
2026-09-22 14:43:51 +02:00
0a770d9173 feat(kubernetes): support HA gateway rebalancing (#1868)
* feat(kubernetes): support HA gateway rebalancing

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* perf(server): cache peer connections, tokens, and owner lookups

Every forwarded relay rebuilt its setup from scratch: an owner lookup, a
blocking read of the peer token, a TLS connect to the owning replica, and
a TokenReview plus Pod GET on the receiving side. Sandbox service routing
does this per HTTP request, so the apiserver calls scaled with traffic.

Cache all of it on ServerState:

- peer channels pooled per endpoint, so relays multiplex over one
  connection instead of redialing
- peer tokens keyed by SHA-256, expiring at min(ttl, token exp) so a hit
  cannot accept an expired token
- owner records for 3s against a 45s ownership TTL, still freshness
  checked before use

Entries are evicted when a relay fails. Also raise HTTP/2
max_concurrent_streams to 1024, since pooling funnels every relay between
two replicas onto one connection and hyper's default of 200 sits below
the 256 pending-relay budget.

Signed-off-by: divesh <dgude@nvidia.com>

* perf(server): pool upstream connections for sandbox services

Each HTTP request to a sandbox service opened its own supervisor relay,
paying a new TCP connection and HTTP/1 handshake every time. Worse, it
counted against the 32 in-flight relay cap, so a service handling more
than 32 concurrent requests failed outright.

Pool idle upstreams per endpoint and port, up to 8 each for 15s. Reuse is
safe because the pool only returns a connection hyper reports as ready,
and HTTP/1 cannot start a request until the previous body has drained.
Upgrades are never pooled since they take the connection over, and a
failed send evicts that endpoint. Pruning is bounded per key, with the
full sweep limited to once per 30s.

Signed-off-by: divesh <dgude@nvidia.com>

* fix(server): address HA gateway review findings (#3449)

- Let a gateway own supervisor sessions without a peer endpoint. Requiring
  one whenever the store is PostgreSQL broke every single-instance
  PostgreSQL deployment, because no sandbox supervisor could connect.
  A cross-replica request to an owner that advertises no endpoint now fails
  immediately naming the cause, instead of retrying until the wait timeout.
- Close a supervisor session on heartbeat only when another replica owns it,
  or after renewals fail for the ownership TTL. A database error no longer
  drops every session heartbeating during an outage.
- Clamp owner record ages at zero so a skewed or corrupt stored timestamp
  cannot produce a negative age.
- Bound the cross-object advisory lock with a lock timeout, so a stuck holder
  fails instead of blocking every mutation in the fleet.
- Refuse to start when a peer endpoint is configured on a multi-replica
  backend but peer authentication is unavailable, and warn when a
  multi-replica backend has no peer endpoint at all.
- Reject a plaintext peer endpoint when the gateway serves TLS.
- Skip the sandbox watch poller on single-replica backends, where the local
  update bus already sees every write.
- Rate-limit the peer owner cache sweep so an insert no longer scans the
  whole map under the lock.
- Retry GET and HEAD on a pooled upstream the sandbox closed, instead of
  returning 502, and drop an emptied endpoint from the pool right away.
- Document the gateway peer environment variables and the post-rollout
  ownership skew operators should expect.

Signed-off-by: divesh <dgude@nvidia.com>

* fix(server): harden HA supervisor ownership

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Signed-off-by: divesh <dgude@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: divesh <dgude@nvidia.com>
Co-authored-by: Divesh Chowdary <47188680+FrostGod@users.noreply.github.com>
2026-09-21 21:22:12 +00:00
Drew Newberry d6f3e1f5f2 fix(packaging): keep SPDX comments out of Debian control (#3483)
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-19 18:55:33 -07:00
Drew Newberry 17ce738bfb fix(ci)!: remove gateway callback listener dependency (#3365)
* fix(ci): repair post-merge release canary

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(packaging): bootstrap canary runtime prerequisites

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci(canary): collect macOS VM diagnostics

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci(canary): pin libkrun-compatible macOS runner

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci(canary): limit macOS smoke test to package startup

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(compute)!: remove gateway callback listeners

Run Docker supervisors on host networking so they use the operator-configured primary gateway endpoint. Remove the unused compute-driver callback listener negotiation and listener-scoped routing machinery.

BREAKING CHANGE: The ComputeDriver API no longer exposes GetGatewayListenerRequirements or GatewayListenerRequirement. External drivers must regenerate bindings and connect supervisors to the configured primary gateway endpoint.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(docker): use sandbox runtime image in launcher

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(podman): exercise production endpoint selection

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(e2e): route supervisors to reachable gateways

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(e2e): align Podman endpoint fixtures

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): preserve host aliases for supervisors

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): align sandbox host gateway pin

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(e2e): address Docker fixtures by bridge IP

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(e2e): serialize sandbox lifecycle cases

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(e2e): host Docker TCP fixture with gateway

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(e2e): use loopback for host-network supervisor

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-18 20:55:55 +00:00
John T. Myers 903d9a0e7a chore(license): align repository compliance text (#3467)
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-09-18 17:05:18 +00:00
Drew NewberryandPiotr Mlocek 8b77925eb7 feat(installer): support prerelease installations (#3364)
* feat(installer): support prerelease installations

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(readme): clarify 0.1.0 production readiness

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(readme): refine production readiness message

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(release): publish rolling prerelease channels

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(release): keep prereleases out of GitHub releases

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(readme): refine 0.1.0 notice

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(readme): expand 0.1.0 notice

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(readme): simplify 0.1.0 notice

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(installer): simplify prerelease alias

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(installer): require successful prerelease runs

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(release): reuse exact prerelease artifacts

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(installer): include prover in prerelease bundle

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci: remove prerelease post-publish validation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(fern): sync announcement configuration

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs: remove stale prerelease canary claim

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(fern): defer announcement synchronization

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(fern): simplify release announcement

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Co-authored-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-09-18 00:53:36 +00:00
Drew Newberry 07d4ac5474 fix(helm): grant secret cleanup permissions (#3363)
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-17 19:36:19 +00:00
Johnny Greco 58b5f8f976 feat(prover): add standalone policy boundary checker (#3289)
* feat(prover): add standalone policy maximum checker

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* refactor(prover): simplify check scope schema

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(prover): align containment and cancellation with runtime

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(prover): stabilize containment checks in CI

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* test(prover): avoid solver in fast-path guard test

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(prover): align string containment with runtime

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(prover): reject ambiguous z3 string escapes

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(prover): align containment with runtime boundaries

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(prover-cli): harden cancellation and invalid input

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* feat(packaging): install policy prover with OpenShell

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(prover): clarify installation and check results

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(build): describe prover distribution directly

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* refactor(prover): rename maximum policy to boundary

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* refactor(prover): localize fail-closed validation

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* test(prover): cover fail-closed CLI surfaces

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* test(prover): allow CI load for REST solver proof

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(prover): use canonical policy schema for containment

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(prover): preserve uncertainty for runtime binary globs

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(prover): bound policy validation work

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(prover): document validation resource limits

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(prover): make containment API extensible

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(prover): define containment API contract

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(ci): integrate prover with consolidated builds

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(ci): declare release packaging dependency

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

---------

Signed-off-by: Johnny Greco <jogreco@nvidia.com>
2026-09-17 17:30:03 +00:00
John T. Myers 0fd3385726 fix(container): use distroless Debian 13 for supervisor (#3393)
* fix(container): use distroless Debian 13 for supervisor

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(docker): preserve supervisor bootstrap ownership

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-09-17 05:29:47 +00:00
krishicks 12a7a35910 chore(tools): upgrade mise to 2026.9.9 (#3385)
Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-09-16 09:04:50 -07:00
Drew Newberry c1f2e7189f feat(isolation): implement the RFC 0012 sandbox architecture (#2942)
* feat(isolation): add RFC 0012 backend contract

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>

* refactor(isolation): name the interface crate explicitly

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): expose trusted host gateway

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(agents): inventory the MXC driver

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): add mediated DNS transport

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): tighten interface error and digest contracts

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(isolation): remove unrelated driver inventory

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): define capability-free launch contract

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): seal confirmed boundary state

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): validate confirmation for external backend implementations

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): clarify mediated DNS identity

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(isolation): generalize loopback connector

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(isolation): unify typed network mediation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): bind launches to sandbox sessions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(mxc): initialize extended sandbox status

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): add boundary protocol and Linux primitives

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): harden signals and separate process status from transport

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): validate remote confirmation through public contract

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): validate wire state and propagate snapshot failures

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(isolation): import owned agent specification explicitly

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(isolation): describe mediated DNS channel

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): bound mediation attach without nested retries

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(isolation): generalize loopback protocol

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): add transport-neutral session authentication

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(isolation): separate sandbox backend protocol

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): harden runtime boundary controls

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): add terminal boundary operation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): split supervisor and sandbox runtimes

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): harden boundary isolation and lifecycle ownership

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): reject private root redirects and adopt typed errors

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): preserve accept thread ownership on musl

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(sandbox): isolate credential probes from filtered threads

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): return retained exec exit status to independent waiters

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): bound network mediation and preserve socket authorization

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): bound control admission and retire stale mediation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci(e2e): select migrated drivers per stack layer

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(sandbox): implement loopback connector

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): authenticate the Sandbox Protocol

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(supervisor): rotate launch-scoped authentication

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(sandbox): consume dedicated backend crate

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(sandbox): align topology session fixture

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): align projected bootstrap bundle

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(auth): validate refreshed credentials before rotation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): fail closed across supervisor disconnects

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): repair rebased sandbox CI

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* build(runtime): publish separate sandbox and supervisor images

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(config): configure the sandbox runtime image

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(ci): validate sandbox binary linkage

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(isolation): use backend and runtime terminology

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(sandbox): use a scratch runtime image

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(ci): refresh schema and dependency policy

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): bind reconnects to supervisor process

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs: align runtime split operational guidance

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* chore(security): document Kubernetes runtime RBAC

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): enforce runtime lifecycle invariants

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(compute): identify sandbox start generations

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(server): restore sandbox launch sessions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): support authenticated runtime replacement

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(auth): bind sandbox session successors

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(auth): retry pending sandbox successors

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(vm): run the supervisor outside the guest workload

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(vm): use sandbox backend protocol

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): repair rebase integration

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): use unified build toolchain

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(vm): use sandbox runtime terminology

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(vm): own guest network bootstrap

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): expose guest init version

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): select native supervisor artifacts

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): guard guest init Linux symbols

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): scope Linux test imports

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): avoid guest interface casts

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): reconcile admitted sandbox identity

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): share resolved sandbox identity

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): surface host supervisor failures

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): include guest logs on supervisor exit

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): rotate and clean runtime generations

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): make sandbox starts generation-aware

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): rotate restored sandbox sessions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(vm): keep shared paths in the base layer

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): bind sandbox session lineage

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(docker): isolate workloads behind the host supervisor

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(docker): rotate launch-scoped authentication

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(docker): use sandbox backend protocol

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(docker): use host networking for supervisor

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): preserve host gateway alias resolution

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(docker): use separate sandbox and supervisor images

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): restore startup validation after rebase

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(docker): name the sandbox runtime directly

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): narrow supervisor CA runtime storage

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): close companion isolation gaps

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(docker): align mediated network expectations

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(docker): exercise mediated network paths

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): attach supervisor to managed network

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): defer supervisor recovery until gateway is ready

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): make sandbox starts generation-aware

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): rotate restored sandbox sessions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): preserve workloads during session rotation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(docker): remove unrelated configuration RFC changes

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): bind sandbox session lineage

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(kubernetes): add proxy-pod isolation topology

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(kubernetes): use sandbox backend protocol

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): use stable sandbox service authority

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(kubernetes): split sandbox and supervisor images

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): adapt proxy pods to current runtime APIs

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(kubernetes): describe the single runtime placement

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(kubernetes): simplify sandbox orchestration

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): validate deployment prerequisites

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): update Trivy Helm profile inventory

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(kubernetes): update Trivy scan inventory count

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): reuse preloaded runtime images in e2e

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): type and clean runtime resources

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): make sandbox restarts recoverable

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): rotate restored sandbox sessions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): preserve supervisor egress

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): bind sandbox session lineage

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(podman): adopt isolated sandbox and supervisor containers

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): stage bootstrap archives at named volume destinations

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(podman): rotate launch-scoped authentication

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(podman): use sandbox backend protocol

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(podman): use host networking for supervisor

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(podman): split sandbox and supervisor images

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): repair rebase integration

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(podman): name the sandbox runtime directly

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): provision supervisor CA runtime storage

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): address isolation review findings

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): inspect Debian supervisor provenance

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): use libpod-compatible tmpfs options

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): bind verified sandbox runtime binary

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): provide external driver data directory

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): start sandbox before joining user namespace

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): separate supervisor user namespace

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): make sandbox starts generation-aware

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): rotate restored sandbox sessions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): bind sandbox session lineage

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* perf(isolation): add TCP and DNS benchmark harnesses

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(perf): align benchmark timing and supported protocols

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(perf): report TCP benchmark metrics accurately

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(perf): cancel failed worker startup

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): build matching local supervisor image

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): make local sandbox smoke test runnable

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): wire local sandbox runtime image

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): narrow sandbox service RBAC

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(ci): validate split runtime artifacts

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): harden runtime session handling

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(supervisor): add standalone network proxy role

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(rfc): remove implementation companion notes

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(vm): standardize runtime release name

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): pin renamed runtime artifacts

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(auth): persist sandbox runtime identity

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(runtime): restore branch validation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): reconcile main after rebase

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(network): close unframed HTTP 1.0 responses

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* chore(isolation): preserve upstream OCSF updates

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(security): close credential and TLS replay paths

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(auth): make sandbox refresh retries idempotent

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-16 00:49:14 +00:00
Seth Jennings dbe36eaf85 fix(security): harden Vault credential transport (#3329)
Reject non-loopback plaintext Vault endpoints, disable redirects, and support private CA bundles without weakening hostname verification. Update Helm configuration, documentation, operator skills, and regression coverage for OSSR-002.

Signed-off-by: Seth Jennings <sjenning@redhat.com>
2026-09-15 19:56:09 +00:00
John T. Myers 607db99915 fix(deps): refresh gateway Debian runtime image (#3350)
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-09-15 19:38:13 +00:00
Mrunal Patel b799fccb8b fix(auth): harden OIDC trust root retrieval (#3332)
* fix(auth): harden OIDC trust root retrieval

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(e2e): pass OIDC HTTP acknowledgement value

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

---------

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
2026-09-15 17:00:45 +00:00
Brandon Squizzato cc780d4e17 feat(helm): add BackendTLSPolicy support (#2728)
* feat(helm): add optional BackendTLSPolicy for e2e TLS

Add grpcRoute.backendTLSPolicy values to optionally create a
BackendTLSPolicy resource that enables end-to-end TLS between the
Gateway proxy and the OpenShell gateway pod. The Gateway proxy
terminates client-facing TLS and re-encrypts when connecting to the
backend, validating the pod's certificate against a user-supplied CA
ConfigMap.

This removes the requirement to set server.disableTls=true when using
HTTPS at the Gateway listener. Supported on OpenShift 4.22+ and other
platforms with BackendTLSPolicy support in the Gateway API
implementation.

Update OpenShift and ingress documentation with e2e TLS instructions.

Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>

* feat(helm,server): auto-create backend CA ConfigMap in certgen hook

Extend the generate-certs command with --backend-ca-configmap-name and
--backend-ca-source-secret flags. When BackendTLSPolicy is enabled, the
certgen pre-install hook creates the CA ConfigMap automatically:

- pkiInitJob mode (default): uses the CA from the generated PKI bundle.
  Fully automatic on first install.
- cert-manager mode: reads ca.crt from the server TLS Secret. On first
  install the Secret does not exist yet (cert-manager reconciles after
  templates are applied), so the ConfigMap is created on the first helm
  upgrade. Logs a warning on the initial skip.

The caCertificateConfigMapName value now defaults to <fullname>-backend-ca
when empty, so users only need to set backendTLSPolicy.enabled=true.

Update certgen RBAC to include configmaps get/create when the feature is
enabled. Add CLI arg parsing tests for the new flags.

Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>

* refactor(helm): add server.tls.enableMtls flag for mTLS control

Replace automatic mTLS disabling based on BackendTLSPolicy with an
explicit server.tls.enableMtls flag that defaults to true. The user is
now responsible for setting this to false when using BackendTLSPolicy,
as ingress proxies cannot present client certificates to backends.

Updated:
- values.yaml: Added server.tls.enableMtls (default true)
- gateway-config.yaml: Check enableMtls instead of backendTLSPolicy
- _gateway-workload.tpl: Check enableMtls for client CA mount
- Tests: Updated to use enableMtls flag
- Docs: Added enableMtls=false to BackendTLSPolicy examples
- README: Document new flag and BackendTLSPolicy requirement

Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>

* docs(helm): clarify cert-manager backend CA ConfigMap workflow

Update documentation to explain the two-step install process required
when using cert-manager with BackendTLSPolicy:

1. helm install - cert-manager issues the server certificate, but the
   certgen hook can't create the backend CA ConfigMap yet (cert-manager
   reconciles after templates are applied)
2. helm upgrade - certgen hook reads the CA from the cert-manager-issued
   certificate and creates the ConfigMap

Previously, the docs said "created on first upgrade" without explaining
why or that the feature won't work until then. The updated docs now:
- Explain the timing issue (cert-manager reconciles after chart install)
- Provide clear steps for the cert-manager workflow
- Note that pkiInitJob (default) creates it immediately on install
- Clarify that users must wait for the Certificate to be Ready before
  running the second upgrade

Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>

* fix(docs): remove incorrect external hostname requirement for BackendTLSPolicy

BackendTLSPolicy validates the backend certificate against the service FQDN
(e.g., openshell.openshell.svc.cluster.local), not the external hostname.
The external hostname only needs to be on the Gateway listener certificate
for client-facing TLS.

The default certManager.serverDnsNames already includes all required service
FQDN variants, so no configuration is needed for BackendTLSPolicy to work.

Fixed incorrect documentation that claimed:
- "The server certificate SAN list must include the external hostname"
- Users need to "configure certManager.serverDnsNames with the external hostname"

Removed the unnecessary pkiInitJob.serverDnsNames override from the example
and clarified that:
- Gateway listener certificate needs the external hostname (for clients)
- Backend certificate needs the service FQDN (for Gateway proxy)
- The service FQDN is already in the defaults

Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>

* docs: clarify ACME with LetsEncrypt reference

Change all references from "ACME issuer" to "LetsEncrypt/ACME issuer"
to help users understand that LetsEncrypt is the most common ACME
provider and what ACME means in practice.

Updated:
- docs/kubernetes/managing-certificates.mdx
- docs/kubernetes/openshift.mdx
- deploy/helm/openshell/values.yaml
- deploy/helm/openshell/README.md
- deploy/helm/openshell/ci/values-openshift-route-cert-manager.yaml

Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>

* docs(openshift): restructure end-to-end TLS options and clarify Gateway hostname

Reorganize the OpenShift production deployment documentation:

1. Changed main section from "Production Deployments" to "Options for
   end-to-end TLS" for better clarity

2. Renamed subsections for consistency and clarity:
   - "End-to-end TLS using Gateway API and BackendTLSPolicy (OpenShift 4.22+)"
   - "End-to-end TLS using pass-through Route (all OpenShift versions)"

3. Clarified that the Gateway hostname is typically a wildcard:
   "typically a wildcard like *.openshell-ingress-gw.example.com"

4. Removed the recommendation to copy the cluster's wildcard certificate
   from openshift-ingress namespace, as this is not a recommended
   security best practice

These changes make it clearer that users have two end-to-end TLS options
and help them understand the typical naming pattern for Gateway hostnames.

Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>

* feat(helm): eliminate two-stage install for BackendTLSPolicy with cert-manager

When using BackendTLSPolicy with cert-manager, the certgen hook now polls
for up to 90 seconds waiting for cert-manager to issue the TLS certificate
before creating the backend CA ConfigMap. This eliminates the need for a
second `helm upgrade` in most cases.

The hook polls every 2 seconds with progress logging every 10 seconds.
If cert-manager takes longer than 90 seconds, the hook times out gracefully
and logs a warning, preserving the fallback to manual ConfigMap creation
or a second upgrade.

The Job's activeDeadlineSeconds is 120s, so the 90s timeout leaves 30s
margin for ConfigMap creation and hook completion.

Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com>

* feat(helm): add configurable timeout for certgen hook

Add `pkiInitJob.timeoutSeconds` Helm value (default 120) to control how long
the certgen hook Job can run. When using cert-manager with BackendTLSPolicy,
the hook polls for (timeoutSeconds - 30) seconds to leave margin for ConfigMap
creation and cleanup.

This allows users to increase the timeout for environments where cert-manager
takes longer than 90 seconds to issue certificates, without requiring code
changes.

Example usage:
```yaml
pkiInitJob:
  timeoutSeconds: 180  # Hook polls for 150 seconds
```

Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com>

* docs(helm): document configurable certgen timeout

Update documentation to mention the pkiInitJob.timeoutSeconds value and
how it affects the cert-manager polling behavior when using BackendTLSPolicy.

Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com>

* feat(helm): add configurable failure behavior for certgen timeout

Add `pkiInitJob.failOnTimeout` Helm value (default false) to control whether
the certgen hook fails or succeeds when cert-manager does not issue a
certificate within the polling timeout.

When false (default), the hook succeeds with a warning and users can run
`helm upgrade` after cert-manager issues the certificate to create the
backend CA ConfigMap. This provides backwards-compatible behavior.

When true, the hook fails immediately if the timeout is reached, providing
clear feedback that BackendTLSPolicy is non-functional. This is useful for
strict validation requirements where incomplete installs should fail fast.

Example usage:
```yaml
pkiInitJob:
  timeoutSeconds: 180
  failOnTimeout: true  # Fail install if cert-manager takes >150s
```

Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com>

* feat(helm): change failOnTimeout default to true and add troubleshooting docs

Change `pkiInitJob.failOnTimeout` default from false to true to provide
immediate feedback when cert-manager does not issue certificates within
the polling timeout. This prevents silent failures where BackendTLSPolicy
is non-functional but the install appears to succeed.

Add comprehensive troubleshooting section to docs/kubernetes/ingress.mdx
documenting the specific error "TLS error: Secret is not supplied by SDS"
that occurs when the backend CA ConfigMap is missing, with step-by-step
resolution instructions.

Updated comments in values.yaml to clearly document the default behavior
and explain when administrators might see connectivity errors if they
override the default to failOnTimeout=false.

BREAKING CHANGE: pkiInitJob.failOnTimeout now defaults to true. Helm
installs will fail if cert-manager takes longer than (timeoutSeconds - 30)
seconds to issue certificates. To restore the old behavior of allowing
installs to succeed with a warning, set `pkiInitJob.failOnTimeout=false`.

Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com>

* fix(helm): make cert-manager resources pre-install hooks to fix ordering

Make Certificate and Issuer resources run as pre-install/pre-upgrade hooks
with weight -30, before the certgen hook (weight -20). This fixes the
chicken-and-egg problem where the certgen hook was waiting for Secrets
created by Certificates that hadn't been created yet.

**Hook ordering:**
1. Certificate and Issuer resources created (weight -30)
2. cert-manager issues certificates and creates Secrets
3. certgen hook runs (weight -20), finds Secrets, creates ConfigMap
4. Main resources (StatefulSet, Service, etc.) created

Previously, the certgen pre-install hook would run before any resources
were created, poll for a non-existent Secret, timeout, and fail. The
Certificate resources would never get created because Helm waits for
all pre-install hooks to succeed before creating main resources.

This fix allows single-stage installs to work reliably as long as
cert-manager can issue certificates within the polling timeout.

Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com>

* feat(helm): add validation to prevent enableMtls with BackendTLSPolicy

Add Helm chart validation that fails the install if both
server.tls.enableMtls=true and grpcRoute.backendTLSPolicy.enabled=true
are set, since this is an invalid configuration.

BackendTLSPolicy requires mTLS to be disabled because the Gateway proxy
cannot present client certificates to the backend. This validation provides
immediate, clear feedback at install time rather than allowing the
misconfiguration to be discovered through runtime errors.

Example error message:
```
Error: grpcRoute.backendTLSPolicy requires mTLS to be disabled because
the Gateway proxy cannot present client certificates to the backend;
set server.tls.enableMtls=false
```

Also updated documentation to mention this validation check.

Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com>

* docs(helm): clarify pkiInitJob.timeoutSeconds polling behavior

Improve documentation to clearly explain that pkiInitJob.timeoutSeconds
controls the Job deadline, but the actual polling timeout is
(timeoutSeconds - 30) to reserve 30 seconds for ConfigMap creation
and cleanup.

Added concrete example: "timeoutSeconds=180 allows 150 seconds of polling"
to make the relationship explicit and avoid confusion where users might
expect the hook to poll for the full timeout value.

Updated both values.yaml inline comments and ingress.mdx documentation
for consistency.

Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com>

* docs(openshift): remove outdated two-stage install instructions

Update OpenShift documentation to reflect that single-stage installs now
work with cert-manager and BackendTLSPolicy. The Certificate resources
run as pre-install hooks (weight -30) before certgen (weight -20),
allowing the hook to poll for and find the issued certificates.

Removed the outdated two-step process:
1. helm install (cert-manager issues cert, hook logs warning)
2. helm upgrade (hook creates ConfigMap)

Replaced with current single-stage behavior:
- Certificate resources created as pre-install hooks
- certgen hook polls for up to 90 seconds (configurable)
- Single helm install succeeds in most cases
- Fails fast by default if timeout reached

This brings openshift.mdx in line with the already-updated ingress.mdx
documentation.

Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com>

* fix(helm): make pkiInitJob.timeoutSeconds the actual polling duration

The timeout value now represents the actual polling time that users
experience when waiting for cert-manager to issue certificates.
The Job activeDeadlineSeconds is set to (timeoutSeconds + 30) to
allow buffer time for ConfigMap creation and cleanup.

Previously, the hook polled for (timeoutSeconds - 30) seconds, which
was confusing when users set timeoutSeconds=180 and only got 150
seconds of actual polling.

Updated documentation in values.yaml, ingress.mdx, and openshift.mdx
to reflect the clearer behavior.

Signed-off-by: Brandon Squizzato <bsquizza@nvidia.com>

* docs(helm): update values.yaml and README with correct polling duration

Updated the caCertificateConfigMapName description to reflect that the
hook polls for exactly pkiInitJob.timeoutSeconds seconds, not
(timeoutSeconds - 30) seconds.

Regenerated README.md with helm-docs.

Signed-off-by: Brandon Squizzato <bsquizza@nvidia.com>

* docs(kubernetes): add OIDC configuration to helm install and CLI examples

Updated all helm install and openshell gateway add examples in ingress.mdx
and openshift.mdx to include OIDC issuer and audience configuration.

Examples now use concrete placeholder values:
- OIDC issuer: https://keycloak.example.com/realms/openshell
- OIDC audience: openshell-cli
- Hostname: gateway.example.com
- ClusterIssuer: letsencrypt-prod

This makes it clearer how to configure OIDC authentication, which is
required when using BackendTLSPolicy or HTTPS termination since the
Gateway proxy cannot present client certificates to the backend.

Signed-off-by: Brandon Squizzato <bsquizza@nvidia.com>

* docs(kubernetes): explicitly list OIDC client ID in gateway add examples

Added --oidc-client-id openshell-cli to all openshell gateway add
commands in ingress.mdx and openshift.mdx, making the default client
ID explicit in the examples even though it's the CLI default.

This improves clarity and helps users understand the complete OIDC
configuration needed for gateway registration.

Signed-off-by: Brandon Squizzato <bsquizza@nvidia.com>

* fix(helm): address PR review feedback for BackendTLSPolicy

- Read backend CA from the authoritative server Secret instead of the
  in-memory PKI bundle so enabling BackendTLSPolicy on an existing
  release uses the CA that actually signed the server certificate.
- Reconcile the backend CA ConfigMap on every hook run (compare and
  update) instead of skipping when it already exists, so CA rotations
  propagate automatically.
- Remove hook annotations from cert-manager Issuer/Certificate resources
  so they remain regular release objects managed by Helm lifecycle. Split
  the cert-manager backend CA ConfigMap creation into a separate
  post-install/post-upgrade hook Job that polls after cert-manager
  Certificate resources are applied.
- Update architecture/gateway.md, docs/reference/gateway-config.mdx,
  debug-openshell-cluster skill, and helm-dev-environment skill with
  BackendTLSPolicy, backend CA ConfigMap, enableMtls, and timeout
  documentation.

Signed-off-by: Brandon Squizzato <bsquizza@nvidia.com>
Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>

* fix(docs): resolve markdown lint errors in helm README and kubernetes docs

Escape inline HTML angle brackets in README.md template placeholders,
remove trailing spaces, and add blank lines around fenced code blocks
in numbered lists.

Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>

* Update docs

* fix(helm): escape inline HTML in values.yaml descriptions and sync mise lockfile

Wrap `<fullname>` and `<namespace>` template placeholders in backticks
so markdownlint does not flag them as inline HTML (MD033). Regenerate
mise.lock to match current mise.toml after rebase onto main.

Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>

* Run 'mise lock'

* fix: align mise.lock with CI mise version output

The lockfile was regenerated locally with mise 2026.8.10 which resolves
uv Linux artifacts to gnu variants and adds provenance_verified fields,
but CI uses v2026.4.25 which produces musl variants without those fields.

Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>

* fix(docs): correct cert-manager hook ordering and clientCaSecretName comment

Update ingress.mdx and openshift.mdx to describe Certificate resources
as regular release objects with a post-install/post-upgrade Job, matching
the current implementation and architecture/gateway.md.

Fix values.yaml clientCaSecretName comment to state that "" disables
client certificate verification, matching the helper and access-control
docs.

Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>

---------

Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>
Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com>
Signed-off-by: Brandon Squizzato <bsquizza@nvidia.com>
2026-09-14 17:39:15 +00:00
krishicks 5b9daab935 fix(ci): restore mise run ci on macOS (#3294)
- Replace BSD-incompatible in-place sed calls with portable temp-file rewrites.
- Remove test-only shell interception and capture generated gateway config
  directly.
- Allow parity tests to use supplied supervisor binaries without resolving a
  Linux target.
- Normalize temporary-directory paths and use portable RPM config installation.
- Set a valid setuptools-scm version for Python protobuf generation in Jujutsu
  checkouts.

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-09-12 00:17:15 +00:00
Brandon Squizzato 99e83a535e feat(helm): scope ClusterRole/ClusterRoleBinding names by release namespace (#2939)
* feat(helm): scope ClusterRole/ClusterRoleBinding names by release namespace

The chart creates cluster-scoped ClusterRole and ClusterRoleBinding
resources with a fixed name derived from the release name. When
multiple Helm releases coexist on the same cluster (multi-tenant),
only one release can own these resources due to Helm ownership
annotations -- the second install fails with a conflict.

Append .Release.Namespace to the ClusterRole and ClusterRoleBinding
names so each release gets its own cluster-scoped resources. The
duplication is harmless (the rules are identical and small) and
eliminates multi-tenant conflicts entirely without requiring external
RBAC management.

Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>

* test(helm): add regression tests for namespace-scoped ClusterRole names

Assert the generated ClusterRole name, ClusterRoleBinding name, and
roleRef all include the release namespace suffix so multi-namespace
installations cannot silently regress to conflicting fixed names.

Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>

---------

Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>
2026-09-11 17:57:42 +00:00
alangou 9b4b63ec69 fix(deps): update DOMPurify and runtime image packages (#3276)
Signed-off-by: Adrien Langou <alangou@nvidia.com>
2026-09-11 10:45:29 +00:00
Jesse JaggarsandDrew Newberry 02b664bb0d refactor(config): normalize and enforce gateway schema v2 (#2814)
* refactor(config): normalize compute driver field names

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* refactor(config): introduce canonical gateway fields

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* refactor(config): enforce gateway schema version 2

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): preserve compute driver runtime guarantees

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): address schema v2 review regressions

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): complete schema v2 migration safeguards

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(config): expand schema v2 regression coverage

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(config): add schema v2 parity manifest

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): correct parity manifest inventory

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* docs(config): record schema v2 intentional changes

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* docs(config): disposition schema v2 parity gaps

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): add dual schema parity harness

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): establish compute lifecycle parity baseline

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): preserve gateway option compatibility

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): record gateway option parity

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* docs(config): close gateway-wide parity gaps

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(podman): apply configured pids limit

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): validate Podman option parity

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): add Kubernetes option parity harness

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): record Kubernetes option parity

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): disposition VM parity lanes

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): add external driver parity lane

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(e2e): preserve external driver pull policy

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): attest parity artifacts and launches

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): require clean parity build sources

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): bind parity runtime artifacts

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(e2e): use isolated supervisor tags

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(e2e): qualify parity image tags

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(e2e): serve parity supervisor locally

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): isolate parity podman services

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): harden parity evidence provenance

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): pin parity sandbox artifacts

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): attest parity runtime inputs

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): bind parity runtime evidence

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): record compute boundary parity

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): disposition cross-cutting parity lanes

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(packaging): preflight gateway config upgrades

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): preserve rebase integration guarantees

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(ci): isolate temporary git signing config

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): update remaining schema v2 consumers

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(ci): provide e2fs tools to VM tests

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): align preflight with gateway startup

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(vm): preserve rootfs tar configuration

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* chore(config): adopt duration unit constructors

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(packaging): preflight RPM gateway config

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): address driver review findings

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(e2e): require fresh semantic parity evidence

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(docker): update tests for renamed sandbox label

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(gateway): preserve selective driver coverage after rebase

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
2026-09-11 05:00:24 +00:00
John T. Myers f4dc6be4b2 refactor(inference): remove managed inference routes (#3195)
* refactor(inference): remove managed inference routes

Closes #3172

Remove the inference route control plane, inference.local data path, built-in router crate, and SDK surface. Move inference workloads to explicitly imported provider profiles and native endpoints, with migration cleanup and updated tests and documentation.

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(policy): preserve alternate upstream isolation

Restore the provider policy activation guard so legacy OpenAI and Anthropic providers configured for alternate base URLs do not grant egress to the built-in public vendor endpoints.

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-09-09 18:47:22 +00:00
Simon Scatton 48c449d8c8 chore(deps): replace ring with AWS-LC (#3243)
* chore(deps): replace ring with AWS-LC

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* fix(lint): address warnings after dependency upgrades

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* fix(tls): limit provider initialization to reqwest clients

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

---------

Signed-off-by: Simon Scatton <sscatton@nvidia.com>
2026-09-09 17:57:42 +00:00
Evie Howard 6e6b3c8905 refactor(cli): remove local Dockerfile image builds (#3214)
Signed-off-by: Evie Howard <evhoward@redhat.com>
2026-09-09 13:28:17 +00:00
Jorge 519e5eb35f feat(e2e): make e2e:kubernetes work transparently on OpenShift (#3183)
* feat(e2e): make e2e:kubernetes work transparently on OpenShift

Running `mise run e2e:kubernetes` on OpenShift required manual namespace
creation, SCC grants, Helm value overrides, and cleanup. A separate
`e2e:openshift` task existed but only checked pod readiness without
running the Rust e2e test suite, and even with the suite wired up the
SSH-relay `sandbox connect` path stalled to the ready timeout because
`kubectl port-forward` cannot carry round-trip-heavy SSH over the
internet.

The harness now auto-detects OpenShift via the `route.openshift.io` API
group and, on OpenShift, both configures the cluster and switches the
gateway transport automatically:

- Drives the gateway through a passthrough OpenShift Route secured with
  mandatory mTLS instead of port-forward, so the connect suites
  (live_policy_update, port_forward, sync, connect-based
  sandbox_lifecycle, settings_management) actually pass. Computes the
  Route host from the cluster ingress domain, extracts client mTLS
  material from the openshell-client-tls secret, waits for the Route to
  serve mTLS, asserts a certless caller is rejected at the TLS
  handshake, and registers an mTLS CLI gateway pointing at the Route.
- Applies an SCC-compatible Helm values overlay that removes hardcoded
  runAsUser/fsGroup, letting OpenShift assign UIDs from the namespace
  range.
- Grants the privileged SCC to openshell-sandbox before Helm install
  and removes it during cleanup.
- Grants the anyuid SCC to the PostgreSQL fixture service account in
  DB scenarios and removes it during cleanup.
- All oc commands use --context to target the correct cluster.

The OpenShift e2e overlay (ci/values-openshift-e2e.yaml) turns TLS back
on, enables the Route, promotes the cert-verified caller to a dev
principal, and forces `image.pullPolicy`/`supervisor.image.pullPolicy`
to Always so runs against the `latest` upstream image use it instead of
a stale copy cached on the cluster nodes. Every OpenShift branch is
gated on OPENSHIFT_DETECTED, so the vanilla-Kubernetes port-forward path
is unchanged.

The Helm template for podSecurityContext is wrapped with {{- with }} so
null values omit the block instead of rendering invalid YAML.

The separate e2e:openshift task and e2e-openshift.sh script are removed
since e2e:kubernetes now covers OpenShift.

TESTING.md is updated with Kubernetes e2e documentation including
OpenShift auto-detection, dropping the e2e-host-gateway feature on
remote clusters, pinning IMAGE_TAG when the CLI and image versions
differ, task variants, and environment variables.

The debug-openshell-cluster skill gains an OpenShift platform row and
two SCC failure patterns (gateway rejected over hardcoded runAsUser,
sandbox missing the privileged SCC) covering the SCC handling and
podSecurityContext behavior this change introduces.

Signed-off-by: Jorge Garcia Oncins <jgarciao@redhat.com>

* fix(e2e): harden OpenShift SCC cleanup, mTLS gate, and Route timeout

Track the anyuid SCC grant for the PostgreSQL fixture with a dedicated
OPENSHIFT_POSTGRES_SCC_GRANTED flag set before the fixture apply, so a
failed apply no longer leaks the binding; cleanup now revokes it whenever
the grant succeeded, independent of deploy state.

Validate the Route server cert in the certless security gate (curl
--cacert instead of -k) and classify curl's exit code so only a TLS
client-auth rejection (35/56) counts as the expected certless rejection;
an unrelated DNS/timeout/TLS failure now fails loudly instead of masking
a potential mTLS hole.

Raise the OpenShift Route timeout in the e2e overlay. The default HAProxy
Route timeout is 30s, which severed long-lived transfers (large sandbox
upload/download, SSH-relay `sandbox connect`) mid-stream and failed the
sync e2e tests. Set both haproxy.router.openshift.io/timeout and
timeout-tunnel to 300s: a passthrough Route proxies in TCP mode, so
timeout-tunnel governs the established tunnel while timeout covers the
pre-tunnel phase.

Document the OpenShift transport exception, oc prerequisites and SCC
grants, and make the skopeo tag-check example copy-safe in TESTING.md.

Signed-off-by: Jorge Garcia Oncins <jgarciao@redhat.com>

---------

Signed-off-by: Jorge Garcia Oncins <jgarciao@redhat.com>
2026-09-08 20:37:19 +00:00
Jorge 457f5dfae7 fix(helm): omit podSecurityContext block when value is null (#3034)
The gateway pod template rendered `securityContext:` unconditionally, so
setting `podSecurityContext: null` (e.g. to let OpenShift's SCC assign the
UID/GID range) produced `securityContext: null` instead of omitting the
block. Wrap the block in `{{- with .Values.podSecurityContext }}` so a null
value omits it and an explicit value renders unchanged.

Add a helm-unittest suite covering the default, explicit, and null cases.

Fixes #3033

Signed-off-by: Jorge Garcia Oncins <jgarciao@redhat.com>
2026-09-08 20:31:04 +00:00
Yuedong Wu 7cc9551677 feat(server): support EC and EdDSA keys in OIDC JWKS validation (#2593)
Signed-off-by: Yuedong Wu <dwcn22@outlook.com>
2026-09-04 00:07:29 +00:00
Dhiraj Bokde cc4ded2088 feat(helm): split gateway and workspace charts (#2643)
* feat(helm): split gateway and workspace charts

Signed-off-by: Dhiraj Bokde <dbokde@nvidia.com>

* fix(helm): preserve split chart upgrade compatibility

Keep workspace manifests valid after value validation and default legacy reused values to the combined resource topology.

* fix(ci): preserve VM runtime for E2E

The Rust cache restores target/ after VM runtime artifacts are staged,
overwriting target/vm-runtime-compressed before openshell-driver-vm is built.
Stage the compressed runtime outside target and pass that location through
OPENSHELL_VM_RUNTIME_COMPRESSED_DIR so build.rs can embed the supervisor.

Also locate the Helm split-ownership test repository root from the script
path rather than git rev-parse. The test runs in a container where the
GitHub checkout can be owned by a different UID and rejected as dubious
ownership.

Signed-off-by: Dhiraj Bokde <dbokde@nvidia.com>

* fix(ci): install yq for Helm ownership test

The split-chart ownership regression uses yq to inspect rendered YAML,
but the Helm CI container installs only tools declared in mise.
Declare and lock yq so mise install --locked provides the test dependency.

Signed-off-by: Dhiraj Bokde <dbokde@nvidia.com>

---------

Signed-off-by: Dhiraj Bokde <dbokde@nvidia.com>
2026-09-02 00:23:39 +00:00
Drew Newberry 9ca19e6c80 refactor(compute): decouple gateway driver composition (#2823)
* refactor(compute): decouple gateway driver composition

Move first-party composition and VM process ownership into openshell-gateway, leaving openshell-server backend-independent. Update packaging and build references with the new crate, simplify the compiled-driver boundary, and keep the driver-free gateway path buildable with bundled Z3 tooling.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(telemetry): bound compute driver categories

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(core): keep runtime transport generic

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(compute): complete server driver decoupling

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(compute): preserve driver integrations after rebase

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>

* fix(compute): preserve docker tracing after decoupling

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>

* fix(compute): preserve driver behavior after extraction

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(compute): remove MXC policy side channel

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(compute): separate policy delivery from readiness

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>
2026-09-01 21:13:45 +00:00
Yuedong Wu e508c169ea fix(helm): honor empty clientCaSecretName for HTTPS-only mode (#2235)
* fix(helm): honor empty clientCaSecretName for HTTPS-only mode

Signed-off-by: Yuedong Wu <dwcn22@outlook.com>

* docs(skills): document HTTPS-only clientCaSecretName in debug-openshell-cluster

Signed-off-by: Yuedong Wu <dwcn22@outlook.com>

---------

Signed-off-by: Yuedong Wu <dwcn22@outlook.com>
2026-09-01 18:10:40 +00:00
Drew Newberry 9b6d904e88 feat(compute): delegate sandbox authentication to drivers (#2968)
* feat(compute): delegate sandbox authentication to drivers

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>

* fix(kubernetes): align sandbox identity annotation

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>

* test(auth): restore sandbox bootstrap coverage

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>

* fix(kubernetes): satisfy ownership test lint

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>

---------

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>
2026-08-31 17:41:56 +00:00
krishicks f68867b869 feat(gateway): identify gateways in exported traces (#2647)
* feat(gateway): add installation name configuration

Add a first-class operator-assigned gateway name with TOML, CLI, environment,
and Helm configuration surfaces. Local gateways default to openshell, while
Helm defaults to the chart fullname; operators sharing a collector across
namespaces or clusters can set a globally distinct name.

Signed-off-by: Kris Hicks <khicks@nvidia.com>

* feat(gateway): identify gateways in exported traces

Attach the configured gateway installation name and compute driver to the
gateway OpenTelemetry resource so operators can filter traces from multiple
installations that share a collector. Forward the gateway name and OTLP
endpoint to managed external drivers so their distinct service resources carry
the same installation identity.

Keep service.name stable per process type, omit blank resource values, and
leave per-span operation names and request attributes unchanged.

Refs #2507

Signed-off-by: Kris Hicks <khicks@nvidia.com>

---------

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-08-27 17:27:34 +00:00
Simon Scatton 981606d2f8 ci: build release binaries with Nix (#2977)
* ci: build release binaries with Nix

Refs #1683

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* ci: build VM artifacts with Nix

Refs #1683

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* ci: build images from Nix artifacts

Refs #1683

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* fix(nix): prevent host header leakage

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* ci: parallelize artifact builds

Refs #1683

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* ci: build external driver test artifacts

Refs #1683

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* fix(nix): disable mold in musl shells

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* ci: key Rust cache by Nix shell derivation

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* ci: refactor end-to-end workflows

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* ci: split platform binary workflows

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* ci: remove obsolete native build workflows

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* ci: replace disallowed mise action

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* ci: fix refactored e2e lanes

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* ci: check out local result action

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* ci: cache mise installations

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* ci: run docker builds on host runners

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* ci: disable unstable kubernetes e2e lanes

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* fix(ci): scope binary builds to cargo packages

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* fix(ci): address zizmor template injection findings

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* fix(ci): resolve remaining zizmor annotations

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

---------

Signed-off-by: Simon Scatton <sscatton@nvidia.com>
2026-08-27 17:25:06 +00:00
krishicks d0dfb22baf feat(kubernetes): export driver traces over OTLP (#2958)
Mirror the VM, Podman, and Docker driver tracing setup for Kubernetes.
Export standalone driver spans through OTLP/gRPC as the distinct
openshell-driver-kubernetes service, preserve gateway trace context, record
lifecycle operations and gRPC failures, and flush spans on shutdown.

Kubernetes currently runs in-process when selected as a built-in gateway
driver. Use the temporary server-boundary shim shared with Podman and Docker
so traces retain the shape they will have when Kubernetes moves to a
separate process. Move the common ComputeDriver RPC tracing layer into
openshell-otel to keep all drivers aligned.

Propagate the active W3C context through the controller-reserved Sandbox
annotation and enable Agent Sandbox OTLP export in the local k3s workflow.
This connects asynchronous controller reconciliation spans to the originating
OpenShell create trace.

Expose gateway OTLP configuration through Helm and add an Aspire collector
to the local k3s workflow. Extend helm:k3s:forward with OTLP ingest and trace
UI forwarding for Kubernetes and local container gateway development.

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-08-26 21:22:51 +00:00
Simon Scatton 455883905a fix(python): remove CLI from wheel (#2321)
The Maturin-based wheel packaging was a historical remnant from when the local gateway launch path and OpenShell CLI were coupled in one binary. The gateway and CLI now ship as standalone artifacts, so the Python distribution should contain only the SDK.

Build a single platform-independent setuptools wheel, verify that it cannot contain native code or an openshell entry point, and simplify the release jobs and documentation for SDK-only PyPI installs.

Signed-off-by: Simon Scatton <sscatton@nvidia.com>
2026-08-25 13:20:42 +00:00
grs 40d1b48666 feat(provider): support for SPIFFE backed token exchange (#1970)
* feat(provider): add ability to request token exchange instead of client credentials as OAuth grant_type

Signed-off-by: Gordon Sim <gsim@redhat.com>

* test(proxy): add further tests for token exchange

Signed-off-by: Gordon Sim <gsim@redhat.com>

* test(provider): add runnable example for token exchange

Signed-off-by: Gordon Sim <gsim@redhat.com>

* test(e2e): cover Podman token exchange grants

Signed-off-by: Gordon Sim <gsim@redhat.com>

* refactor(oauth): extract duplicated functionality from server and supervisor

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(provider): evict nearest-to-expiry entry from intermediate token cache

Signed-off-by: Gordon Sim <gsim@redhat.com>

* doc(supervisor): add podman example for token exchange

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(provider): withhold token-exchange subject credentials

Signed-off-by: Gordon Sim <gsim@redhat.com>

---------

Signed-off-by: Gordon Sim <gsim@redhat.com>
2026-08-24 05:42:30 +00:00
Evan Lezar 3be2cd8a29 fix(helm): preflight Agent Sandbox APIs (#2867)
* fix(helm): preflight Agent Sandbox APIs

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(kubernetes): share Agent Sandbox setup

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(e2e): wait for Agent Sandbox CRD status

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* ci(canary): sparse-checkout sandbox helper

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-08-21 14:32:19 +00:00