40 Commits
Author SHA1 Message Date
Evan Lezar cb193ef1c2 ci: use package installers consistently in integration tests (#4056)
* test(tmachine): add Fedora RPM package installer

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* ci: qualify Ubuntu branch installs with DEB packages

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* ci: align package installers across integration matrices

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-10-01 15:05:35 +00:00
Evan Lezar 82e889374f test(tmachine): add Fedora RPM package installer (#4025)
Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-10-01 14:29:09 +00:00
Simon ScattonandMrunal Patel 5698c4f746 fix(ci): qualify protobuf compatibility by release train (#4049)
* feat(ci): detect breaking protobuf changes

Compare the proto module against the PR or merge-group base and report Buf violations in Branch Checks. Add local reproduction and fixture coverage.

Closes #3794

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(ci): pin protobuf check container image

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(ci): qualify protobuf compatibility by release train

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* refactor(ci): reuse protobuf compatibility action

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* refactor(ci): run protobuf checks as a Nix app with one ref

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

---------

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
Signed-off-by: Simon Scatton <sscatton@nvidia.com>
Co-authored-by: Mrunal Patel <mrunalp@gmail.com>
2026-10-01 14:00:45 +00:00
Simon Scatton fde79f1aa6 fix(ci): align integration inputs with release candidate source (#4048)
Signed-off-by: Simon Scatton <sscatton@nvidia.com>
2026-10-01 13:01:20 +00:00
Simon Scatton 21fea95935 test(tmachine): add K3s conformance scenario (#3848)
* test(tmachine): add K3s conformance scenario

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* refactor(tmachine): use Helm values file for K3s installer

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* ci(tmachine): run K3s conformance in integration jobs

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* ci(tmachine): verify installer scripts and document version baseline

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

---------

Signed-off-by: Simon Scatton <sscatton@nvidia.com>
2026-09-30 14:49:09 +00:00
Piotr Mlocek c0eb3dbd30 fix(ci): restore repository permission vetters (#3875)
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-09-29 18:05:21 +00:00
Piotr Mlocek aead95b7ab fix(policy): propose rules for unknown DNS hosts (#3707)
* fix(policy): propose rules for unknown DNS hosts

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(policy): clarify synthetic DNS use across protocols

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(policy): harden unknown-host DNS observations

- Emit the policy_dns_ineligible denial for every unknown name and
  report observation staging failures as DNS failure events.
- Refuse unknown names during fail-closed quarantine and after the
  observation budget, now a quarter of each address family's pool.
- Pin transparent TCP to the mapping of the deciding policy generation
  so a reload between DNS and authorization fails closed.
- Stop Docker workloads from inheriting host DNS search domains, which
  let the first expanded short name claim an observation address.
- Share mechanistic draft polling in conformance, register
  new-hostname-proposal in the installed suite, and update docs.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* test(policy): build policy DNS proxy tests on every target

The proxy tests name PolicyEndpointId, which proxy.rs imported only on
Linux, so the macOS test build failed. Import it for test builds too.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(policy): name DNS queries and mapped hosts in OCSF denials

DNS denial and failure events attached port 53 to the queried name,
which read as a connection to that host. They now carry only the name.
Transparent TCP denials for a policy DNS address show the mapped
hostname and keep the synthetic address in dst_endpoint.ip.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

---------

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-09-25 08:28:41 +00:00
Piotr Mlocek c9257c8447 fix(policy): restore policy.local and proposal conformance (#3689)
* fix(policy): restore policy.local and proposal conformance

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* test(conformance): select policy scenarios by name

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* test(conformance): reduce policy scenario timing flakes

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* test(conformance): assert proposals target Bash

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

---------

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-09-24 21:03:26 -07:00
Jim Meyer 48725c5fcc ci(release): move CodeQL, Trivy, and Zizmor to advisory (#3693)
* ci(release): Move CodeQL, Trivy, and Zizmor to advisory

* docs(ci): describe advisory static findings for tagged releases

Signed-off-by: Jim Meyer <jimeyer@nvidia.com>

---------

Signed-off-by: Jim Meyer <jimeyer@nvidia.com>
2026-09-25 01:26:59 +00:00
Jim Meyer a00ea31c66 ci: restrict copy-pr-bot manual vetters (#3678)
Signed-off-by: Jim Meyer <jimeyer@nvidia.com>
2026-09-24 22:42:51 +00:00
Oliver Calder 6c864ec9ab fix(install): avoid installing incompatible docker snap (#3666)
* fix(install): avoid installing incompatible docker snap

The work to land RFC-0012 added new restrictions when interacting with
Docker by setting `NoNewPrivs`. This prevents the `docker` snap from
transitioning its AppArmor profile from `snap.docker.dockerd` to
`docker-default` when it tries to launch a container. Thus, the `docker`
snap is currently incompatible with OpenShell.

This commit prevents `install.sh` from installing the `docker` snap
before installing the `openshell` snap, and instead requires the user to
install a non-snap Docker daemon before proceeding with installing the
snap. Systems without the `snap` command are unaffected, since they
install native packages without checking for the presence of Docker or
other compute providers.

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* fix(snap): only snapd 2.76 for openshell snap since store installs work

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* fixup! fix(install): avoid installing incompatible docker snap

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* fixup! fix(snap): only snapd 2.76 for openshell snap since store installs work

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

---------

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>
2026-09-24 19:46:00 +00:00
Oliver Calder e60098d748 fix(snap): install openshell snap via install.sh when snap available (#3656)
* fix(snap): update stale snap docs and tests

Previously, the `openshell` snap required the `docker` snap. Now, it
works with any Docker daemon running on the system. Furthermore, the
`snap-declaration` assertion on the `openshell` snap when installed from
the Snap Store causes the `openshell` snap to always connect to the
system `:docker` slot, rather than a slot provided by the `docker` snap.

This commit updates the documentation, including the `description` field in
`snapcraft.yaml`, to ensure that all information is correct and
up-to-date.

Additionally, some tests connected the `openshell:docker` plug to the
`docker` snap's `docker:docker-daemon` slot, which is inconsistent with
how the `openshell` snap operates when installed from the store. Update
those tests to connect to the system `:docker` slot as well.

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* fix(snap): require snapd 2.76 for openshell snap

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* fix(nix): align snap gateway reproducer timeout with release-canary

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* fix(snap): require snapd 2.77 for openshell snap

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* fixup! fix(snap): require snapd 2.77 for openshell snap

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* feat(snap): install openshell snap via install.sh when snap available

Change `install.sh` to install the `openshell` snap by default when
snapd is installed on the host. This installs the snap from the
`latest/stable` channel, which should match the most up-to-date release
tag on github.

If `OPENSHELL_VERSION=dev` is set for `install.sh`, then it will install
the `openshell` snap from the `latest/edge` channel, which matches the
latest dev release available on github.

The `openshell` snap currently requires Docker in order to function. If
a Docker daemon is already installed on the system, it will be used by
the `openshell` snap. Otherwise, `install.sh` will install the `docker`
snap first, wait for the Docker daemon to be ready, and then install the
`openshell` snap.

Also, update the `release-canary.yml` to split the `ubuntu-snap` job
into `ubuntu-snap-system-docker` and `ubuntu-snap-provisions-docker`,
which test the two aforementioned scenarios. Previously, `ubuntu-snap`
manually installed a given snap artifact as built from CI, but with
these new jobs, it instead uses `install.sh` to install the published
`openshell` snap from the `latest/edge` track, thus matching the
behavior of the other release canary jobs.

Make corresponding changes to the `nix` guest reproducer, documentation,
and tests.

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* fix(snap): configure local gateway authentication

The `openshell` snap runs the gateway as a systemd system service, which
runs as root. Thus, the mTLS certs are generated by root and stored in a
root-owned directory to which non-root users do not have access. For
this reason, the snap's `openshell-gateway-wrapper` script sets
`OPENSHELL_DISABLE_TLS=true`.

This commit ensures that the `openshell` snap's gateway allows
unauthenticated local access by writing a default `gateway.toml`
configuration file during the install hook, which runs after the snap is
first installed but before services are started. The config file
contains sets `allow_unauthenticated_users = true`.

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* fix(install): configure snap gateway authentication

Recently, a new install hook was added which writes a default config
file for the `openshell` snap to allow unauthenticated local access to
the gateway. This is because the gateway service runs as root and the
mTLS certificates are not accessible to non-root users.

However, the `install.sh` script installs the `openshell` snap from the
snap store, and the published version may not yet have that new install
hook. Or, the user may already have the snap installed, in which case
the install hook does not run. In either case, we need `install.sh` to
ensure that the config file is written to set the gateway auth to
`allow_unauthenticated_users = true`, and then restart the openshell
gateway service.

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* fixup! feat(snap): install openshell snap via install.sh when snap available

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

---------

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>
2026-09-24 15:10:08 +00:00
Evan Lezar 907f894ebc ci(release): publish prereleases with qualification summary (#3593)
Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-09-23 13:18:30 +00:00
Evan Lezar 3107ff1f82 ci(security): stage release finding enforcement (#3552)
Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-09-22 17:21:23 +00:00
Evan Lezar 251f77e2b8 ci(security): gate tagged releases on scans (#3523)
Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-09-21 18:51:37 +00:00
Simon Scatton fc03bffead ci: consolidate multi-platform image builds (#3408)
* ci: remove release smoke tests

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* ci: consolidate image builds

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

---------

Signed-off-by: Simon Scatton <sscatton@nvidia.com>
2026-09-17 13:46:31 +00:00
Drew Newberry c1f2e7189f feat(isolation): implement the RFC 0012 sandbox architecture (#2942)
* feat(isolation): add RFC 0012 backend contract

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>

* refactor(isolation): name the interface crate explicitly

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): expose trusted host gateway

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(agents): inventory the MXC driver

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): add mediated DNS transport

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): tighten interface error and digest contracts

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(isolation): remove unrelated driver inventory

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): define capability-free launch contract

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): seal confirmed boundary state

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): validate confirmation for external backend implementations

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): clarify mediated DNS identity

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(isolation): generalize loopback connector

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(isolation): unify typed network mediation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): bind launches to sandbox sessions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(mxc): initialize extended sandbox status

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): add boundary protocol and Linux primitives

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): harden signals and separate process status from transport

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): validate remote confirmation through public contract

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): validate wire state and propagate snapshot failures

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(isolation): import owned agent specification explicitly

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(isolation): describe mediated DNS channel

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): bound mediation attach without nested retries

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(isolation): generalize loopback protocol

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): add transport-neutral session authentication

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(isolation): separate sandbox backend protocol

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): harden runtime boundary controls

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): add terminal boundary operation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): split supervisor and sandbox runtimes

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): harden boundary isolation and lifecycle ownership

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): reject private root redirects and adopt typed errors

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): preserve accept thread ownership on musl

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(sandbox): isolate credential probes from filtered threads

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): return retained exec exit status to independent waiters

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): bound network mediation and preserve socket authorization

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): bound control admission and retire stale mediation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci(e2e): select migrated drivers per stack layer

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(sandbox): implement loopback connector

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): authenticate the Sandbox Protocol

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(supervisor): rotate launch-scoped authentication

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(sandbox): consume dedicated backend crate

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(sandbox): align topology session fixture

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): align projected bootstrap bundle

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(auth): validate refreshed credentials before rotation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): fail closed across supervisor disconnects

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): repair rebased sandbox CI

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* build(runtime): publish separate sandbox and supervisor images

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(config): configure the sandbox runtime image

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(ci): validate sandbox binary linkage

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(isolation): use backend and runtime terminology

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(sandbox): use a scratch runtime image

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(ci): refresh schema and dependency policy

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): bind reconnects to supervisor process

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs: align runtime split operational guidance

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* chore(security): document Kubernetes runtime RBAC

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): enforce runtime lifecycle invariants

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(compute): identify sandbox start generations

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(server): restore sandbox launch sessions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): support authenticated runtime replacement

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(auth): bind sandbox session successors

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(auth): retry pending sandbox successors

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(vm): run the supervisor outside the guest workload

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(vm): use sandbox backend protocol

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): repair rebase integration

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): use unified build toolchain

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(vm): use sandbox runtime terminology

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(vm): own guest network bootstrap

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): expose guest init version

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): select native supervisor artifacts

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): guard guest init Linux symbols

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): scope Linux test imports

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): avoid guest interface casts

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): reconcile admitted sandbox identity

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): share resolved sandbox identity

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): surface host supervisor failures

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): include guest logs on supervisor exit

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): rotate and clean runtime generations

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): make sandbox starts generation-aware

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): rotate restored sandbox sessions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(vm): keep shared paths in the base layer

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): bind sandbox session lineage

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(docker): isolate workloads behind the host supervisor

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(docker): rotate launch-scoped authentication

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(docker): use sandbox backend protocol

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(docker): use host networking for supervisor

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): preserve host gateway alias resolution

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(docker): use separate sandbox and supervisor images

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): restore startup validation after rebase

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(docker): name the sandbox runtime directly

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): narrow supervisor CA runtime storage

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): close companion isolation gaps

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(docker): align mediated network expectations

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(docker): exercise mediated network paths

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): attach supervisor to managed network

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): defer supervisor recovery until gateway is ready

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): make sandbox starts generation-aware

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): rotate restored sandbox sessions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): preserve workloads during session rotation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(docker): remove unrelated configuration RFC changes

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): bind sandbox session lineage

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(kubernetes): add proxy-pod isolation topology

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(kubernetes): use sandbox backend protocol

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): use stable sandbox service authority

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(kubernetes): split sandbox and supervisor images

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): adapt proxy pods to current runtime APIs

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(kubernetes): describe the single runtime placement

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(kubernetes): simplify sandbox orchestration

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): validate deployment prerequisites

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): update Trivy Helm profile inventory

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(kubernetes): update Trivy scan inventory count

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): reuse preloaded runtime images in e2e

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): type and clean runtime resources

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): make sandbox restarts recoverable

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): rotate restored sandbox sessions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): preserve supervisor egress

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): bind sandbox session lineage

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(podman): adopt isolated sandbox and supervisor containers

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): stage bootstrap archives at named volume destinations

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(podman): rotate launch-scoped authentication

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(podman): use sandbox backend protocol

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(podman): use host networking for supervisor

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(podman): split sandbox and supervisor images

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): repair rebase integration

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(podman): name the sandbox runtime directly

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): provision supervisor CA runtime storage

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): address isolation review findings

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): inspect Debian supervisor provenance

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): use libpod-compatible tmpfs options

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): bind verified sandbox runtime binary

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): provide external driver data directory

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): start sandbox before joining user namespace

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): separate supervisor user namespace

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): make sandbox starts generation-aware

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): rotate restored sandbox sessions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): bind sandbox session lineage

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* perf(isolation): add TCP and DNS benchmark harnesses

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(perf): align benchmark timing and supported protocols

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(perf): report TCP benchmark metrics accurately

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(perf): cancel failed worker startup

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): build matching local supervisor image

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): make local sandbox smoke test runnable

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): wire local sandbox runtime image

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): narrow sandbox service RBAC

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(ci): validate split runtime artifacts

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): harden runtime session handling

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(supervisor): add standalone network proxy role

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(rfc): remove implementation companion notes

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(vm): standardize runtime release name

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): pin renamed runtime artifacts

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(auth): persist sandbox runtime identity

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(runtime): restore branch validation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): reconcile main after rebase

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(network): close unframed HTTP 1.0 responses

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* chore(isolation): preserve upstream OCSF updates

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(security): close credential and TLS replay paths

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(auth): make sandbox refresh retries idempotent

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-16 00:49:14 +00:00
alangou d99f12a33b fix(ci): align Trivy change detection and scan baselines (#3277)
Signed-off-by: Adrien Langou <alangou@nvidia.com>
2026-09-11 10:55:52 +00:00
alangou 3eb81beba9 ci(security): orchestrate security scans with severity gating (#3255)
Signed-off-by: Adrien Langou <alangou@nvidia.com>
2026-09-11 10:26:56 +00:00
Piotr Mlocek 0569c3a20a ci(windows): make PR checks opt-in and main jobs advisory (#3268)
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-09-11 01:29:34 +00:00
alangou 3693b32841 ci(trivy): add artifact and PR configuration scans (#3185)
* ci(trivy): add artifact and PR configuration scans

Signed-off-by: Adrien Langou <alangou@nvidia.com>

* fix(ci): harden Trivy gate detection and finding diff

Signed-off-by: Adrien Langou <alangou@nvidia.com>

* feat(ci): scan released artifacts in release pipelines

Signed-off-by: Adrien Langou <alangou@nvidia.com>

* fix(ci): harden and simplify Trivy scans

Signed-off-by: Adrien Langou <alangou@nvidia.com>

* fix(ci): consolidate Trivy reports and prevent collisions

Signed-off-by: Adrien Langou <alangou@nvidia.com>

---------

Signed-off-by: Adrien Langou <alangou@nvidia.com>
2026-09-09 13:58:06 +00:00
alangou f7180c0fd6 feat(ci): add Codex Security release qualification (#3087)
* feat(ci): add Codex Security release qualification

Scan cumulative release-train diffs through NVIDIA inference and publish findings to Code Scanning.

Signed-off-by: alangou <alangou@nvidia.com>

* fix(ci): disable package cache for security scan

Prevent cache poisoning in the tag-triggered Codex Security workflow.

Signed-off-by: alangou <alangou@nvidia.com>

* refactor(ci): simplify Codex Security reporting

Remove custom inference cost accounting so the workflow remains focused on scanning and SARIF publication.

Signed-off-by: alangou <alangou@nvidia.com>

---------

Signed-off-by: alangou <alangou@nvidia.com>
2026-09-01 13:57:14 +00:00
alangou 4c9437b63e ci(codeql): run nightly scans on main (#3007)
Signed-off-by: Adrien Langou <alangou@nvidia.com>
2026-08-28 13:26:39 +00:00
Simon Scatton 981606d2f8 ci: build release binaries with Nix (#2977)
* ci: build release binaries with Nix

Refs #1683

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* ci: build VM artifacts with Nix

Refs #1683

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* ci: build images from Nix artifacts

Refs #1683

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* fix(nix): prevent host header leakage

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* ci: parallelize artifact builds

Refs #1683

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* ci: build external driver test artifacts

Refs #1683

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* fix(nix): disable mold in musl shells

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* ci: key Rust cache by Nix shell derivation

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* ci: refactor end-to-end workflows

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* ci: split platform binary workflows

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* ci: remove obsolete native build workflows

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* ci: replace disallowed mise action

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* ci: fix refactored e2e lanes

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* ci: check out local result action

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* ci: cache mise installations

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* ci: run docker builds on host runners

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* ci: disable unstable kubernetes e2e lanes

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* fix(ci): scope binary builds to cargo packages

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* fix(ci): address zizmor template injection findings

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* fix(ci): resolve remaining zizmor annotations

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

---------

Signed-off-by: Simon Scatton <sscatton@nvidia.com>
2026-08-27 17:25:06 +00:00
alangou 5f90c8579c ci(security): add informational security checks (#2930)
Signed-off-by: Adrien Langou <alangou@nvidia.com>
2026-08-27 13:32:16 +00:00
Evan Lezar 9f88f8ff9b ci: remove rootless podman e2e lane (#2981)
Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-08-27 08:48:44 +00:00
5548405fcb feat(credentials): add provider credential storage drivers (#2437)
* feat(credentials): add provider credential storage drivers

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(credentials): harden credential update handling

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(credentials): harden credential driver security, correctness, and performance

Address review findings from the credential storage drivers PR:

- Route additional_credentials through the driver on refresh to prevent
  silent data loss for multi-credential providers (e.g. AWS STS)
- Clean up stored credential handles on CAS failure during refresh to
  prevent orphaned secrets in external backends
- Enforce namespace validation in the Kubernetes Secrets driver to
  prevent cross-namespace credential access when allow_reference_namespace
  is not enabled
- Cache Vault Kubernetes auth tokens with 80% TTL to avoid re-authenticating
  on every credential operation
- Parallelize resolve_credentials in all three drivers using try_join_all
  for faster sandbox startup
- Add existingSecret support for the KEK Secret to fix helm template/GitOps
  workflows where lookup returns empty and regenerates the key
- Document RBAC blast radius for the Kubernetes Secrets credential driver
  and recommend a dedicated namespace

Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com>
Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>

* fix(credentials): add optimistic concurrency, fix thundering herd, parallelize operations

Use resourceVersion optimistic concurrency with retry loop for K8s
Secret ownership checks to prevent TOCTOU races. Switch Vault token
cache from RwLock to Mutex with double-check pattern to prevent
thundering herd on cache miss. Parallelize credential store and delete
operations across independent keys using try_join_all.

Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com>
Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>

* fix(credentials): handle partial failures, add delete retry, consolidate cleanup

Replace try_join_all with join_all in credential store/delete operations
to handle partial failures — successfully-stored handles are cleaned up
when another key fails. Add retry loop with conflict detection to
db-credstore delete_credential, matching the K8s driver pattern.
Consolidate 4 manual cleanup_pre_stored_provider_credentials call sites
into a single error handler using an async block. Remove inconsistent
.trim() from db-credstore validate_handle_owner.

Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com>
Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>

* fix(credentials): fix retry loop guard and remove unprotected validation

Remove attempt-count guard from 409/Aborted match arms in retry loops
so the post-loop Status::aborted error is reachable after exhausting
retries. Previously, last-attempt conflicts fell through to the
catch-all error arm, producing misleading Status::unavailable errors.

Remove duplicate validation calls that ran after
prepare_provider_credential_update but outside the cleanup-protected
async block, which would leak pre-stored handles on failure.

Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com>
Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>

* fix(credentials): add workspace/provider UUID to credential backend paths

Include workspace and provider ID in credential backend object paths to ensure
cross-workspace uniqueness and prevent credential collision (GATOR-1806c9be-01).

- Updated credential driver proto to include workspace and provider_id fields
- Modified Vault driver to include workspace/provider_id in managed_secret_path
- Modified Kubernetes Secrets driver to include workspace/provider_id in
  credential_owner_id and managed_secret_name
- Updated all credential runtime calls to pass workspace/provider_id
- Updated tests to use the new signatures

This prevents two workspaces sharing the same external credential store from
colliding on provider names, which was a critical security issue (CWE-639).

* fix(credentials): preserve provider-level expiration for handle-backed credentials

Compute effective expiration from both provider and driver values using the
earliest non-zero timestamp and skip expired values before insertion
(GATOR-1806c9be-02).

- Modified resolve_provider_handles to check provider credential_expires_at_ms
- Skip expired credentials during resolution instead of returning them
- Use effective expiration (min of provider and driver) in resolution results
- Fix inference.rs to preserve earliest expiration when merging

This ensures handle-backed credentials respect the same expiration semantics
as inline credentials.

* fix(credentials): stage refresh changes under new handles before validation

Stage credential replacements under new immutable handles instead of reusing
existing handles to prevent overwriting committed values before validation/CAS
(GATOR-1806c9be-03).

- Stage credentials with empty existing_handles map to force new handle creation
- Validate and CAS before the new values are committed to backend storage
- Delete old handles only after successful CAS
- On CAS failure, delete only the newly staged handles
- This prevents CWE-362/CWE-367 race conditions where failed refreshes could
  still modify or delete the active credential

The fix ensures that a rejected refresh cannot modify the backend object still
referenced by the committed provider record.

* fix(credentials): add timeouts to credential driver RPCs

Apply configured timeouts to both startup capability negotiation and runtime
RPCs to prevent indefinite hangs (GATOR-1806c9be-05).

- Add DEFAULT_CREDENTIAL_DRIVER_RPC_TIMEOUT_SECS constant (30s)
- Apply timeout to GetCapabilities during startup connection
- Apply timeout to all runtime RPCs (store, delete, resolve)
- Use tokio::time::timeout to bound the entire GetCapabilities operation
  during startup, not just the socket connection
- Return contextual deadline errors on timeout

This prevents a faulty or overloaded driver from hanging gateway operations
indefinitely.

* fix(credentials): fix test to use consistent workspace/provider identity

The Kubernetes auth Vault resolve test was constructing a managed path with
test-workspace/test-provider-id but sending default/prov-123 in the request,
causing validation to reject the request (GATOR-18e32351-01).

- Update test to use test-workspace and test-provider-id in the request to
  match the logical_path construction
- This ensures the test exercises the intended code path and validates
  Kubernetes auth resolution properly

The test now passes and correctly validates identity enforcement.

* fix(credentials): use unique staging ID for refresh to avoid overwrites

Stage refresh replacements under genuinely distinct immutable handles using
a unique staging ID to prevent overwriting committed values (GATOR-1806c9be-03).

- Generate a unique staging ID using UUID for each refresh operation
- Use this staging ID when storing credentials instead of the real provider ID
- Pass the same staging ID during cleanup on failure to delete only staged objects
- This ensures deterministic paths (Vault) and object names (K8s) don't collide
  with the committed provider's credentials

The fix prevents failed refreshes from silently replacing active credentials
or breaking providers by deleting still-referenced backend objects.

* fix(credentials): wrap credential driver RPCs in local timeouts

Add local tokio::time::timeout wrappers around credential driver RPCs to
bound non-compliant or stalled UDS peers (GATOR-1806c9be-05).

- Wrap StoreCredential, DeleteCredential, and ResolveCredentials in local timeouts
- Return contextual deadline_exceeded errors when timeouts occur
- Keep existing gRPC timeout metadata for compliant implementations
- GetCapabilities during startup was already wrapped in previous commit

This ensures a faulty local driver cannot hang gateway operations indefinitely,
even if it accepts the connection but never responds to the RPC.

* fix(credentials): preserve ownership for staged refreshes

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(credentials): bound startup capability probe

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* test(provider): authenticate credential handler requests

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(ci): grant actions read to credential driver e2e

Signed-off-by: Seth Jennings <sjenning@redhat.com>

---------

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com>
Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>
Signed-off-by: Seth Jennings <sjenning@redhat.com>
Co-authored-by: Taylor Mutch <taylormutch@gmail.com>
Co-authored-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>
2026-08-05 16:29:41 +00:00
Evan Lezar 339eae5ad1 ci(e2e): reuse prebuilt CLI and gateway artifacts (#2311)
* ci(e2e): reuse prebuilt CLI artifacts

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* ci(e2e): reuse prebuilt gateway artifacts

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* ci(e2e): reuse prebuilt VM driver artifact

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-07-20 16:24:02 +00:00
Drew Newberry 5402551797 test(e2e): run VM suite in CI (#2305)
* test(e2e): run VM suite in CI

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(ci): configure KVM permissions directly

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(e2e): flush VM overlay before restart

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs: simplify VM test documentation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(e2e): include gateway resume in VM run

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-07-16 16:46:39 -07:00
Evan Lezar b4be33e541 feat(ci): introduce merge queue (#2024)
* feat(ci): introduce merge queue

Closes #1946

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(ci): run GPU E2E for merge groups

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* docs(ci): clarify merge queue GPU gate

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-07-16 00:03:52 +02:00
Taylor Mutch 269dbc6d8a ci(kubernetes): add HA e2e workflow (#1598)
Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
2026-06-01 13:54:30 -04:00
Taylor Mutch 52389370b4 ci: deduplicate e2e workflows (#1512)
Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
2026-05-21 17:44:05 -07:00
Piotr Mlocek f8e3f9b3cd fix(ci): resolve mirror gate statuses for fork PRs (#1504)
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-05-21 11:25:07 -07:00
Taylor Mutch c600b11ff2 docs(agents): add release canary testing skill (#1440)
* ci(canary): add kind-based helm chart smoke test

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* docs(agents): add release canary testing skill

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

---------

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
2026-05-20 09:34:57 -07:00
Piotr Mlocek 2a5a44989a fix(ci): require PR checks to pass (#1461) 2026-05-19 14:46:23 -07:00
John T. Myers 8ace316f8d fix(ci): allowlist dependabot for DCO (#1202)
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-05-06 09:33:43 -07:00
jtoelke2 4803889cc7 ci: cut over non-release workflows to shared runners (#1131)
Signed-off-by: Jonas Toelke <jtoelke@nvidia.com>
2026-05-04 14:34:16 -05:00
Piotr Mlocek c49ae09d5b fix(ci): grant actions:read and contents:read to E2E label helper (#995) 2026-04-28 10:16:27 -07:00
Piotr Mlocek c4286648bb ci(e2e): replace label dispatcher with comment-only helper (#990) 2026-04-27 11:11:52 -07:00
Piotr Mlocek e703b597cd ci(e2e): add label dispatcher and contributor CI docs (#975) 2026-04-27 09:49:01 -07:00