Commit Graph
1622 Commits
Author SHA1 Message Date
Shiju ecf8d1973b test(vm): inherit durable identity restart fixture
Integrate the identity parent correction while preserving cleanup behavior.

Signed-off-by: Shiju <shiju@nvidia.com>
2026-10-03 07:52:03 +05:30
Shiju e62095dbd5 test(vm): flush identity fixture before restart
Persist the canonical identity file before readiness and report the observed
exec, canonical and file-owner identities before comparing them.

Signed-off-by: Shiju <shiju@nvidia.com>
2026-10-03 07:49:17 +05:30
Shiju 6ed7c1a9c6 fix(vm): limit preparation locking to cache publication
Allow independent image and private overlay workers to progress concurrently.
Serialize only destination checks and atomic cache publication, preserving
worker cancellation ownership. Integrate inactive identity restoration.

Signed-off-by: Shiju <shiju@nvidia.com>
2026-10-03 07:32:48 +05:30
Shiju fd1a801cfe fix(vm): restore inactive sandbox workload identity
Recover the persisted overlay owner before publishing stopped and terminal
sandboxes. Keep resources manageable when identity metadata is invalid.
Clarify fixed MicroVM ownership in policy-generation guidance.

Signed-off-by: Shiju <shiju@nvidia.com>
2026-10-03 07:21:56 +05:30
Shiju 2757ae7a35 fix(vm): inherit the supervisor startup call-site correction
Merge the corrected workload-identity parent into the cleanup branch.
Preserve the owned-worker cleanup changes while repairing inherited
supervisor compilation.

Signed-off-by: Shiju <shiju@nvidia.com>
2026-10-03 03:53:46 +05:30
Shiju 042bc23704 fix(supervisor): align VM identity startup with current APIs
Pass the optional rejection-log key for VM identity failures and keep
generic startup-write regressions free of VM identity constraints.

Repair the call sites after the branch rebase so the identity and cleanup
proposals compile against the current startup helpers.

Signed-off-by: Shiju <shiju@nvidia.com>
2026-10-03 03:50:17 +05:30
Shiju 16275912cb fix(vm): stop image workers before cleaning staging files
Run image preparation in an owned worker process, reserve its process
identity until cleanup completes, and protect staging with leases so
cancellation and recovery cannot race with another preparation attempt.

Signed-off-by: Shiju <shiju@nvidia.com>
2026-10-03 02:36:25 +05:30
Shiju 40e8e7d7a8 test(sandbox): clarify VM identity rejection fixtures
Name invalid user and group fixtures distinctly and move the final
workload identity into its group mismatch test.

Signed-off-by: Shiju <shiju@nvidia.com>
2026-10-03 02:36:24 +05:30
Shiju 036acbc112 fix(vm): enforce the configured workload identity
Reject conflicting policy users and groups before VM image preparation
and before guest attach or process startup changes state. Validate
supervisor policy updates against the protected VM workload identity.

Preserve the gateway CA transport and capability-free sandbox launcher.

Signed-off-by: Shiju <shiju@nvidia.com>
2026-10-03 02:36:24 +05:30
Drew Newberry ec49209da2 fix(providers): stabilize provider environment revisions (#4122)
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-10-02 18:42:14 +00:00
alangou 046fd2a024 ci: pin CI images by digest and add native architecture smoke checks (#4114)
* ci: pin CI images by digest and add native architecture smoke checks

Signed-off-by: Adrien Langou <alangou@nvidia.com>

* docs: resolve CI monitoring guidance conflict with main

Signed-off-by: Adrien Langou <alangou@nvidia.com>

---------

Signed-off-by: Adrien Langou <alangou@nvidia.com>
2026-10-02 16:22:27 +00:00
Florent BENOIT 8e9136bbc3 fix(vm): codesign macOS driver-vm with hypervisor entitlement in CI (#3507)
The release tarball ships an unsigned binary that fails at runtime
when Hypervisor.framework rejects the caller. Sign with the existing
entitlements plist during the build, before artifact upload.

Closes #3506

Signed-off-by: Florent Benoit <fbenoit@redhat.com>
2026-10-02 16:09:09 +00:00
Eric Curtin a48920ac04 feat(helm): add sandbox UID and GID values (#3947)
* feat(helm): add sandbox UID and GID values

Closes #2697

Signed-off-by: Eric Curtin <eric.curtin@docker.com>

* fix(helm): reject boolean sandbox UID and GID

Signed-off-by: Eric Curtin <eric.curtin@docker.com>

---------

Signed-off-by: Eric Curtin <eric.curtin@docker.com>
2026-10-02 15:36:29 +00:00
Philippe Martin f7273e48f6 fix(providers): restore supervisor-backed GCP metadata discovery (#3973)
* fix(providers): restore supervisor-backed GCP metadata discovery

Relay the reserved metadata endpoint to the supervisor and restore project, account, and placeholder token responses from live provider state. Cover Google SDK discovery and repeated refresh with provider E2E tests.

Closes #3860

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(providers): preserve default metadata account without an email

Use the default account identifier when the optional service account email is missing or empty. Cover repeated SDK refresh for missing, empty, and configured email values with distinct providers for parallel E2E execution.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* test(sandbox): fix metadata relay lint and timeout

Signed-off-by: Philippe Martin <phmartin@redhat.com>

---------

Signed-off-by: Philippe Martin <phmartin@redhat.com>
2026-10-02 15:33:37 +00:00
alangou 36819f476d fix(cli): stop uploads when Git filtering fails or selects no files (#3957)
Signed-off-by: Adrien Langou <alangou@nvidia.com>
2026-10-02 14:24:32 +00:00
alangou 88afd36de5 fix(deps): upgrade russh to address Dependabot alert 40 (#4116)
Signed-off-by: Adrien Langou <alangou@nvidia.com>
2026-10-02 14:12:25 +00:00
Matthew Grossman 5d6b3b8120 fix(kubernetes): serialize lifecycle cleanup with sandbox restart (#4078)
Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>
2026-10-02 12:25:46 +00:00
Oliver Calder 6e865df349 feat(snap): ship the standalone prover binary in the snap (#3717)
Signed-off-by: Oliver Calder <oliver.calder@canonical.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>
v0.1.3-pre.3
2026-10-02 11:51:56 +00:00
Matthew GrossmanandEvan Lezar 6048bed368 fix(ci): retry Nix shell and app dependency preparation (#4066)
* fix(ci): prepare Nix development shells in setup-nix

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(ci): retry Nix builds before executing apps once

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Evan Lezar <elezar@nvidia.com>
Co-authored-by: Evan Lezar <elezar@nvidia.com>
2026-10-02 11:27:20 +00:00
Drew Newberry 8719fc9f37 fix(sandbox): restrict provider file mode (#4093)
* fix(sandbox): restrict provider file mode (fixes #4091)

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(sandbox): set provider file mode with safe API

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-10-02 06:35:15 +00:00
Matthew Grossman 76cfd0e31d test(python): synchronize interactive exec TTY readiness (#4076)
* test(python): synchronize interactive exec TTY readiness

Wait for the complete readiness marker before streaming stdin so PTY echo cannot split the separately written TTY flags. Preserve pipe stream separation and verify consumed stdin and both output sentinels in TTY mode.

Fixes #4075

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* test(python): reuse interactive exec readiness marker

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

---------

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>
2026-10-01 21:47:40 +00:00
Drew Newberry 348a1fc625 feat(examples): run Jupyter notebooks in an OpenShell sandbox (#2253)
* feat(examples): add Jupyter sandbox fleet

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(examples): simplify Jupyter sandbox demo

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(examples): refresh Jupyter sandbox for current SDK

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(examples): simplify Jupyter sandbox demo

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(examples): execute notebooks on sandbox Jupyter kernel

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(examples): use CLI for Jupyter sandbox setup

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(examples): use published Jupyter image and gateway CLI

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(examples): execute Jupyter demo notebook in place

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(examples): update Jupyter base image

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-10-01 19:26:13 +00:00
John T. MyersandJohn Myers 1b77cd4e93 chore(gator): default to GPT-6.1 Sol (fixes #4064) (#4065)
Signed-off-by: John Myers <johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <johntmyers@users.noreply.github.com>
2026-10-01 18:02:20 +00:00
krishicks 8091f66877 feat(sandbox): write agent output to the container log (#4005)
Since #2726 the canonical main process's stdout and stderr are captured
in pipes that feed only the in-memory replay buffer used by sandbox
connect. Agent output therefore never reaches the container's own stdout
and stderr, so it is missing from kubectl logs, docker logs, and podman
logs and from anything that collects container logs. Before #2726 the
entrypoint inherited the container's descriptors and its output appeared
there.

Copy the main process's output to the launcher's stdout and stderr in
addition to the replay buffer, restoring the earlier behavior:

- Output is copied byte for byte to the matching stream from a
  forwarder thread per stream, after it is published to the replay
  buffer. When the container runtime falls behind on a stream, that
  stream's reader waits instead of dropping output, so backpressure
  reaches the agent as it did with inherited descriptors, while the
  other stream and attachments keep receiving output.
- Before the main process's exit is published, the output readers
  finish and queued output is drained to the container log, so an
  agent's final lines are not lost at shutdown. A 30 second deadline
  covers both; when it expires, readers waiting on the container log
  are released and drain the pipes into the replay buffer only, so a
  stalled container log cannot block exit reporting.
- PTY-mode processes are not copied. The terminal stream carries escape
  sequences and echoed input, and terminal commands never reached the
  container log before #2726.
- Exec, SSH, and SFTP sessions are not copied.

Launcher log lines keep their existing format and remain in the
container's stderr. They are written as whole lines, and a newline is
inserted first when the agent left stderr mid-line, so launcher and
agent lines do not merge.

The Docker and VM drivers appended the tail of the workload's output to
failure messages: Docker the workload container's log, and the VM driver
the guest console, which carries the launcher's stdout and stderr. Those
messages land in the sandbox's Ready condition and in platform events
that the gateway republishes to the sandbox event stream. With agent
output in that log, those messages would carry arbitrary agent output,
including anything sensitive the agent prints, into gateway status and
events. The supervisor starts its health endpoint only after the agent
starts, so every Docker failure path could include agent output, and the
VM driver reports one whenever the VM or host supervisor exits. Forward
only the supervisor's log tail, matching the Podman driver, which reads
the workload log solely to match fixed launcher markers and never
forwards raw workload output. The workload's output remains available
through docker logs and the VM's rootfs-console.log.

Document where main process output appears in the logging docs and the
cluster debugging skill.

Closes #3928

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-10-01 17:23:09 +00:00
Fede Kamelhar 71440b28f4 fix(policy): refresh pending proposals when the sandbox policy changes (#3923)
* fix(policy): refresh pending proposals when the sandbox policy changes

Approving, removing, or undoing a rule, or updating the sandbox policy,
changes the inputs every other pending proposal was evaluated against.
Only proposals the new policy covered were reconciled; the rest kept
their old prover result and review token. The review surface
(GetDraftPolicy) therefore showed a stale evaluation, and the first
approval of the next proposal refreshed it and failed with
FAILED_PRECONDITION, so approving proposals one after another always
failed once.

Re-evaluate the remaining pending proposals at each policy change,
reusing the cached prover result unless the proposal's inputs changed.
Approval still rejects a review token that does not match the stored
evaluation, so a reviewer holding a pre-refresh evaluation must still
refetch it.

When a refresh does happen at approval time (inputs changed between
fetch and approve), the CLI now explains that the rule was re-evaluated
and how to review it, instead of printing the raw gRPC status.

Closes #3884

Signed-off-by: fede-kamel <fkamelhar@gmail.com>

* fix(policy): make pending proposal refresh race-safe and bounded

Store refreshed evaluations with a compare-and-swap: the store re-reads
the proposal, refuses when its rule name, proposed rule, or review token
changed since the evaluation read it, copies only the evaluation fields
onto the stored record, and updates only if the payload is still the one
it read. A refresh can no longer revert a concurrent edit or observation,
and the edit path uses the same guard against a concurrent refresh.

Bound each refresh to the 32 newest pending proposals; the rest keep the
approval-time recheck, which still refuses a stale review token. Operator
decisions (approve, approve-all, remove, undo) refresh before responding.
UpdateConfig, which holds the gateway-wide sandbox sync guard, and
agent-driven auto-approval refresh in a background task instead, one per
sandbox with later changes coalesced into a single rerun.

Refs #3884

Signed-off-by: fede-kamel <fkamelhar@gmail.com>

* docs(policy): describe proposal rechecks after approvals and approve-all

Explain that approving, removing, or undoing a rule rechecks the other
pending proposals so they can be approved one after another, when the
recheck is deferred or bounded, and what rule approve reports when a
proposal changed after it was listed. Show rule approve-all in Run Your
First Agent with its security-flag behavior.

Refs #3884

Signed-off-by: fede-kamel <fkamelhar@gmail.com>

* fix(policy): refresh pending proposals after a full policy replacement

A full policy UpdateConfig (openshell policy set) re-reads the latest
revision after its atomic write, finds the revision it just committed,
and returns before reaching the pending-proposal refresh at the end of
the handler. Pending proposals kept their stale evaluation, so rule get
showed the old candidate and the next approval failed with the refresh
precondition. Schedule the background refresh right after the commit.

Refs #3884

Signed-off-by: fede-kamel <fkamelhar@gmail.com>

---------

Signed-off-by: fede-kamel <fkamelhar@gmail.com>
2026-10-01 16:24:51 +00:00
Shiju 8d418f1f62 fix(supervisor): bound pending exec stdin and cancel stalled writers (#3846)
* fix(supervisor): bound pending exec stdin and cancel stalled writers

Signed-off-by: Shiju <shiju@nvidia.com>

* docs(supervisor): separate pending stdin guidance from CLI modes

Keep the pending-input limit beside the RPC lifecycle contract so the streaming CLI documentation can merge independently.

Signed-off-by: Shiju <shiju@nvidia.com>

---------

Signed-off-by: Shiju <shiju@nvidia.com>
2026-10-01 16:23:11 +00:00
John T. MyersandJohn Myers 6e369f2396 chore(agents): simplify contributor instructions and workflows (#3987)
* chore(agents): simplify contributor instructions and workflows

Closes #3980

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* docs(contributing): scope verification to affected components

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* docs(contributing): standardize issue branch naming

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <johntmyers@users.noreply.github.com>
2026-10-01 16:20:35 +00:00
Fede Kamelhar ffcbe6280c fix(cli): start sandbox exec without waiting for piped stdin EOF (#4006)
With a non-terminal stdin, sandbox exec read stdin to EOF before it sent
the exec request. A pipe that never closes (CI runners, supervisors, agent
harnesses) blocked the CLI forever in read(2) without the gateway ever
seeing the request, and a slow producer delayed the command until EOF.

Collect piped stdin on a detached reader thread for at most 200 ms. Input
that reaches EOF within that window still travels in the single request
that older gateways need. If the pipe is still open, start the command
through the streaming RPC and forward the collected prefix plus the rest of
stdin as it arrives, closing remote stdin at EOF. The 4 MiB cap covers the
prefix and the streamed remainder together.

Closes #3993

Signed-off-by: Federico Kamelhar <federico.kamelhar@oracle.com>
2026-10-01 16:11:49 +00:00
Shiju f2901393e6 fix(supervisor): wait for repair when the gateway refuses a startup policy write (#3785)
* fix(supervisor): wait for repair when the gateway refuses a startup policy write

Startup writes the sandbox policy to the gateway in two cases: it
uploads a discovered image policy when the gateway has none, and it
writes the policy back after adding the proxy baseline filesystem paths.
When the gateway refused either write with FAILED_PRECONDITION or
INVALID_ARGUMENT, for example because the policy binds a provider that
is not attached, startup treated the refusal as a permanent error and
the supervisor exited. The sandbox never reached the ConfigurationInvalid
repair state that other startup rejections use.

Report such a refusal as a configuration rejection carrying the
gateway's message, log it once per write and error code, and keep
polling, so attaching the provider or replacing the policy completes
startup. Other error codes keep their current handling: transient codes
are retried, and permission, not-found and authentication failures
still end startup.

Skip the baseline-path write-back while a global policy is active. The
gateway refuses every sandbox policy write in that state, so startup
exited whenever a global policy lacked a baseline path. The supervisor
now adds the paths to its own copy of the policy without saving a
revision.

Signed-off-by: Shiju <shiju@nvidia.com>

* test(supervisor): stabilize startup refusal log capture

Keep a second tracing dispatcher alive while capturing startup refusal
logs. With only one dispatcher, a parallel test thread without a default
subscriber can cache Interest::never for the shared OCSF callsite after
the capture thread rebuilds the cache.

Preserve the exact log-count, diagnostic, configuration-generation and
repair assertions. Production startup behavior is unchanged.

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(supervisor): reconcile stale startup rejection reports

Refetch desired configuration immediately when a rejection report is aborted because its generation changed. Preserve acknowledged rejection pacing and all other report error handling.

Signed-off-by: Shiju <shiju@nvidia.com>

* test(supervisor): box startup repair race futures

Keep the repair regressions below the large-future lint threshold without changing their inputs, scheduling, or assertions.

Signed-off-by: Shiju <shiju@nvidia.com>

---------

Signed-off-by: Shiju <shiju@nvidia.com>
2026-10-01 16:11:05 +00:00
Oliver Calder 1ad4e428a6 fix(snap): simplify snap hooks (#3988)
* fix(snap): simplify snap hooks

The `post-refresh` hook runs after initial snap installation as well, so
there is no need to call the `install` hook from within the
`post-refresh` hook; instead, the logic can simply be moved into the
`post-refresh` hook directly, and the `install` hook removed.

Also, the existing `install` hook logic looked for an insecure
configuration, and if found, replaced the entire configuration file with
a minimal default in the current format. But OpenShell does that default
behavior without any config file, so we may as well simply remove the
configuration file entirely to keep up-to-date with the current default
behavior. Let OpenShell create a configuration file if it needs to,
rather than auto-create one via the packaging scripts.

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* fix(snap): remove the connect-plug-docker hook

The `openshell:docker` is auto-connected to the system `:docker` slot,
so there should not be a need to separately restart the gateway service
when the interface is connected.

For locally-built test snaps which were not published to the store, the
autoconnection is not made, but when the snap is installed, the gateway
will attempt to start anyway and fail to find any available compute
driver, so quickly restart until it hits the systemd start-limit, after
which systemd prevents the service from being started again. If a user
tries to manually connect their locally-built `openshell` snap to the
`:docker` slot, then the `connect-plug-docker` hook runs and triggers a
restart of the gateway, which will usually fail because the start limit
has already been hit. An error in the hook will thus cause the interface
connection to be undone, which is undesirable.

Thus, we can remove this hook entirely, and instead allow interface
connections to succeed as intended. The user still needs to manually
restart the gateway service after making a manual connection (as was the
case previously) and probably needs to `systemctl reset-failed` first,
but at least connection will succeed beforehand so they can proceed with
these steps.

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* fix(snap): set refresh-mode: endure again, with manual restart

Return to the previous behavior before commit a67567e58, where the
gateway is not stopped before refreshes. The `post-refresh` hook
now restarts the gateway if the TLS configuration was corrected, so we
don't have to enforce restarting the gateway on every refresh even when
not necessary. Thus, set `refresh-mode: endure`, and let the hook decide
when the gateway needs to be restarted.

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* fix(snap): update docs and tests to reflect snap hook changes

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

* docs(snap): remove verbose explanation of snap gateway refresh behavior

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>

---------

Signed-off-by: Oliver Calder <oliver.calder@canonical.com>
2026-10-01 15:10:38 +00:00
Evan Lezar cb193ef1c2 ci: use package installers consistently in integration tests (#4056)
* test(tmachine): add Fedora RPM package installer

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* ci: qualify Ubuntu branch installs with DEB packages

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* ci: align package installers across integration matrices

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-10-01 15:05:35 +00:00
Evan Lezar 82e889374f test(tmachine): add Fedora RPM package installer (#4025)
Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-10-01 14:29:09 +00:00
bornav b3e9201316 feat(flake): add packages required to run mise command allowing us to compile vm driver from source (#2151) 2026-10-01 14:23:04 +00:00
Simon ScattonandMrunal Patel 5698c4f746 fix(ci): qualify protobuf compatibility by release train (#4049)
* feat(ci): detect breaking protobuf changes

Compare the proto module against the PR or merge-group base and report Buf violations in Branch Checks. Add local reproduction and fixture coverage.

Closes #3794

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(ci): pin protobuf check container image

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(ci): qualify protobuf compatibility by release train

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* refactor(ci): reuse protobuf compatibility action

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* refactor(ci): run protobuf checks as a Nix app with one ref

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

---------

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
Signed-off-by: Simon Scatton <sscatton@nvidia.com>
Co-authored-by: Mrunal Patel <mrunalp@gmail.com>
2026-10-01 14:00:45 +00:00
Evan Lezar 0e8d9e55f9 test(e2e): remove schema parity campaign (#3864)
Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-10-01 13:17:20 +00:00
Simon Scatton fde79f1aa6 fix(ci): align integration inputs with release candidate source (#4048)
Signed-off-by: Simon Scatton <sscatton@nvidia.com>
2026-10-01 13:01:20 +00:00
Drew Newberry fe38637533 fix(runtime): recover SSH relays and bound startup diagnostics (#4011)
* fix(runtime): recover SSH relays and bound startup diagnostics

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(server): deliver pending relays once per supervisor session

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(server): satisfy relay delivery clippy diagnostics

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(server): bound relay setup with one absolute deadline

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-10-01 12:47:20 +00:00
Simon Scatton e21b7fd8cf chore(build): remove bundled Z3 support (#3275)
* chore(build): remove bundled Z3 support

Signed-off-by: Simon Scatton <sscatton@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(build): preserve vendored Z3 for local gateway artifacts

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

---------

Signed-off-by: Simon Scatton <sscatton@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-10-01 12:08:33 +00:00
Drew Newberry 021400be8a refactor(auth): separate sandbox identity from TLS (#3110)
* refactor(auth): separate sandbox identity from TLS

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(auth): clarify gateway mTLS behavior

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(auth): include workspace scope in TLS authorization checks

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(e2e): bound service auth sandbox names for large PIDs

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
v0.1.3-pre.2
2026-10-01 04:33:25 +00:00
John T. Myers 2935e9731b fix(gateway): delete finalized ephemeral sandboxes while connected (#3984)
Start driver cleanup after terminal finalization and retain disconnect fallback. Add detached success and failure e2e coverage across supervisor-based drivers.

Closes #3938

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-10-01 00:09:08 +00:00
Jim Meyer 5a91572be7 fix(gator): require full head SHA for /ok to test (#4007)
* fix(gator): require full head SHA for /ok to test

copy-pr-bot will stop accepting abbreviated SHAs in /ok to test comments.
Tell gator to read the full 40-character head SHA immediately before
posting, and make the gh wrapper reject any /ok to test comment that is
not exactly the command with the current full head SHA.

Signed-off-by: Jim Meyer <jimeyer@nvidia.com>

* fix(gator): drop gh wrapper /ok to test guard

Keep the change to the gator-gate skill instructions only.

Signed-off-by: Jim Meyer <jimeyer@nvidia.com>

* fix(gator): unify /ok to test SHA placeholder

Use <full-head-sha> for every /ok to test reference in the gator-gate
skill and state the full-SHA requirement directly.

Signed-off-by: Jim Meyer <jimeyer@nvidia.com>

---------

Signed-off-by: Jim Meyer <jimeyer@nvidia.com>
2026-10-01 00:02:25 +00:00
krishicks 9912d21d30 fix(e2e): keep locally built Kubernetes images off the chart's default registry (#3991)
679b19067 added global.image.registry (ghcr.io/nvidia) as the fallback for
empty per-image registries and split e2e image references into registry and
repository. Locally built images such as openshell/gateway:<tag> have no
registry host, so the chart rewrote them to ghcr.io/nvidia/openshell/* and
the k3d cluster could not pull them. Clear global.image.registry in the
Kubernetes e2e wrapper, which sets every image's registry explicitly.

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-09-30 23:06:03 +00:00
Matthew Grossman 4784e79451 refactor(sandbox): remove unreachable root-side identity and workspace code (#3979)
* refactor(sandbox): remove unreachable root-side identity and workspace code

RFC 0012 moved the workload into its own capability-free container that
starts as the final sandbox identity. The sandbox no longer runs a root
supervisor that prepares the filesystem, rewrites account files, resolves
OCI USER entries, or drops privileges before launching the workload, so
that code had no production callers.

Remove the unreachable paths and their tests:

- prepare_filesystem / prepare_filesystem_with_identity, the /sandbox and
  OCI workspace chown preparation, and the root-side workspace validation
  (validate_oci_workspace and its privilege-dropped subprocess)
- the hidden validate-workspace subcommand
- drop_privileges / drop_privileges_with_identity, capability bounding set
  clearing, validate_sandbox_user/group, and /etc/passwd and /etc/group
  rewriting
- the sandbox-side OCI USER resolver (identity.rs) and
  ResolvedProcessIdentity; the boundary now writes the driver-resolved
  UID/GID into the policy directly

The workspace check that still runs inside the capability-free boundary
(validate_oci_workspace_as_effective_identity) is unchanged.

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* chore(sandbox): remove unused capability dependency and refresh Landlock comments

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

---------

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>
2026-09-30 23:04:50 +00:00
krishicks dde8a9a57f fix(server): log polled request responses at debug (#3974)
log_response has always logged every gateway response at INFO, including
health probes and the GetSandboxConfig and provider-readiness polls each
supervisor makes. #3915 demoted the request spans for those polled paths
to DEBUG, which stripped the request{method path} prefix from the log
line at INFO but left the line itself, so the gateway log fills with
bare 'response status=200' lines several times per second.

Follow the span's level: polled requests log their response at DEBUG,
or WARN on a 5xx so probe and poll failures stay visible.

Signed-off-by: Kris Hicks <khicks@nvidia.com>
v0.1.3-pre.1
2026-09-30 22:54:42 +00:00
Derek Carr 912a077bd6 feat(service): add bearer authorization passthrough (#3796)
* feat(service): add bearer authorization passthrough

Signed-off-by: Derek Carr <decarr@redhat.com>

* docs(sdk): add service authorization migration guide

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(server): remove stale version import

Signed-off-by: Derek Carr <decarr@redhat.com>

* docs(upgrade): remove service authorization SDK guide

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(e2e): relabel provider readiness TLS mount

Signed-off-by: Derek Carr <decarr@redhat.com>

* test(e2e): stabilize exposed service routing

Signed-off-by: Derek Carr <decarr@redhat.com>

* test(e2e): support HTTPS service routing

Signed-off-by: Derek Carr <decarr@redhat.com>

---------

Signed-off-by: Derek Carr <decarr@redhat.com>
2026-09-30 20:22:46 +00:00
Shiju 374c035962 fix(network): refuse protocol upgrades on GraphQL endpoints (#3841)
* fix(network): refuse protocol upgrades on GraphQL endpoints

Refuse Upgrade headers before forwarding GraphQL-over-HTTP requests.
Share the protocol refusal table with JSON-RPC and MCP, and close
unexpected protocol switches before relaying frames.

Keep GraphQL-over-WebSocket inspection on separate WebSocket endpoints.
Cover upgrade refusal, audit mode, subscription handshakes, and ordinary
HTTP and WebSocket controls. Update the current policy documentation.

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(network): refuse GraphQL upgrades before reading bodies

Validate the HTTP head and endpoint authority before upgrade refusal, then inspect ordinary GraphQL bodies. Preserve missing-authority credential rejection after body inspection.

Signed-off-by: Shiju <shiju@nvidia.com>

---------

Signed-off-by: Shiju <shiju@nvidia.com>
2026-09-30 20:10:04 +00:00
Eric CurtinandDrew Newberry 07a486d751 fix(cli): accept sandbox name before -- in exec (#3901)
* fix(cli): accept sandbox name before -- in exec

Closes #3882

Signed-off-by: Eric Curtin <eric.curtin@docker.com>

* fix(cli): define exec grammar in clap

Signed-off-by: Eric Curtin <eric.curtin@docker.com>

* docs(sandboxes): remove exec overview change from PR

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Eric Curtin <eric.curtin@docker.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
2026-09-30 18:50:04 +00:00
krishicks 7caff12d3c perf(otel): stop exporting spans from steady-state polling (#3915)
Store operation spans and request spans for supervisor-polled RPCs
(GetSandboxConfig, ReportProviderReadiness) use DEBUG level, so the
default INFO filter no longer exports them. The provider credential
refresh worker opens its span only when a state has work.

Refs #2698

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-09-30 15:19:39 +00:00
Simon Scatton 21fea95935 test(tmachine): add K3s conformance scenario (#3848)
* test(tmachine): add K3s conformance scenario

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* refactor(tmachine): use Helm values file for K3s installer

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* ci(tmachine): run K3s conformance in integration jobs

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* ci(tmachine): verify installer scripts and document version baseline

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

---------

Signed-off-by: Simon Scatton <sscatton@nvidia.com>
2026-09-30 14:49:09 +00:00
Evan Lezar 5acaaba192 test(conformance): verify deletion through sandbox list (#3792)
Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-09-30 09:15:45 +00:00