209 Commits
Author SHA1 Message Date
araza008 8f22fe84f6 ci(windows): run MXC host probe, WebSocket agent, and OpenClaw forwar… (#3826)
* ci(windows): run MXC host probe, WebSocket agent, and OpenClaw forward examples checks

Add separate hosted CI tasks for the MXC host probe, WebSocket agent,
and OpenClaw forward examples. Use mock workloads to verify gateway,
CLI, driver, and sandbox lifecycle wiring without requiring wxc-exec.

Document that mock passes do not validate forwarding or MXC enforcement.

* fix(tests): enhance environment isolation for WebSocket and OpenClaw mock tests

* fix(mxc): enhance OpenClaw mock validation to require 'Ready' sandbox state

* fix(mxc): restore native process helper in WebSocket example

Restore Invoke-NativeCaptured for the PowerShell argument regression test and delegate CLI execution through it while preserving explicit gateway endpoint selection.

Signed-off-by: Akber Raza <akberr@nvidia.com>

---------

Signed-off-by: Akber Raza <akberr@nvidia.com>
2026-09-30 14:13:26 -07:00
Prekshi Vyas d46a814141 ci(windows): exercise MXC credential, audit, and aggregate E2E flows (#3787)
* ci(windows): exercise MXC provider credential example

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* ci(windows): exercise MXC OCSF audit example

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* ci(windows): enable aggregate MXC example E2E

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* ci(windows): select native MXC mock target

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* fix(windows): isolate MXC example CI harnesses

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

---------

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>
2026-09-29 19:50:53 -07:00
Prekshi Vyas 7ba7a39d09 ci(windows): exercise MXC inference demos with mock API (#3780)
* ci(windows): exercise MXC Ollama demo with mock API

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* ci(windows): cover both MXC inference demos

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* fix(mxc): preserve executable extension resolution

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* test(mxc): use absolute PowerShell in lifecycle checks

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* test(mxc): make lifecycle write probes deterministic

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* test(mxc): assert stable lifecycle completion

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

---------

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>
2026-09-29 13:39:23 -07:00
Prekshi VyasandShailendra Singh 651ed7e03d NVBug 6783374: make MXC HTTPS L7 qualification authoritative (#3479)
* test(mxc): qualify HTTPS L7 enforcement (NVBug 6783374)

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* docs(mxc): clarify real qualification failures

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* docs(mxc): align real test skip semantics

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>

---------

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>
Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Co-authored-by: Shailendra Singh <shailendras@nvidia.com>
2026-09-22 12:38:07 -07:00
Prekshi Vyas a69d0319f2 test(windows): define GB300 MXC qualification contract (NVBug 6643699) (#3471)
* test(windows): define GB300 MXC qualification contract

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* test(mxc): harden GB300 qualification provenance

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* test(windows): package portable GB300 qualification

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* test(windows): honor external qualification checkout

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* revert: remove portable GB300 packaging

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* test(mxc): bind GB300 evidence to exact inputs

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

---------

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>
2026-09-22 10:52:21 -07:00
Prekshi Vyas fb2980e077 fix(windows): restore MXC qualification and cold-start readiness (#3468)
* fix(windows): restore MXC qualification gates

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* fix(mxc): poll target for full readiness budget

---------

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>
2026-09-18 23:01:55 +05:30
Drew Newberry 49b4f0eb7f feat(mxc): add UI policy, credentials, relay lifecycle, and proxy auth
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-17 12:06:06 -07:00
Mrunal Patel b799fccb8b fix(auth): harden OIDC trust root retrieval (#3332)
* fix(auth): harden OIDC trust root retrieval

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(e2e): pass OIDC HTTP acknowledgement value

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

---------

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
2026-09-15 17:00:45 +00:00
26f2f96393 feat(mxc): add Windows host proxy for MXC sandbox network egress (#3163)
* Implement Windows host proxy integration and update dependencies for OpenShell

* Update README and gateway config to clarify egress proxy address handling and allocation

* Refactor ProxyIdentityMode to return Result for static_binary and add tests for binary path and SHA256 hash

* Enhance platform_hosts_path for Windows to use SystemRoot and improve error handling for hosts file reading

* Refactor FileFingerprint to use Option for mtime and ctime, simplifying metadata handling

* Add conditional compilation for Windows host module

* add unit tests for OPA policy evaluation and identity handling

* remove openshell-supervisor-network from unsupported driver package test exclusion list

* feat(mxc): enable host proxy TLS state generation

Generate per-sandbox TLS state for the MXC host proxy so HTTPS L7 enforcement can use the same MITM path as Linux. Grant generated CA material to the MXC process and inject standard trust env vars, while matching Linux behavior by disabling TLS termination on CA setup failure and relying on proxy fail-closed handling.

* fix(docs): remove outdated notes on governed egress from docs

* fix(tests): update TLS environment variable paths to use temporary directory

* fix(examples): make run-mxc-e2e harness correct and orphan-free

The MXC e2e harness never actually exercised the fs scenarios: it started
the gateway once and patched agent_command per scenario AFTERWARDS, so the
running gateway kept launching the default demo agent (not shipped in the
kit) and every fs scenario failed with CreateProcessW error:2. It also
scored on the `sandbox create` exit code (non-zero due to the harmless
interactive attach), wrote sandbox records to the persistent gateway DB
(leaving orphans that collided on later runs), and its deny scenarios never
proved denial.

Changes:
- Start a FRESH gateway per scenario so each scenario's agent_command is
  actually loaded (root cause of CreateProcessW error:2).
- Score by on-disk artifact / expected outcome, not `sandbox create` exit.
- Real deny assertions: a control write to a granted path must succeed
  (proves the agent ran) while the denied write must be absent. fs-empty
  probes an ungranted out-of-share path (share_dir is mapped rw by design).
- Run the gateway on an ephemeral in-memory DB (sqlite::memory:) so the
  harness never writes to the persistent store and cannot leave orphan
  sandbox records; also use unique per-run sandbox names + pre-delete.
- Fix the process_container probe: use a real cwd + absolute cmd.exe
  (canonical wxc-exec does not expand %TEMP% -> 0x8007010B).
- Fix summary counts (@() so a single FAIL is counted and exit is non-zero).

Verified PASS=4 FAIL=0 on 7F203-MXC-003 (no BaseContainer velocity keys)
using a canonical wxc-exec build (AppContainer fallback).

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(e2e): probe timeout is milliseconds (10ms->30000ms)

MXC process.timeout is wall-clock ms (wire.rs). The 10 value meant 10ms,
which the base-container tier (7F203-MXC-001/.181) enforced strictly and
timed the probe out. AppContainer path (.18/-003) happened to slip under
it. Bump to 30000ms so the process_container preflight probe is reliable
across both tiers.

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc): use native paths in real runtime probes

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc): make processcontainer work with mxc-latest-released wxc-exec

Three fixes to support the release wxc-exec binary (BaseContainer dispatcher)
in addition to mxc-fixes-env-vars:

1. Seed process env from host (driver.rs)
   ProcessContainer starts with a completely blank environment -- no PATH,
   SystemRoot, or anything.  Seed the process env from the gateway host
   environment so the agent binary can locate DLLs and run.  Skip internal
   Windows drive-letter variables (keys starting with '=') which cause
   CreateProcessW to return ERROR_ENVVAR_NOT_FOUND.  User agent_env entries
   and TLS CA vars are applied as overrides on top of the host env.

2. Remove TLS readonly_paths grant (driver.rs)
   The release wxc-exec (BaseContainer dispatcher) requires write-DAC
   permission on every path in readonly_paths to set up AppContainer ACLs.
   Adding the proxy's temp TLS directory caused a DACL error and exit -1.
   The CA cert paths remain available to the agent via TLS env vars.

3. Remove allowedHosts from network JSON (mxc.rs)
   The release wxc-exec rejects network.allowedHosts / network.blockedHosts
   on Windows with "not yet supported".  Removed the loopback exemption
   attempt (127.0.0.1, ::1, localhost) from the network section.
   Intra-container loopback works natively in the release binary without
   it -- the spawner can connect to the server at 127.0.0.1:22000 directly.

Additional changes:
- mxc-ws-agent.rs: add relay-debug.txt error capture and relay-ready.txt
  marker for reliable timing of host client connections.
- mxc-ws-gateway.toml: debug = true for JSON config dump during diagnosis.
- run-ws-agent-test.ps1: default port changed to 17670 (gateway default);
  relay-ready.txt polling before ws-echo to avoid connecting before the
  spawner has established the proxy bridge.

Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(e2e): address CodeRabbit review on run-mxc-e2e.ps1 (MR !46)

Four robustness/correctness fixes from CodeRabbit:

1. Start-Gw: kill the spawned gateway before the "did not start within 30s"
   throw. If the process is alive but never binds the port, $gw is not yet
   assigned in the caller, so the finally block cannot reap it -> orphan
   gateway holding the port for the next run.

2. create-fail scoring: a non-zero `sandbox create` exit alone is not proof
   of a policy rejection (gateway-registration/transport/fixture errors also
   exit non-zero and would false-pass). PASS now requires a genuine
   rejection signal (network / invalid_argument / network_policies) AND that
   it is not an infrastructure failure; other non-zero exits go to FAIL with
   output captured.

3. deny scenarios (ControlTarget path): snapshot the deny target AFTER
   Wait-File lands the control artifact, so a late denied write (enforcement
   regression racing the control write) can no longer be recorded as PASS.

4. -KeepRunning: break out of the scenario loop after the first scenario so
   a later scenario does not start a second gateway on the same port
   (previously a reliable port collision instead of a usable debug mode).

Re-verified PASS=4 FAIL=0 on both boxes (7F203-MXC-001 base-container and
7F203-MXC-003 AppContainer fallback); network-policy-rejected correctly
scores as "policy rejection".

Signed-off-by: Akber Raza <akberr@nvidia.com>

* feat(mxc-e2e): collect run-mxc-e2e output into a results bundle

Mirror the sibling run-*.ps1 scripts by collecting every run's logs into a
timestamped results-e2e-<stamp>\ folder and zipping it. The bundle contains the
console transcript, per-scenario gateway stdout/stderr, the exact TOML rendered
for each scenario, the policy fixture used, and a summary.txt with the verdict
table.

Per-scenario gateway logs now land in gateway.<scenario>.log/.err.log inside the
bundle instead of a single fixed gateway.e2e.log in the script directory.

Wrap pre-flight, mode setup, scenario definitions, and the scenario loop in a
single try/catch/finally so the finally always writes the summary, stops the
transcript, and zips the bundle -- even on a pre-flight failure. The existing
per-scenario gateway-cleanup try/finally stays nested inside. All scenario
logic, scoring rules, and comments are preserved.

Signed-off-by: Prashant S Khodade <pkhodade@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc-e2e): address CodeRabbit review on run-mxc-e2e.ps1

- Require -Scenario when -KeepRunning: the loop breaks after the first
  scenario, so a full-suite run would execute only one scenario yet still
  report the suite as PASS. Fail fast so a partial run can't be mislabeled
  complete.
- Start-Transcript now runs inside the guarded try block with a
  $transcriptStarted flag; Stop-Transcript is only called when it actually
  started, so a Start-Transcript failure still yields the results bundle.
- Wrap the -Scenario filter in @() so a single exact match stays an array
  (reliable .Count and a proper array for the scenario loop on PS 5.1).

Signed-off-by: Prashant S Khodade <pkhodade@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(examples): pass gateway config via OPENSHELL_GATEWAY_CONFIG for spaced paths

Start-Process -ArgumentList does not quote array elements, so launching the
gateway with a bare --config <path> token split on any space in the install
path (e.g. C:\Users\First Last\...), and clap rejected the fragment with
'unrecognized subcommand'. Every MXC example launcher that started the gateway
hit this when the kit was unzipped under a path containing a space.

Pass the config path through the OPENSHELL_GATEWAY_CONFIG env var (which the
gateway already reads via clap) and drop the --config token. Env vars carry
spaces safely.

Affected: run-ocsf-audit, run-mxc-e2e, run-demo, run-inference-test,
run-ollama-test. run-mtls-test was not affected (its launch passes no config
path). Root-caused and fix-verified on 7F203-MXC-003 from a spaced path.

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(run-mxc-e2e): improve scoring logic and enhance command execution handling

* fix(mxc): reconcile proxy support after rebase

Restore the proxy-enabled OCSF audit example removed by 13185f6e now that the host CONNECT proxy is present. Adapt the proxy lifecycle test to the target branch's DriverSandboxSpec policy delivery contract.

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc): grant sandbox access to proxy CA

- Share the per-sandbox public CA bundle with the AppContainer
- Add real wxc-exec HTTPS proxy coverage and document trust isolation

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc): reject unsupported network middleware

- Reject middleware-bearing MXC policies before sandbox lifecycle begins
- Guard host proxy startup and document the unsupported registry path
- Add mapper, lifecycle, and host proxy regression coverage

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc): reconcile host proxy with main

- remove obsolete inference routing from the host proxy adapter
- use the workspace AWS-LC provider in host-proxy tests
- adapt the forward-proxy test to ProxyIdentityMode

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(build): switch to bundled Z3 for Windows MSVC builds

* fix(mxc): reconcile host proxy after rebase

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Akber Raza <akberr@nvidia.com>
Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
Signed-off-by: Prashant S Khodade <pkhodade@nvidia.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Jamie King <jamiek@nvidia.com>
Co-authored-by: Shailendra Singh <shailendras@nvidia.com>
Co-authored-by: Prashant Khodade <pkhodade@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
2026-09-14 22:58:49 +00:00
krishicks 5b9daab935 fix(ci): restore mise run ci on macOS (#3294)
- Replace BSD-incompatible in-place sed calls with portable temp-file rewrites.
- Remove test-only shell interception and capture generated gateway config
  directly.
- Allow parity tests to use supplied supervisor binaries without resolving a
  Linux target.
- Normalize temporary-directory paths and use portable RPM config installation.
- Set a valid setuptools-scm version for Python protobuf generation in Jujutsu
  checkouts.

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-09-12 00:17:15 +00:00
Piotr Mlocek 5b57f0d154 fix(ci): restore Windows test portability (#3288)
* fix(ci): restore Windows test portability

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(tasks): skip Unix lockfile check on Windows

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(tasks): use buf shim on Windows

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

---------

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-09-11 21:40:08 +00:00
Piotr Mlocek b92620e838 ci(rust): reject stale Cargo lockfiles (#3227)
* ci(rust): reject stale Cargo lockfiles

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(ci): clarify lockfile validation policy

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* refactor(ci): structure and test Cargo lockfile validation

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* chore(ci): remove lockfile validator regression tests

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(rust): complete locked validation and lint examples

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(rust): check lockfile diffs after validation

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(build): shorten lockfile validation notes

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(rust): skip lockfile check after failures

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

---------

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-09-11 08:52:41 +00:00
Jesse JaggarsandDrew Newberry 02b664bb0d refactor(config): normalize and enforce gateway schema v2 (#2814)
* refactor(config): normalize compute driver field names

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* refactor(config): introduce canonical gateway fields

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* refactor(config): enforce gateway schema version 2

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): preserve compute driver runtime guarantees

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): address schema v2 review regressions

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): complete schema v2 migration safeguards

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(config): expand schema v2 regression coverage

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(config): add schema v2 parity manifest

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): correct parity manifest inventory

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* docs(config): record schema v2 intentional changes

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* docs(config): disposition schema v2 parity gaps

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): add dual schema parity harness

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): establish compute lifecycle parity baseline

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): preserve gateway option compatibility

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): record gateway option parity

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* docs(config): close gateway-wide parity gaps

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(podman): apply configured pids limit

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): validate Podman option parity

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): add Kubernetes option parity harness

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): record Kubernetes option parity

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): disposition VM parity lanes

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): add external driver parity lane

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(e2e): preserve external driver pull policy

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): attest parity artifacts and launches

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): require clean parity build sources

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): bind parity runtime artifacts

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(e2e): use isolated supervisor tags

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(e2e): qualify parity image tags

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(e2e): serve parity supervisor locally

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): isolate parity podman services

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): harden parity evidence provenance

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): pin parity sandbox artifacts

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): attest parity runtime inputs

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): bind parity runtime evidence

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): record compute boundary parity

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): disposition cross-cutting parity lanes

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(packaging): preflight gateway config upgrades

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): preserve rebase integration guarantees

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(ci): isolate temporary git signing config

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): update remaining schema v2 consumers

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(ci): provide e2fs tools to VM tests

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): align preflight with gateway startup

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(vm): preserve rootfs tar configuration

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* chore(config): adopt duration unit constructors

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(packaging): preflight RPM gateway config

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): address driver review findings

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(e2e): require fresh semantic parity evidence

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(docker): update tests for renamed sandbox label

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(gateway): preserve selective driver coverage after rebase

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
2026-09-11 05:00:24 +00:00
Drew Newberry 38f2aef930 feat(gateway): support selective compute driver builds (#3118)
* feat(gateway): support selective compute driver builds

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(gateway): support selective Windows MXC builds

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-11 00:36:09 +00:00
Piotr Mlocek ddc8bba967 ci(windows): add Windows MSVC CI jobs (#2738)
* fix(ci): preserve Windows Rust build cache

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(ci): invalidate empty Windows caches

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* perf(ci): cache Windows builds with sccache

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(windows): restore target directory caching

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* perf(ci): use prebuilt Z3 on Windows

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* perf(ci): layer sccache on Windows target cache

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(windows): split PR checks from main validation

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(windows): separate checks builds and cache seeding

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(ci): simplify Windows build dependency

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(ci): rely on Windows job dependency status

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(ci): use valid opt-in Windows ARM runner

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(windows): keep ARM64 validation local

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(ci): install Clippy for Windows validation

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(windows): focus platform lint coverage

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(licenses): explain bzip2 allowance

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(windows): simplify workflow name

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(windows): allow async platform stub

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(network): make file fingerprints portable

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(windows): lint supported deliverables

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(windows): allow platform-gated lint

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* chore(ci): align Windows cache action with main

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(ci): align Windows validation with prerequisites

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(windows): pin Rust toolchain action

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(ci): use enterprise-approved Windows actions

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(windows): restore strict MSVC validation

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(windows): run Rust tests with nextest

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(ci): normalize nextest lock provenance

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(windows): add native arm64 validation

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(ci): lock nextest for Windows ARM64

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(windows): resolve duplicate MXC authentication method

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* test(conformance): use native absolute paths on Windows

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(windows): address MSVC review feedback

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(windows): simplify cache key names

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(ci): isolate Windows Rust toolchains for stable caches

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(ci): configure Rustup home in runner setup

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(windows): surface sccache server write diagnostics

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(windows): remove temporary cache diagnostics

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(windows): address review feedback

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(windows): reconcile merged driver capabilities

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(deps): preserve AWS-LC-only lockfile

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* chore(deps): allow z3 prebuilt TLS wrapper

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

---------

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-09-10 23:53:01 +00:00
Piotr Mlocek ce25acca5a fix(build): honor Cargo target directory when staging binaries (#3262)
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-09-10 23:19:13 +00:00
Piotr Mlocek a0814443f1 feat(docs): publish versioned release snapshots (#3149)
* feat(docs): add version availability labels

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(docs): use supported Python for sync

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* feat(docs): publish versioned docs from releases

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(docs): format dev version label

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* chore(docs): upgrade Fern CLI to 5.112.0

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(docs): make release publishing monotonic

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(docs): preserve snapshot release identity

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(fern): document versioned publishing

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* test(docs): cover explicit snapshot rollback

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* refactor(docs): use Fern refs for versions

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(docs): bundle components for ref versions

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* revert(docs): keep complete version copies

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(docs): publish latest and dev channels

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

---------

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-09-09 23:34:34 +00:00
John T. Myers f4dc6be4b2 refactor(inference): remove managed inference routes (#3195)
* refactor(inference): remove managed inference routes

Closes #3172

Remove the inference route control plane, inference.local data path, built-in router crate, and SDK surface. Move inference workloads to explicitly imported provider profiles and native endpoints, with migration cleanup and updated tests and documentation.

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(policy): preserve alternate upstream isolation

Restore the provider policy activation guard so legacy OpenAI and Anthropic providers configured for alternate base URLs do not grant egress to the built-in public vendor endpoints.

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-09-09 18:47:22 +00:00
Gaizka Menendez 8af79a7f4b fix(podman): resolve macOS Podman socket dynamically (#3135)
* docs(podman): document macOS socket path mismatch and dynamic lookup

On macOS, Homebrew-installed Podman does not create the default socket
path that the Podman driver probes. Document the OPENSHELL_PODMAN_SOCKET
override and the podman machine inspect lookup in both the compute
drivers reference and the debug-openshell-cluster skill.

Fixes #1690

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* fix(podman): resolve macOS Podman socket dynamically

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* chore: restore debug skill file

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* chore: drop legacy debug skill path

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* fix(podman): trim unrelated e2e changes

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* fix(e2e): harden shell array expansion

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* chore: remove unrelated skill note

* ci: retrigger checks

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

---------

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>
2026-09-09 14:10:59 +00:00
alangou 3693b32841 ci(trivy): add artifact and PR configuration scans (#3185)
* ci(trivy): add artifact and PR configuration scans

Signed-off-by: Adrien Langou <alangou@nvidia.com>

* fix(ci): harden Trivy gate detection and finding diff

Signed-off-by: Adrien Langou <alangou@nvidia.com>

* feat(ci): scan released artifacts in release pipelines

Signed-off-by: Adrien Langou <alangou@nvidia.com>

* fix(ci): harden and simplify Trivy scans

Signed-off-by: Adrien Langou <alangou@nvidia.com>

* fix(ci): consolidate Trivy reports and prevent collisions

Signed-off-by: Adrien Langou <alangou@nvidia.com>

---------

Signed-off-by: Adrien Langou <alangou@nvidia.com>
2026-09-09 13:58:06 +00:00
Jorge 519e5eb35f feat(e2e): make e2e:kubernetes work transparently on OpenShift (#3183)
* feat(e2e): make e2e:kubernetes work transparently on OpenShift

Running `mise run e2e:kubernetes` on OpenShift required manual namespace
creation, SCC grants, Helm value overrides, and cleanup. A separate
`e2e:openshift` task existed but only checked pod readiness without
running the Rust e2e test suite, and even with the suite wired up the
SSH-relay `sandbox connect` path stalled to the ready timeout because
`kubectl port-forward` cannot carry round-trip-heavy SSH over the
internet.

The harness now auto-detects OpenShift via the `route.openshift.io` API
group and, on OpenShift, both configures the cluster and switches the
gateway transport automatically:

- Drives the gateway through a passthrough OpenShift Route secured with
  mandatory mTLS instead of port-forward, so the connect suites
  (live_policy_update, port_forward, sync, connect-based
  sandbox_lifecycle, settings_management) actually pass. Computes the
  Route host from the cluster ingress domain, extracts client mTLS
  material from the openshell-client-tls secret, waits for the Route to
  serve mTLS, asserts a certless caller is rejected at the TLS
  handshake, and registers an mTLS CLI gateway pointing at the Route.
- Applies an SCC-compatible Helm values overlay that removes hardcoded
  runAsUser/fsGroup, letting OpenShift assign UIDs from the namespace
  range.
- Grants the privileged SCC to openshell-sandbox before Helm install
  and removes it during cleanup.
- Grants the anyuid SCC to the PostgreSQL fixture service account in
  DB scenarios and removes it during cleanup.
- All oc commands use --context to target the correct cluster.

The OpenShift e2e overlay (ci/values-openshift-e2e.yaml) turns TLS back
on, enables the Route, promotes the cert-verified caller to a dev
principal, and forces `image.pullPolicy`/`supervisor.image.pullPolicy`
to Always so runs against the `latest` upstream image use it instead of
a stale copy cached on the cluster nodes. Every OpenShift branch is
gated on OPENSHIFT_DETECTED, so the vanilla-Kubernetes port-forward path
is unchanged.

The Helm template for podSecurityContext is wrapped with {{- with }} so
null values omit the block instead of rendering invalid YAML.

The separate e2e:openshift task and e2e-openshift.sh script are removed
since e2e:kubernetes now covers OpenShift.

TESTING.md is updated with Kubernetes e2e documentation including
OpenShift auto-detection, dropping the e2e-host-gateway feature on
remote clusters, pinning IMAGE_TAG when the CLI and image versions
differ, task variants, and environment variables.

The debug-openshell-cluster skill gains an OpenShift platform row and
two SCC failure patterns (gateway rejected over hardcoded runAsUser,
sandbox missing the privileged SCC) covering the SCC handling and
podSecurityContext behavior this change introduces.

Signed-off-by: Jorge Garcia Oncins <jgarciao@redhat.com>

* fix(e2e): harden OpenShift SCC cleanup, mTLS gate, and Route timeout

Track the anyuid SCC grant for the PostgreSQL fixture with a dedicated
OPENSHIFT_POSTGRES_SCC_GRANTED flag set before the fixture apply, so a
failed apply no longer leaks the binding; cleanup now revokes it whenever
the grant succeeded, independent of deploy state.

Validate the Route server cert in the certless security gate (curl
--cacert instead of -k) and classify curl's exit code so only a TLS
client-auth rejection (35/56) counts as the expected certless rejection;
an unrelated DNS/timeout/TLS failure now fails loudly instead of masking
a potential mTLS hole.

Raise the OpenShift Route timeout in the e2e overlay. The default HAProxy
Route timeout is 30s, which severed long-lived transfers (large sandbox
upload/download, SSH-relay `sandbox connect`) mid-stream and failed the
sync e2e tests. Set both haproxy.router.openshift.io/timeout and
timeout-tunnel to 300s: a passthrough Route proxies in TCP mode, so
timeout-tunnel governs the established tunnel while timeout covers the
pre-tunnel phase.

Document the OpenShift transport exception, oc prerequisites and SCC
grants, and make the skopeo tag-check example copy-safe in TESTING.md.

Signed-off-by: Jorge Garcia Oncins <jgarciao@redhat.com>

---------

Signed-off-by: Jorge Garcia Oncins <jgarciao@redhat.com>
2026-09-08 20:37:19 +00:00
Piotr Mlocek 039b265096 feat(middleware): define HTTP response pre-return interface (#3073)
* feat(middleware): define HTTP response pre-return interface

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(middleware): clarify HTTP response interface

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* refactor(middleware): align response result actions

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* refactor(middleware): expose response reason codes

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* refactor(middleware): share session end reasons

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* feat(middleware)!: finalize HTTP response pre-return contract

Replace the separate body_end event with HttpResponseBodyUnit.end_of_stream.
Every body-inspecting stage receives exactly one flagged unit, which may be
empty; a zero-byte body is one empty flagged unit and OpenShell never reads
ahead to set the flag.

Defer response trailers from V1 and reserve their field numbers. HTTP/1.0
clients and Content-Length bodies cannot carry trailers and that behavior was
undefined.

Add HttpResponsePreflight.permitted_body_modes, computed once from the
original upstream head so every stage sees the same list, and make an
unlisted selection a failure rather than a downgrade. Add the block_delivery
preflight action as a successful decision enforced regardless of on_error.

Expose Content-Length, Content-Encoding, and Content-Range read-only in
preflight. Cap STREAM_BYTES input units at half of max_payload_bytes and
permit deferring bytes across replacements only for fail_closed bindings,
surfaced as deferral_permitted.

Split PEER_DISCONNECT into DOWNSTREAM_DISCONNECT and UPSTREAM_DISCONNECT and
attribute WebSocket relay failures by direction instead of a generic peer
error. Compile the content-guard example in lint and branch checks so proto
renames cannot break it silently.

BREAKING CHANGE: WebSocketSessionEndReason and WebSocketSessionEnd are
replaced by the shared MiddlewareSessionEndReason and MiddlewareSessionEnd.
NORMAL_CLOSE is now NORMAL, UPSTREAM_REJECTED is now UPSTREAM_FAILURE, and
PEER_DISCONNECT is split into DOWNSTREAM_DISCONNECT and UPSTREAM_DISCONNECT.
Enum numbers are unchanged so binary wire compatibility is preserved;
generated symbols and JSON names change.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(middleware): describe skip as opting out of inspection

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* feat(middleware): add body-phase block_delivery and skip_remaining actions

Body results may now stop delivery or opt out of inspecting the rest of the
response after a prefix. One HttpResponseBlockDelivery message is shared by
preflight and body results and documents the difference between blocking
before and after head commitment. Drop the field reservations, since nothing
in this contract has shipped, and renumber session_end to close the gap.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* refactor(middleware): share HTTP body leaf messages across directions

HttpBodyUnit, HttpBodyPassThrough, HttpBodyTransform, HttpBodySkipRemaining,
and HttpBodyMode carry no response-specific semantics, so name them for reuse
by the streaming request hook. Envelopes, results, preflight, and
block_delivery stay response-specific because commitment semantics differ.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* refactor(middleware): keep HTTP body leaf messages response-specific

Reverts the shared HttpBody* naming. A direction-specific payload such as a
response-only semantic mode would otherwise add unreachable variants to the
other direction or force a source-breaking fork after 0.1.0. The streaming
request hook defines its own HttpRequestBody* messages and copies the shape;
SDKs present a direction-neutral body handler over both.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(middleware): simplify response proto comments

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): reject undispatched response bindings

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(middleware): simplify phase field comment

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* refactor(middleware): rename HTTP response preflight result

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* feat(middleware): add HTTP response trailer results

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(middleware): define response block delivery

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(middleware): define stage-local response body modes

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(middleware): define final response body units

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(middleware): define streaming response deferral

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(middleware): defer whole-body accumulation timeout

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(middleware): align response result diagnostics

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* test(middleware): cover upstream WebSocket disconnect

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): keep response streams unit-local

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(middleware): trim disconnect compatibility note

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

---------

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-09-04 21:18:02 +00:00
Piotr Mlocek d7cb6e456d fix(dev): inherit non-expiring sandbox JWT in local gateway scripts (#2636)
* fix(dev): inherit non-expiring sandbox JWT in local gateway scripts

The local gateway launcher scripts hardcode gateway_jwt.ttl_secs = 3600,
which overrides the non-expiring default introduced in #1721. Local
Docker, Podman, and VM sandboxes are still unrecoverable when the gateway
is down longer than that TTL: the on-disk token expires and only the
Kubernetes ServiceAccount path can rebootstrap, so the supervisor
crash-loops on policy fetch and the sandbox never leaves Provisioning.

Drop the override so local drivers inherit the default. gateway.sh also
serves the kubernetes driver, which is a shared deployment and must keep
a positive TTL, so it now emits ttl_secs only for that driver.

The e2e regression test added in #1721 does not catch this because the
e2e harness uses its own configs, which already set ttl_secs = 0.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(dev): expand sandbox JWT TTL in gateway config

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

---------

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-09-04 17:16:44 +00:00
alangou 64a858dade fix(ci): restore Codex Security scan execution (#3124)
* refactor(ci): resolve Codex Security range in Python

Signed-off-by: Adrien Langou <alangou@nvidia.com>

* fix(ci): allow unprivileged userns for Codex sandbox

Signed-off-by: Adrien Langou <alangou@nvidia.com>

---------

Signed-off-by: Adrien Langou <alangou@nvidia.com>
2026-09-03 07:38:51 +00:00
Dhiraj Bokde cc4ded2088 feat(helm): split gateway and workspace charts (#2643)
* feat(helm): split gateway and workspace charts

Signed-off-by: Dhiraj Bokde <dbokde@nvidia.com>

* fix(helm): preserve split chart upgrade compatibility

Keep workspace manifests valid after value validation and default legacy reused values to the combined resource topology.

* fix(ci): preserve VM runtime for E2E

The Rust cache restores target/ after VM runtime artifacts are staged,
overwriting target/vm-runtime-compressed before openshell-driver-vm is built.
Stage the compressed runtime outside target and pass that location through
OPENSHELL_VM_RUNTIME_COMPRESSED_DIR so build.rs can embed the supervisor.

Also locate the Helm split-ownership test repository root from the script
path rather than git rev-parse. The test runs in a container where the
GitHub checkout can be owned by a different UID and rejected as dubious
ownership.

Signed-off-by: Dhiraj Bokde <dbokde@nvidia.com>

* fix(ci): install yq for Helm ownership test

The split-chart ownership regression uses yq to inspect rendered YAML,
but the Helm CI container installs only tools declared in mise.
Declare and lock yq so mise install --locked provides the test dependency.

Signed-off-by: Dhiraj Bokde <dbokde@nvidia.com>

---------

Signed-off-by: Dhiraj Bokde <dbokde@nvidia.com>
2026-09-02 00:23:39 +00:00
Drew Newberry 9ca19e6c80 refactor(compute): decouple gateway driver composition (#2823)
* refactor(compute): decouple gateway driver composition

Move first-party composition and VM process ownership into openshell-gateway, leaving openshell-server backend-independent. Update packaging and build references with the new crate, simplify the compiled-driver boundary, and keep the driver-free gateway path buildable with bundled Z3 tooling.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(telemetry): bound compute driver categories

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(core): keep runtime transport generic

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(compute): complete server driver decoupling

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(compute): preserve driver integrations after rebase

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>

* fix(compute): preserve docker tracing after decoupling

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>

* fix(compute): preserve driver behavior after extraction

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(compute): remove MXC policy side channel

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(compute): separate policy delivery from readiness

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>
2026-09-01 21:13:45 +00:00
Russell BryantandJohn Myers 5b925dd8af feat(build): add defaults-without-telemetry feature alias (#2843)
* feat(build): add defaults-without-telemetry feature alias

Cargo cannot subtract a single default feature, so compiling telemetry out
meant `--no-default-features` plus a hand-maintained keep-list of the crate's
other defaults. That keep-list was already wrong for operators: telemetry is
the only default on openshell-server and openshell-driver-vm, but
openshell-sandbox also defaults to `bundled-ca-roots`, so a bare
`--no-default-features` silently swapped the supervisor onto the platform
trust store.

Add a `defaults-without-telemetry` alias to each of the three telemetry-
carrying binary crates, enumerating every default except `telemetry`.
Telemetry-free builds become `--no-default-features --features
defaults-without-telemetry` and stay correct as the default set grows.

The alias is a keep-list, not a switch. Enabling it on top of the defaults
would otherwise produce a telemetry-on binary that reads as telemetry-free, so
each crate root carries a `compile_error!` for the `telemetry` +
`defaults-without-telemetry` combination.

Add `rust:verify:defaults-without-telemetry` to guard both properties: each
alias still equals its crate's defaults minus `telemetry`, and the
mutual-exclusion error is wired up. The additive-misuse check matches on the
`compile_error!` text rather than a nonzero exit code so it cannot pass
vacuously on hosts where openshell-driver-vm fails to build for unrelated
reasons. `rust:verify:telemetry-off` now builds through the alias.

Signed-off-by: Russell Bryant <rbryant@redhat.com>

* fix feature alias for openshell-server

Signed-off-by: Russell Bryant <rbryant@redhat.com>

* fix(ci): run Rust verification in Nix shell

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: Russell Bryant <rbryant@redhat.com>
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-09-01 19:55:42 +00:00
Simon Scatton a4f9c762ce fix(release): handle prerelease tag builds (#3094)
Signed-off-by: Simon Scatton <sscatton@nvidia.com>
2026-09-01 16:12:01 +00:00
Simon Scatton c8f13205e3 ci(release): publish prerelease artifacts (#3093)
Signed-off-by: Simon Scatton <sscatton@nvidia.com>
2026-09-01 15:31:18 +00:00
alangou f7180c0fd6 feat(ci): add Codex Security release qualification (#3087)
* feat(ci): add Codex Security release qualification

Scan cumulative release-train diffs through NVIDIA inference and publish findings to Code Scanning.

Signed-off-by: alangou <alangou@nvidia.com>

* fix(ci): disable package cache for security scan

Prevent cache poisoning in the tag-triggered Codex Security workflow.

Signed-off-by: alangou <alangou@nvidia.com>

* refactor(ci): simplify Codex Security reporting

Remove custom inference cost accounting so the workflow remains focused on scanning and SARIF publication.

Signed-off-by: alangou <alangou@nvidia.com>

---------

Signed-off-by: alangou <alangou@nvidia.com>
2026-09-01 13:57:14 +00:00
Evan Lezar bb70461878 test(e2e): run conformance in gateway lanes (#2925)
* test(e2e): isolate VM-specific smoke assertions

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(e2e): add portable CLI conformance baseline

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* feat(conformance): add standalone CLI runner

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(e2e): run conformance in gateway lanes

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-09-01 12:36:51 +00:00
krishicks 197b41371d fix(dev): harden local cluster and gateway startup (#2993)
* fix(helm): refresh kubeconfig for existing k3d clusters

Docker can recreate the k3d load balancer on a new API port. Start existing
clusters and prefer fresh k3d entries so create does not retain a stale
endpoint.

Signed-off-by: Kris Hicks <khicks@nvidia.com>

* fix(dev): conditionally enable local OTLP export

Probe port 4317 before adding OTLP configuration for the VM, Docker,
and Podman gateway tasks. Document the startup behavior and troubleshooting
for local collector availability.

Signed-off-by: Kris Hicks <khicks@nvidia.com>

---------

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-08-27 17:55:01 +00:00
krishicks f68867b869 feat(gateway): identify gateways in exported traces (#2647)
* feat(gateway): add installation name configuration

Add a first-class operator-assigned gateway name with TOML, CLI, environment,
and Helm configuration surfaces. Local gateways default to openshell, while
Helm defaults to the chart fullname; operators sharing a collector across
namespaces or clusters can set a globally distinct name.

Signed-off-by: Kris Hicks <khicks@nvidia.com>

* feat(gateway): identify gateways in exported traces

Attach the configured gateway installation name and compute driver to the
gateway OpenTelemetry resource so operators can filter traces from multiple
installations that share a collector. Forward the gateway name and OTLP
endpoint to managed external drivers so their distinct service resources carry
the same installation identity.

Keep service.name stable per process type, omit blank resource values, and
leave per-span operation names and request attributes unchanged.

Refs #2507

Signed-off-by: Kris Hicks <khicks@nvidia.com>

---------

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-08-27 17:27:34 +00:00
alangou 37072ee81c feat(build): publish OCI SBOM and provenance attestations (#2836)
* feat(build): embed auditable Rust dependency metadata

Signed-off-by: Adrien Langou <alangou@nvidia.com>

* feat(build): publish OCI SBOM and provenance attestations

Signed-off-by: Adrien Langou <alangou@nvidia.com>

---------

Signed-off-by: Adrien Langou <alangou@nvidia.com>
2026-08-27 14:53:33 +00:00
Evan Lezar 9f88f8ff9b ci: remove rootless podman e2e lane (#2981)
Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-08-27 08:48:44 +00:00
bcd517bbe0 feat(driver-mxc): native Windows MXC compute driver + server wiring (#2721)
* feat(driver): add MXC compute driver for Windows isolation sessions

Introduces the openshell-driver-mxc crate implementing ComputeDriver
backed by Microsoft MXC isolation sessions (Windows only). Wires the
new driver into the server's build_compute_runtime dispatch and adds
the Mxc variant to ComputeDriverKind.

Also adds a local protobuf-src stub (tools/protobuf-src-local) to
unblock Windows builds that lack MSYS2/MinGW, and pins the zig
Windows x64 toolchain in mise.lock.

(cherry picked from commit 4f7012224efb18fbfeb47aa87e0cfd3f036f32f0)
Signed-off-by: Jamie King <jamiek@nvidia.com>

* wip(mxc): checkpoint hung-agent work (recon, policy_map embed, A1 wiring, demo artifacts)

Safety checkpoint of uncommitted work from the background agent run that stalled mid-Step-7. Includes: mxc-driver-recon.md (Step 0.5), policy_map.rs (~876L embedded mapper), A1 policy-threading edits across driver.rs/policy.rs/mxc.rs/compute/mod.rs, and examples/ (demo.yaml + mxc-gateway.toml). Not yet verified to compile end-to-end; to be reorganized into the skill's Step 11 commit sequence.

(cherry picked from commit 38e42c03870be3d10e984a54f17a3b61122ff510)
Signed-off-by: Jamie King <jamiek@nvidia.com>

* test(mxc): fix lifecycle and policy unit-test compile drift

- Bring futures::StreamExt into scope for the watch-stream `.next()` call in
  driver::lifecycle_tests so the negative policy proof test compiles.
- Bind a local `mapper` and drop the unused/deprecated NetworkBinary in the
  embedded-mapper network-policy rejection test.

Signed-off-by: Jamie King <jamiek@nvidia.com>
(cherry picked from commit 039b0baf98735ca672dae52be8d3af2417dc0c1a)
Signed-off-by: Jamie King <jamiek@nvidia.com>

* fix(mxc): downgrade missing sandbox_token to debug log

The gateway mints `sandbox_token` only when a sandbox-JWT issuer is
configured. There is no in-sandbox supervisor on MXC (supervisor-removal
design — D1/D4), so no component ever consumes the token; requiring it
on the driver side blocks the demo's `--disable-tls` smoke gateway with a
spurious `invalid_argument`. Log the absence and proceed instead.

Signed-off-by: Jamie King <jamiek@nvidia.com>
(cherry picked from commit cea209797d0edcb1d152251e748900b0a63cca62)
Signed-off-by: Jamie King <jamiek@nvidia.com>

* fix(mxc): keep sandbox Ready after a successful one-shot agent exec

monitor_exec demoted Ready->Error on exit 0 (reason ExecCompleted), so the positive demo (write hello.txt + exit) landed in Error phase. Keep Ready=True (reason AgentCompleted) on success; only non-zero exits go to ExecFailed. Tighten the positive lifecycle test to assert the terminal condition stays Ready=True/AgentCompleted. Verified live via gateway mock round-trip: phase now Provisioning->Ready with no demotion.

(cherry picked from commit 54ab030f03ca0f83d0050d8dd843b633651684ad)
Signed-off-by: Jamie King <jamiek@nvidia.com>

* feat(mxc): add processContainer backend for default-deny enforcement

Add a backend selector to the MXC driver (isolation_session default | process_container). process_container drives a one-shot AppContainer that is genuinely default-deny: a write to any ungranted path is denied by the OS, unlike isolation_session which is grant-only and cannot deny. The lifecycle forks on the flag - isolation_session keeps provision/start/exec, process_container runs a single ephemeral container via run_oneshot.

Also: run-demo.ps1 gains -Backend and hardens the CLI register/create calls; docs corrected to state isolation_session does NOT deny out-of-policy writes and that the negative proof requires process_container.

Verified end-to-end on a real demo box (gateway -> CLI -> driver -> MXC): in-policy write succeeds, out-of-policy write denied (PermissionDenied), OVERALL: PASS.

(cherry picked from commit c6cde3860bbe1b8edb3147d3e840f6bf0ece32d8)
Signed-off-by: Jamie King <jamiek@nvidia.com>

* refactor(driver-mxc): embed policy mapper as a module; remove standalone crate

Adopt the proto-based mapper (map_to_mxc) as the single source of truth,
embedded in openshell-driver-mxc as a Windows-gated `policy_map` module.
Rewire EmbeddedPolicyMapper to call it directly on the typed SandboxPolicy,
deleting the serde_yaml proto->YAML bridge. Move the CLI to a windows-gated
example and the parity tests into the crate; delete openshell-policy-mapper.

- gate policy_map + seam Windows-only (MXC is Windows-only)
- drop serde_yaml; add dev-deps openshell-policy, clap, anyhow
- normalize mapped paths to Windows form in the seam, in one place
- docs: add driver-mxc to AGENTS.md table; correct design doc section 17 test lane

Signed-off-by: Giedrius Burachas <gburachas@nvidia.com>
(cherry picked from commit f22f9c7a25b9a651c5c5cc73f62fb01c4d6c1a8d)
Signed-off-by: Jamie King <jamiek@nvidia.com>

* feat(driver-mxc): implement lossless split_policy for proxy-delegated egress

Signed-off-by: Giedrius Burachas <gburachas@nvidia.com>
(cherry picked from commit 96d6afa0e2dc6a1d54edd12c34a0ceb0a30dadd0)
Signed-off-by: Jamie King <jamiek@nvidia.com>

* feat(driver-mxc): implement Pattern-C governed-egress split through the policy seam

- split_policy: SocketAddr proxy_redirect (replaces bare port), processcontainer
  containment guard naming MXC M1, version preserved in the trimmed proxy_policy,
  delegation reported as an info loss item
- seam: MappedConfig carries trimmed_policy + proxy_addr; MapCtx.egress selects
  the split path; coarse path unchanged when egress is disabled
- driver: [openshell.drivers.mxc] egress_proxy / egress_proxy_addr config,
  validated at create (isolation_session rejected until M1); lifecycle threads
  the redirect into provision and stores the trimmed policy per sandbox,
  emitting an EgressRedirect platform event
- mxc: optional MxcNetwork block (defaultPolicy=block + proxy) in provision and
  one-shot configs; mock records configs for test assertions
- tests: lossless-invariant suite over all example policies (validate +
  serialize round-trip), split lifecycle proof, M1 rejection; example gains
  --split --proxy-addr writing mxc-config.json / trimmed-policy.yaml /
  loss-report.json

Signed-off-by: Giedrius Burachas <gburachas@nvidia.com>
(cherry picked from commit 34d54ad9f25dc6034c3ba15668555ff0d22cddd8)
Signed-off-by: Jamie King <jamiek@nvidia.com>

* fix(driver-mxc): emit MXC network.proxy as {localhost: port}

Verified against the real wxc-exec 0.6.0-alpha via --dry-run: MXC accepts
only the {localhost: N} proxy shape (the form the design doc specifies)
and rejects {host, port} with a parse error. Schema 0.6.0-alpha can
express only a loopback port, so non-127.0.0.1 redirect addresses are now
rejected: split_policy emits an error loss (no proxy block) and the driver
refuses egress_proxy_addr values off 127.0.0.1. Per-sandbox attribution
must use per-sandbox ports until the schema widens.

Signed-off-by: Giedrius Burachas <gburachas@nvidia.com>
(cherry picked from commit edde8d5434571fd5398409204fcf6862672c0793)
Signed-off-by: Jamie King <jamiek@nvidia.com>

* fix(driver-mxc): serialize isolation_session stop/deprovision as unit variants

Empirical contract finding from the real test lane (build 26300.8553,
wxc-exec 2026-06-10): the stop and deprovision experimental blocks are
unit variants in the wxc-exec schema and must serialize as null; sending
{} is rejected with malformed_request (invalid type: map, expected unit),
while provision/start accept maps. The production invoker, the real-lane
test, the probe script, and the e2e runner all sent {} - the driver could
provision and run an agent but never stop or delete an isolation-session
sandbox against this build. Pinned by a unit test.

Signed-off-by: Giedrius Burachas <gburachas@nvidia.com>
(cherry picked from commit 0df39ca0b22ebc21eb965b2b567a5b2cac26af32)
Signed-off-by: Jamie King <jamiek@nvidia.com>

* test(driver-mxc): add Tier-0 mapper coverage matrix with schema drift guard

Three-quadrant, table-driven matrix (38 tests): mappable fields assert
exact MXC output; every OpenShell field MXC cannot express asserts a loss
item with the expected severity (and seam rejection on error); an empty
policy asserts the restrictive default-deny posture for every MXC knob
OpenShell does not control. The handled_fields_inventory drift guard
serializes a fully-populated policy and compares its YAML keys against
the mapper-handled field lists, so a new openshell-policy field fails the
suite until consciously mapped, delegated, or reported as loss.

Re-exports the policy seam types for integration tests; adds serde_yml,
base64, serde_json as dev-dependencies.

Signed-off-by: Giedrius Burachas <gburachas@nvidia.com>
(cherry picked from commit 91807f984a3b16846e35d6ca0d5ec41057aafa3a)
Signed-off-by: Jamie King <jamiek@nvidia.com>

* feat(driver-mxc): inject agent_env into sandbox process.env

Add MxcComputeConfig.agent_env: each entry is either KEY=VALUE (verbatim) or a bare KEY resolved from the gateway host environment at launch, keeping secrets (e.g. inference API keys) out of the config file. Wire it into the agent process so gateway-launched agents can authenticate to cloud endpoints (process.env was previously hardcoded empty). Unit-tested via resolve_agent_env_passthrough_and_host_lookup.

Also add a gateway-driven cloud-inference (T1) test harness: mxc-inference.toml (agent_env + curl agent), inference.yaml policy, and run-inference-test.ps1 which starts the gateway, creates an isolation_session sandbox, runs an authenticated Nemotron call, and bundles redacted results. Documented agent_env in mxc-gateway.toml. Validated end-to-end on the test box (chat HTTP 200 + completion via the gateway).

(cherry picked from commit 94d9e827b8b77e0af9dc654943ea8e7d8400cc9f)
Signed-off-by: Jamie King <jamiek@nvidia.com>

* test(driver-mxc): avoid unsafe env mutation in resolve_agent_env test

Replace std::env::{set,remove}_var (unsafe + racy under parallel test
execution in edition 2024) with a read-only PATH lookup. Preserves all
three behaviors under test and drops the #[allow(unsafe_code)].

(cherry picked from commit ac5766eeab2db4e8cc6fcd8d8a97809edaf3df30)
Signed-off-by: Jamie King <jamiek@nvidia.com>

* fix(mxc): adapt MXC driver to current GitHub OpenShell API

The MXC driver crate was authored on GitLab against an earlier proto/core
API. Adapt it to the API on GitHub main:

- build_capabilities_response no longer takes supports_interactive_session
- DriverSandboxSpec.gpu (bool) is now resource_requirements; detect GPU via
  effective_driver_gpu_count(driver_gpu_requirements(..))
- DriverSandbox gained a `workspace` field
- SandboxPolicy gained `network_middlewares`: pass it through the proxy split,
  emit a loss item on the coarse MXC path, and account for it in the mapper
  drift-guard test

Verified: cargo check + 75 mock-based tests pass (lib 27, examples 10,
policy_mapper_matrix 38).

Signed-off-by: Jamie King <jamiek@nvidia.com>

* feat(server): wire the MXC compute driver into the gateway on Windows

Register openshell-driver-mxc as the Windows-only in-process compute
backend so compute_driver = "mxc" resolves to a working runtime:

- ComputeRuntime::new_mxc, adapted to the current 11-arg from_driver
- mxc_policy_sink A1 side channel, staged in create_sandbox before dispatch
- mxc_config_from_context loader and the Mxc dispatch arm (Windows
  constructs; other targets return an explicit "Windows-only" error)
- Windows-gated openshell-driver-mxc dependency
- Mxc arms for the telemetry, config-file required-fields, and CLI
  reserved-builtin matches to keep them exhaustive/correct

Verified with cargo check --workspace --features openshell-prover/bundled-z3
on x86_64-pc-windows-msvc, stacked on PR #2496.

Signed-off-by: Jamie King <jamiek@nvidia.com>

* fix(driver-mxc): implement GetGatewayListenerRequirements for #2496 base

Signed-off-by: Jamie King <jamiek@nvidia.com>

* test(driver-mxc): add probe-gated real wxc-exec test lane (no mocks)

- tests/wxc_exec_real.rs: ignored-by-default integration tests against a
  real wxc-exec. Six --dry-run contract tests run wherever the binary
  exists (they caught the network.proxy shape mismatch); enforcement
  tests (processcontainer default-deny positive/negative, isolation
  session lifecycle round trip with a deprovision drop-guard) probe the
  backend and SKIP with a recorded reason where it is not live.
- examples/probe-mxc-host.ps1: classifies a host (OS build, --probe,
  per-backend trial) and emits a JSON capability verdict.
- examples/run-mxc-e2e.ps1 + e2e-policies/: scenario runner generalizing
  run-demo.ps1 (fs-rw, fs-readonly, fs-default-deny-empty,
  network-policy-rejected) with PASS/FAIL/SKIP gating and a stale
  OPENSHELL_MXC_MOCK_WXC guard in real mode.
- tasks/windows.toml: windows:test:mxc-real:x64, windows:e2e:mxc,
  windows:e2e:mxc:mock.

Signed-off-by: Giedrius Burachas <gburachas@nvidia.com>
(cherry picked from commit 49afafe892caded59c4df50a9b652011ad97f41c)
Signed-off-by: Jamie King <jamiek@nvidia.com>

* fix(driver-mxc): address PR review feedback

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>

* docs: defer public MXC documentation

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>

* fix(driver-mxc): build Windows capabilities response

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>

* fix(server): gate in-tree tracing on Windows

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>

* test(windows): fix cross-platform test assumptions

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>

* test(windows): verify process and PEM portably

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>

* test(windows): keep lifecycle command in policy

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>

---------

Signed-off-by: Jamie King <jamiek@nvidia.com>
Signed-off-by: Giedrius Burachas <gburachas@nvidia.com>
Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>
Co-authored-by: Prashant K <pkhodade@nvidia.com>
Co-authored-by: Giedrius Burachas <gburachas@nvidia.com>
Co-authored-by: Shailendra Singh <shailendras@nvidia.com>
Co-authored-by: Drew Newberry <385+drew@users.noreply.github.com>
2026-08-27 07:24:04 +00:00
krishicks d0dfb22baf feat(kubernetes): export driver traces over OTLP (#2958)
Mirror the VM, Podman, and Docker driver tracing setup for Kubernetes.
Export standalone driver spans through OTLP/gRPC as the distinct
openshell-driver-kubernetes service, preserve gateway trace context, record
lifecycle operations and gRPC failures, and flush spans on shutdown.

Kubernetes currently runs in-process when selected as a built-in gateway
driver. Use the temporary server-boundary shim shared with Podman and Docker
so traces retain the shape they will have when Kubernetes moves to a
separate process. Move the common ComputeDriver RPC tracing layer into
openshell-otel to keep all drivers aligned.

Propagate the active W3C context through the controller-reserved Sandbox
annotation and enable Agent Sandbox OTLP export in the local k3s workflow.
This connects asynchronous controller reconciliation spans to the originating
OpenShell create trace.

Expose gateway OTLP configuration through Helm and add an Aspire collector
to the local k3s workflow. Extend helm:k3s:forward with OTLP ingest and trace
UI forwarding for Kubernetes and local container gateway development.

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-08-26 21:22:51 +00:00
krishicks c399342649 feat(dev): unify local Kubernetes gateway workflow (#2914)
Make the local k3s gateway workflow match the Docker and Podman flows by
registering and selecting successful plaintext Skaffold deployments with the
OpenShell CLI. Derive the registration name from the worktree-specific k3d
cluster name so parallel worktrees retain independent gateway metadata.

Add helm:k3s:forward as the standard way to expose the Kubernetes gateway on
localhost:8090, and update the development and debugging guidance to use the
active registered gateway instead of one-off endpoint flags.

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-08-26 14:44:47 +00:00
Mrunal Patel 18ce13b9b1 feat(providers): expose actionable OAuth refresh failures (#2887)
* fix(providers): classify OAuth refresh failures

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* test(providers): add Keycloak refresh e2e lane

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(providers): harden OAuth refresh recovery

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(providers): classify post-mint refresh failures

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(providers): tolerate malformed OAuth subtypes

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

---------

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
2026-08-25 18:00:00 +00:00
alangou fb6610df39 feat(build): embed auditable Rust dependency metadata (#2734)
Signed-off-by: Adrien Langou <alangou@nvidia.com>
2026-08-25 14:36:15 +00:00
Simon Scatton 455883905a fix(python): remove CLI from wheel (#2321)
The Maturin-based wheel packaging was a historical remnant from when the local gateway launch path and OpenShell CLI were coupled in one binary. The gateway and CLI now ship as standalone artifacts, so the Python distribution should contain only the SDK.

Build a single platform-independent setuptools wheel, verify that it cannot contain native code or an openshell entry point, and simplify the release jobs and documentation for SDK-only PyPI installs.

Signed-off-by: Simon Scatton <sscatton@nvidia.com>
2026-08-25 13:20:42 +00:00
Philippe Martin 0a1f246587 feat(sandbox,podman): trust corporate CA for https:// proxies and intercepted TLS (#2512)
* feat(sandbox,podman): trust corporate CA for https:// proxies and intercepted TLS

The corporate proxy chaining only accepted plain http:// proxy URLs, so
operators whose forward proxy terminates TLS with a private corporate CA
had no way to reach it, and TLS-intercepting proxies (mitmproxy, squid
ssl-bump) that re-sign tunneled server certificates broke every upstream
handshake after CONNECT.

The supervisor now accepts https:// proxy URLs: it wraps the connection to
the proxy in TLS before the CONNECT handshake, verifying the proxy
certificate against the built-in Mozilla roots, the system CA bundle, and an
optional operator corporate CA bundle. The upstream dial returns a
Plain/Tls stream enum consumed generically by the relay paths.

The corporate CA is delivered as a driver-supplied command-line argument
(--upstream-proxy-ca-bundle), never an environment variable, matching the
hardened proxy-config model where a sandbox image cannot influence the
operator's egress boundary. It is folded into the sandbox combined trust
bundle (write_ca_files) and the L7 upstream verification store
(build_upstream_client_config) at startup, so intercepted upstream
handshakes succeed and sandbox workloads trust the re-signed certificates.
Configuration is fail-closed: a CA bundle set without a proxy, or an
unreadable or certificate-free file, is fatal rather than silently
weakening the trust boundary.

The shared parse_upstream_proxy_url validator accepts https:// (recording
the scheme so the driver and supervisor agree), keeping the explicit-port
requirement. The Podman driver gains a proxy_ca_bundle operator setting
(TOML, --sandbox-proxy-ca-bundle, OPENSHELL_SANDBOX_PROXY_CA_BUNDLE) that
bind-mounts the host PEM read-only into the sandbox (a CA certificate is not
secret) and points --upstream-proxy-ca-bundle at it, with a create-time
readability check.

The standalone dev gateway task passes OPENSHELL_SANDBOX_PROXY_CA_BUNDLE
through to the generated podman config, so a local gateway can be pointed at
a TLS-intercepting proxy without hand-editing the regenerated TOML.

Refs #1792

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(sandbox): reject CA bundles with valid PEM framing but invalid X.509 DER

The proxy CA bundle validation counted PEM blocks that base64-decoded
successfully, but did not verify the decoded bytes were accepted as
trust anchors by RootCertStore. A bundle with syntactically valid PEM
framing but invalid DER would pass the startup check while contributing
zero usable anchors, causing opaque TLS failures at runtime instead of
a fail-closed startup error.

Validate decoded certificates through RootCertStore::add_parsable_certificates
and reject the bundle unless at least one is accepted.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(sandbox,podman): allow proxy auth without insecure acknowledgement for https:// proxies

For an https:// proxy the Proxy-Authorization credential travels inside
the verified TLS session, so the proxy_auth_allow_insecure
acknowledgement is unnecessary. Previously both http:// and https://
proxies required it, producing a misleading cleartext-risk diagnostic
for a path that is already encrypted.

Skip the requirement when the proxy URL uses https://; the
acknowledgement is still tolerated if set. Updated in both the
supervisor and Podman driver validation paths, with docs and tests.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix: format

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(kubernetes): use truly unsupported scheme in proxy validation test

https:// is now a supported proxy scheme after a13c4dce, so the
unsupported-scheme test must use a genuinely unsupported scheme.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(e2e): sign the https proxy fixture listener cert with a CA

The corporate-proxy E2E fixture served a single `openssl req -x509`
certificate as its TLS listener identity. OpenSSL marks that certificate
`basicConstraints: critical, CA:TRUE`, and rustls refuses a CA
certificate presented as an end-entity certificate (CaUsedAsEndEntity).
The supervisor's TLS handshake with the proxy therefore failed, the
upstream dial errored, and the workload's CONNECT was dropped without a
response, so podman_corporate_proxy_trusts_ca_bundle_for_https_proxy
failed on the approved destination while policy denial still worked.

Generate a corporate CA and a separate listener leaf signed by it, serve
the leaf chain, and publish only the CA as the bundle the supervisor
trusts. This is what an intercepting proxy actually presents, and it
exercises the corporate-CA trust path rather than pinning the listener
certificate itself.

Refs #1792

Signed-off-by: Philippe Martin <phmartin@redhat.com>

---------

Signed-off-by: Philippe Martin <phmartin@redhat.com>
2026-08-24 21:10:25 +00:00
krishicks 6c38646c59 feat(dev): add dedicated gateway:podman task (#2880)
Previously, Podman could be selected through automatic driver detection or with
`mise run gateway -- --driver podman`, but it did not have a dedicated task
like the Docker and VM drivers.

This adds a gateway:podman task and moves the Podman-specific setup into its
own script. The generic gateway task now delegates Podman launches to that
script.

Additionally:

Unlike Docker, which rebuilds and bind-mounts the supervisor binary, Podman
uses a dev-tagged supervisor image that can become stale. The default Podman
supervisor image is therefore rebuilt on each launch.

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-08-21 19:11:19 +00:00
John T. MyersandJohn Myers 679fe4c334 fix(policy): validate the applicable advisor candidate (#2850)
* fix(policy): bind reviews to applicable candidates

Build and validate the exact effective-policy candidate before approval, bind review to live policy/provider/credential inputs, and preserve inspected endpoint contracts during mechanistic expansion.

Closes #2821

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* fix(policy): canonicalize advisor review inputs

Serialize nested protobuf maps in stable key order for proposal review tokens and effective-policy hashes. Narrow reused multi-port endpoint contracts to the denied port so advisor proposals cannot widen binary access. Add regressions for both cases.

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* test(e2e): keep advisor sandbox running

Create the issue 2821 regression sandbox detached with a durable canonical main process so policy denial, approval, and hot-reload checks run before lifecycle exit.

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* fix(policy): apply reviewed draft batches atomically

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <johntmyers@users.noreply.github.com>
2026-08-21 19:05:53 +00:00
Evan Lezar 3be2cd8a29 fix(helm): preflight Agent Sandbox APIs (#2867)
* fix(helm): preflight Agent Sandbox APIs

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(kubernetes): share Agent Sandbox setup

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(e2e): wait for Agent Sandbox CRD status

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* ci(canary): sparse-checkout sandbox helper

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-08-21 14:32:19 +00:00
Drew NewberryandEvan Lezar 40f822906c feat(compute): add standalone first-party drivers (#2822)
* feat(compute): add standalone first-party drivers

Build Docker, Podman, Kubernetes, and VM drivers as external binaries and
exercise each through the public compute-driver API. Keep the external E2E
setup complete at introduction, including VM image selection, Kubernetes
post-renderer isolation, supervisor reuse, and scoped Podman coverage.

External Kubernetes endpoints support shared and managed workspace modes.
Operator mode remains restricted to the in-process driver because gateway
authentication and the driver must share a dynamic namespace allowlist.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci(e2e): run managed and external drivers independently

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(compute): cover external driver socket contract

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* ci(e2e): install bundled Z3 build dependency

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(server): gate in-tree driver tracing

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>
Co-authored-by: Evan Lezar <elezar@nvidia.com>
2026-08-21 02:49:55 +00:00
Drew Newberry ef296806f5 feat(sandbox): add canonical main process (#2726)
* feat(sandbox): add canonical main process

Closes #2710

Persist and supervise one canonical workload per sandbox, attach sandbox connect to its retained session, and make every unexpected main-process exit terminal.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(sandbox): simplify canonical main process contract

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): preserve legacy VM main compatibility

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): preserve main status across driver updates

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): satisfy macOS process lint

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): gate Linux exit acknowledgement publisher

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(sandbox): simplify main process plumbing

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(supervisor): make controlling tty ioctl portable

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(supervisor): initialize canonical process environment

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* perf(supervisor): optimize retained main session

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(sandbox): detach main session on ctrl-c

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): use explicit main detach keys

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(dev): atomically stage Docker supervisor

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(test): align Docker main environment assertion

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(sdk): expose canonical main process fields

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-08-20 22:18:23 +00:00
Derek Carr 59479f492a feat(k8s): add namespace-per-workspace support (RFC 0011 Phase 3) (#2656)
* feat(k8s): add namespace-per-workspace support (RFC 0011 Phase 3)

Implement three workspace namespace modes for the Kubernetes compute
driver: shared (default, preserves current single-namespace behavior),
managed (auto-creates/deletes namespaces per workspace), and operator
(pre-provisioned namespaces with dynamic discovery via label selector
or drop-in allowlist file).

Key changes:
- WorkspaceMode enum and namespace resolution in driver config
- Managed namespace lifecycle with ServiceAccount and OpenShift SCC
  annotation propagation
- Cluster-wide sandbox CR watchers for managed/operator modes
- NamespaceValidator (Exact/Prefix/Allowlist) for SA token auth
- Workspace-aware credential secret storage
- Helm ClusterRole for multi-namespace RBAC
- Gateway config, architecture, and reference docs

Signed-off-by: Derek Carr <decarr@redhat.com>

* test(k8s): add e2e tests for workspace namespace modes

Add end-to-end tests for managed and operator workspace modes
introduced in RFC 0011 Phase 3. The managed mode tests verify
namespace creation with correct labels, ServiceAccount provisioning,
sandbox CR placement, and namespace survival with remaining sandboxes.
The operator mode tests verify rejection of unlabeled and nonexistent
namespaces. The positive operator path (sandbox in labeled namespace)
is known to fail due to an RBAC gap and will be addressed separately.

Also fixes Helm 4 compatibility: move SPDX license headers inside
conditional guards in 8 chart templates to prevent empty comment-only
documents, and fix a trailing whitespace trimmer in clusterrole.yaml
that concatenated the license header with apiVersion.

Adds cleanup sweep in with-kube-gateway.sh to remove managed and
operator namespaces before Helm uninstall, and mise tasks for running
each mode independently.

Signed-off-by: Derek Carr <decarr@redhat.com>

* feat(k8s): add operator namespace label watcher

Spawn a background kube::runtime::watcher in the K8s driver that
watches namespaces matching the configured label selector and populates
the OperatorNamespaceAllowlist at runtime. The driver owns the
allowlist and exposes its Arc so the server can share the same set with
the SA token authenticator.

create_sandbox now gates pod creation on the allowlist in operator
mode — workspaces whose namespace is not yet labeled are rejected at
resource render time rather than silently proceeding. Workspace
lifecycle itself is unaffected; only sandbox (resource) creation is
gated.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(k8s): harden operator mode and address review findings

Close the fail-open gap in operator mode when only
operator_namespace_file is configured: the allowlist is now created
unconditionally in operator mode (fail-closed from startup).

Implement the namespace file watcher using the notify crate, following
the TLS hot-reload pattern (parent-directory watch, 1s debounce,
ConfigMap symlink-swap safe). The file format is a JSON array of
namespace name strings.

Additional fixes from the 10-reviewer audit:
- Change allowlist rejection from InvalidArgument to FailedPrecondition
  so callers know the request may succeed later once the namespace is
  provisioned.
- NamespaceValidator::Allowlist now holds the OperatorNamespaceAllowlist
  newtype instead of a raw Arc<RwLock<BTreeSet>>, eliminating silent
  denial on RwLock poison.
- Verify LABEL_MANAGED_BY and LABEL_GATEWAY_ID ownership before
  deleting a managed namespace.
- Replace fixed 5s sleep in operator e2e test with a 30s poll loop.
- Add Helm validation for workspaceMode values.
- Fix Helm README type column and description for operator fields.
- Add insert/remove methods to OperatorNamespaceAllowlist; label
  watcher now uses them instead of reaching through shared().
- Reject configs with both operator_namespace_label and
  operator_namespace_file set.

Signed-off-by: Derek Carr <decarr@redhat.com>

* feat(k8s): add workspace-level compute driver RPCs and harden RBAC

Decouple namespace lifecycle from sandbox lifecycle by adding
EnsureWorkspace/DeleteWorkspace RPCs to the ComputeDriver service.
Namespace creation now happens before credential storage and namespace
deletion happens on workspace delete, fixing credential storage in
managed workspace mode.

- Add EnsureWorkspace and DeleteWorkspace proto RPCs with
  implementations across all compute drivers (K8s managed delegates to
  ensure_namespace/delete_namespace_if_empty; others no-op)
- Wire ensure_workspace into provider create/update/refresh paths so
  the namespace exists before the credential driver writes secrets
- Wire delete_workspace into workspace deletion for cleanup
- Remove delete_namespace_if_empty from sandbox deletion path
- Scope ClusterRole secrets access to non-shared workspace modes
- Add TODO for TLS cert hot-reload in sandbox gRPC client
- Harden e2e tests with control-plane sandbox resolution assertions
- Fix docker image save --platform flag for OCI index manifests

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(k8s): address re-review findings and add test coverage

- Use server-side apply for TLS secret sync (fixes second sandbox
  creation failure when TLS is enabled)
- Scope gateway-ID label selector unconditionally across all workspace
  modes (fixes operator reads/watches/deletes seeing foreign sandboxes)
- Validate operator allowlist in EnsureWorkspace and DeleteWorkspace
  RPCs (prevents credential writes to namespaces outside the allowlist)
- Extend ClusterRole secrets patch+delete to all non-shared modes with
  credential driver enabled (fixes operator credential storage RBAC)
- Validate namespace ownership on 409 conflict in ensure_namespace
  (prevents adopting unowned namespaces in managed mode)
- Replace delete_namespace_if_empty with unconditional delete_namespace
  letting Kubernetes cascade cleanup (fixes stuck terminating CRs)
- Strengthen NetworkPolicy TODO to cover both managed and operator modes
- Extract selector and ownership logic into testable free functions
- Add unit tests for gateway-ID selectors and namespace ownership
- Add Helm ClusterRole RBAC tests for operator credential driver

Signed-off-by: Derek Carr <decarr@redhat.com>

* ci(k8s): add workspace managed and operator mode e2e to CI

Wire the existing e2e:kubernetes:workspace-managed and
e2e:kubernetes:workspace-operator mise tasks into the branch-e2e
workflow so they run alongside the other core Kubernetes e2e suites.
Both are gated by run_core_e2e and included in the Core E2E result
gate.

Signed-off-by: Derek Carr <decarr@redhat.com>

* test(k8s): add e2e tests for workspace namespace modes

Add 7 new e2e tests covering workspace namespace lifecycle, TLS secret
copying, ownership conflict detection, DNS-1123 validation, operator
namespace preservation, and dynamic label watcher behavior. Fix async
sandbox deletion race condition in existing tests by polling sandbox
list instead of asserting immediately after delete.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(k8s): grant secrets/patch unconditionally and backfill gateway-id labels

Address two review findings:

1. RBAC: server-side apply (PATCH) is used for TLS secret sync in
   multi-namespace modes, but the ClusterRole only granted patch when
   the kubernetes-secrets credential driver was enabled. Grant patch
   unconditionally for non-shared modes since TLS sync always needs it;
   keep delete gated on the credential driver.

2. Upgrade safety: the new gateway-id label selector would orphan
   legacy Sandbox CRs that predate its introduction. Add a startup
   backfill in shared mode that patches any managed Sandbox CR missing
   the gateway-id label before the driver begins serving requests.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(k8s): address workspace namespace review findings

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(k8s): address follow-up review findings

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(k8s): preserve workspace lookup after rebase

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(helm): allow managed secret creation

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(k8s): stop pods in workspace namespace

Signed-off-by: Derek Carr <decarr@redhat.com>

* test(k8s): scope pod deletion check to v1alpha1

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(k8s): address workspace namespace review findings

Signed-off-by: Derek Carr <decarr@redhat.com>

---------

Signed-off-by: Derek Carr <decarr@redhat.com>
2026-08-14 21:22:41 +00:00
krishicks ae40cf6744 fix(gateway): respect OPENSHELL_BIND_ADDRESS in dev task (#2756)
This is convenient when you want to run a local gateway pointed at a remote
compute driver so that the supervisor can reach across the network to the
gateway which is listening on 0.0.0.0.

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-08-14 21:04:16 +00:00
Piotr Mlocek bdabb54cb3 fix(security): authenticate extension services (#2638)
* fix(security): authenticate extension services

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* feat(extension-core): verify gateway JWTs

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* refactor(extension-core): keep inbound verification external

This should become an extension SDK package.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(security): harden the extension authentication contract

Follow-up hardening on the alpha extension authentication mechanism.

Claim contract:
- Extension tokens carry an explicit `typ` of `openshell-ext+jwt`. They
  share a signing key with sandbox-to-gateway admission tokens and were
  otherwise separated by audience alone, so a verifier that neglects to
  check `aud` could accept a gateway credential. The header is a second,
  independent discriminator.
- Publish OIDC-shaped discovery at `/.well-known/openid-configuration`
  so a service configured with only the gateway URL can learn the exact
  expected issuer and the JWKS location. It is shaped, not compliant:
  `issuer` is the gateway identity, not the serving URL.

Audience agreement:
- `MiddlewareManifest` and `InterceptorManifest` gain `expected_audience`.
  The audience is otherwise configured independently on each side of the
  boundary, where a mismatch surfaces only as an opaque authentication
  failure on every call. OpenShell now compares the two and fails at
  startup. An empty field keeps the check off for existing services.

Compatibility:
- Add `allow_insecure_transport` per registration. Enabling gateway JWT
  signing previously made any plaintext endpoint a hard startup failure,
  including the endpoint form used in our own documentation. The opt-out
  attaches no credential, is refused by the gateway if a supervisor asks
  for one, and warns at every startup.
- Make the transport requirement kind-aware. A middleware endpoint must
  be reachable from every sandbox supervisor, so only interceptors may
  use a gateway-local Unix socket.

Credential lifecycle:
- Replace the process-global slot map with a supervisor-owned
  `ExtensionCredentialStore` shared explicitly across the gateway
  connections the supervisor opens, removing test-order coupling.
- Rotate only when a credential is missing or has passed four fifths of
  its lifetime. Configuration polling ran every ten seconds against
  fifteen-minute credentials, so each poll re-ran gateway effective-policy
  resolution and re-minted the gateway token.
- Bound credential minting per sandbox, since each request resolves the
  caller's effective policy.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs: record alpha extension authentication in RFC appendices

Restore the RFC 0009 and 0010 bodies to their accepted text and move
every extension-authentication update into appendices instead. An RFC
records a decision at a point in time; superseding detail belongs
alongside it rather than rewritten into it.

RFC 0009's appendix carries the shared contract: claims, authorization,
key distribution, the `allow_insecure_transport` replacement for the
body's `allow_insecure`, and residual risks. RFC 0010's records only
what differs for interceptors and links to it. The existing
protocol-extensions appendix, which parked the phase 2 transport
question, now points forward to what was built.

Also document the audience handshake, the discovery endpoint, the
`typ` requirement, and `jti` replay guidance in the extensibility and
gateway configuration pages, and correct the middleware transport
guidance: middleware endpoints must be reachable from sandbox
supervisors, so Unix sockets are not an option there.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* refactor(extension-core): abstract extension server trust

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(core): update middleware manifest example

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(extension-auth): preserve unsigned gateway compatibility

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(extension-auth): reject cross-domain token replay

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

---------

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-08-14 20:34:10 +00:00