Commit Graph
143 Commits
Author SHA1 Message Date
araza008 8f22fe84f6 ci(windows): run MXC host probe, WebSocket agent, and OpenClaw forwar… (#3826)
* ci(windows): run MXC host probe, WebSocket agent, and OpenClaw forward examples checks

Add separate hosted CI tasks for the MXC host probe, WebSocket agent,
and OpenClaw forward examples. Use mock workloads to verify gateway,
CLI, driver, and sandbox lifecycle wiring without requiring wxc-exec.

Document that mock passes do not validate forwarding or MXC enforcement.

* fix(tests): enhance environment isolation for WebSocket and OpenClaw mock tests

* fix(mxc): enhance OpenClaw mock validation to require 'Ready' sandbox state

* fix(mxc): restore native process helper in WebSocket example

Restore Invoke-NativeCaptured for the PowerShell argument regression test and delegate CLI execution through it while preserving explicit gateway endpoint selection.

Signed-off-by: Akber Raza <akberr@nvidia.com>

---------

Signed-off-by: Akber Raza <akberr@nvidia.com>
2026-09-30 14:13:26 -07:00
Prekshi Vyas d46a814141 ci(windows): exercise MXC credential, audit, and aggregate E2E flows (#3787)
* ci(windows): exercise MXC provider credential example

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* ci(windows): exercise MXC OCSF audit example

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* ci(windows): enable aggregate MXC example E2E

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* ci(windows): select native MXC mock target

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* fix(windows): isolate MXC example CI harnesses

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

---------

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>
2026-09-29 19:50:53 -07:00
Prekshi Vyas 7ba7a39d09 ci(windows): exercise MXC inference demos with mock API (#3780)
* ci(windows): exercise MXC Ollama demo with mock API

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* ci(windows): cover both MXC inference demos

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* fix(mxc): preserve executable extension resolution

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* test(mxc): use absolute PowerShell in lifecycle checks

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* test(mxc): make lifecycle write probes deterministic

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* test(mxc): assert stable lifecycle completion

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

---------

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>
2026-09-29 13:39:23 -07:00
Prekshi Vyas a69d0319f2 test(windows): define GB300 MXC qualification contract (NVBug 6643699) (#3471)
* test(windows): define GB300 MXC qualification contract

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* test(mxc): harden GB300 qualification provenance

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* test(windows): package portable GB300 qualification

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* test(windows): honor external qualification checkout

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* revert: remove portable GB300 packaging

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* test(mxc): bind GB300 evidence to exact inputs

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

---------

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>
2026-09-22 10:52:21 -07:00
Prekshi Vyas fb2980e077 fix(windows): restore MXC qualification and cold-start readiness (#3468)
* fix(windows): restore MXC qualification gates

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* fix(mxc): poll target for full readiness budget

---------

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>
2026-09-18 23:01:55 +05:30
Drew Newberry 49b4f0eb7f feat(mxc): add UI policy, credentials, relay lifecycle, and proxy auth
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-17 12:06:06 -07:00
Mrunal Patel b799fccb8b fix(auth): harden OIDC trust root retrieval (#3332)
* fix(auth): harden OIDC trust root retrieval

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(e2e): pass OIDC HTTP acknowledgement value

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

---------

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
2026-09-15 17:00:45 +00:00
26f2f96393 feat(mxc): add Windows host proxy for MXC sandbox network egress (#3163)
* Implement Windows host proxy integration and update dependencies for OpenShell

* Update README and gateway config to clarify egress proxy address handling and allocation

* Refactor ProxyIdentityMode to return Result for static_binary and add tests for binary path and SHA256 hash

* Enhance platform_hosts_path for Windows to use SystemRoot and improve error handling for hosts file reading

* Refactor FileFingerprint to use Option for mtime and ctime, simplifying metadata handling

* Add conditional compilation for Windows host module

* add unit tests for OPA policy evaluation and identity handling

* remove openshell-supervisor-network from unsupported driver package test exclusion list

* feat(mxc): enable host proxy TLS state generation

Generate per-sandbox TLS state for the MXC host proxy so HTTPS L7 enforcement can use the same MITM path as Linux. Grant generated CA material to the MXC process and inject standard trust env vars, while matching Linux behavior by disabling TLS termination on CA setup failure and relying on proxy fail-closed handling.

* fix(docs): remove outdated notes on governed egress from docs

* fix(tests): update TLS environment variable paths to use temporary directory

* fix(examples): make run-mxc-e2e harness correct and orphan-free

The MXC e2e harness never actually exercised the fs scenarios: it started
the gateway once and patched agent_command per scenario AFTERWARDS, so the
running gateway kept launching the default demo agent (not shipped in the
kit) and every fs scenario failed with CreateProcessW error:2. It also
scored on the `sandbox create` exit code (non-zero due to the harmless
interactive attach), wrote sandbox records to the persistent gateway DB
(leaving orphans that collided on later runs), and its deny scenarios never
proved denial.

Changes:
- Start a FRESH gateway per scenario so each scenario's agent_command is
  actually loaded (root cause of CreateProcessW error:2).
- Score by on-disk artifact / expected outcome, not `sandbox create` exit.
- Real deny assertions: a control write to a granted path must succeed
  (proves the agent ran) while the denied write must be absent. fs-empty
  probes an ungranted out-of-share path (share_dir is mapped rw by design).
- Run the gateway on an ephemeral in-memory DB (sqlite::memory:) so the
  harness never writes to the persistent store and cannot leave orphan
  sandbox records; also use unique per-run sandbox names + pre-delete.
- Fix the process_container probe: use a real cwd + absolute cmd.exe
  (canonical wxc-exec does not expand %TEMP% -> 0x8007010B).
- Fix summary counts (@() so a single FAIL is counted and exit is non-zero).

Verified PASS=4 FAIL=0 on 7F203-MXC-003 (no BaseContainer velocity keys)
using a canonical wxc-exec build (AppContainer fallback).

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(e2e): probe timeout is milliseconds (10ms->30000ms)

MXC process.timeout is wall-clock ms (wire.rs). The 10 value meant 10ms,
which the base-container tier (7F203-MXC-001/.181) enforced strictly and
timed the probe out. AppContainer path (.18/-003) happened to slip under
it. Bump to 30000ms so the process_container preflight probe is reliable
across both tiers.

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc): use native paths in real runtime probes

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc): make processcontainer work with mxc-latest-released wxc-exec

Three fixes to support the release wxc-exec binary (BaseContainer dispatcher)
in addition to mxc-fixes-env-vars:

1. Seed process env from host (driver.rs)
   ProcessContainer starts with a completely blank environment -- no PATH,
   SystemRoot, or anything.  Seed the process env from the gateway host
   environment so the agent binary can locate DLLs and run.  Skip internal
   Windows drive-letter variables (keys starting with '=') which cause
   CreateProcessW to return ERROR_ENVVAR_NOT_FOUND.  User agent_env entries
   and TLS CA vars are applied as overrides on top of the host env.

2. Remove TLS readonly_paths grant (driver.rs)
   The release wxc-exec (BaseContainer dispatcher) requires write-DAC
   permission on every path in readonly_paths to set up AppContainer ACLs.
   Adding the proxy's temp TLS directory caused a DACL error and exit -1.
   The CA cert paths remain available to the agent via TLS env vars.

3. Remove allowedHosts from network JSON (mxc.rs)
   The release wxc-exec rejects network.allowedHosts / network.blockedHosts
   on Windows with "not yet supported".  Removed the loopback exemption
   attempt (127.0.0.1, ::1, localhost) from the network section.
   Intra-container loopback works natively in the release binary without
   it -- the spawner can connect to the server at 127.0.0.1:22000 directly.

Additional changes:
- mxc-ws-agent.rs: add relay-debug.txt error capture and relay-ready.txt
  marker for reliable timing of host client connections.
- mxc-ws-gateway.toml: debug = true for JSON config dump during diagnosis.
- run-ws-agent-test.ps1: default port changed to 17670 (gateway default);
  relay-ready.txt polling before ws-echo to avoid connecting before the
  spawner has established the proxy bridge.

Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(e2e): address CodeRabbit review on run-mxc-e2e.ps1 (MR !46)

Four robustness/correctness fixes from CodeRabbit:

1. Start-Gw: kill the spawned gateway before the "did not start within 30s"
   throw. If the process is alive but never binds the port, $gw is not yet
   assigned in the caller, so the finally block cannot reap it -> orphan
   gateway holding the port for the next run.

2. create-fail scoring: a non-zero `sandbox create` exit alone is not proof
   of a policy rejection (gateway-registration/transport/fixture errors also
   exit non-zero and would false-pass). PASS now requires a genuine
   rejection signal (network / invalid_argument / network_policies) AND that
   it is not an infrastructure failure; other non-zero exits go to FAIL with
   output captured.

3. deny scenarios (ControlTarget path): snapshot the deny target AFTER
   Wait-File lands the control artifact, so a late denied write (enforcement
   regression racing the control write) can no longer be recorded as PASS.

4. -KeepRunning: break out of the scenario loop after the first scenario so
   a later scenario does not start a second gateway on the same port
   (previously a reliable port collision instead of a usable debug mode).

Re-verified PASS=4 FAIL=0 on both boxes (7F203-MXC-001 base-container and
7F203-MXC-003 AppContainer fallback); network-policy-rejected correctly
scores as "policy rejection".

Signed-off-by: Akber Raza <akberr@nvidia.com>

* feat(mxc-e2e): collect run-mxc-e2e output into a results bundle

Mirror the sibling run-*.ps1 scripts by collecting every run's logs into a
timestamped results-e2e-<stamp>\ folder and zipping it. The bundle contains the
console transcript, per-scenario gateway stdout/stderr, the exact TOML rendered
for each scenario, the policy fixture used, and a summary.txt with the verdict
table.

Per-scenario gateway logs now land in gateway.<scenario>.log/.err.log inside the
bundle instead of a single fixed gateway.e2e.log in the script directory.

Wrap pre-flight, mode setup, scenario definitions, and the scenario loop in a
single try/catch/finally so the finally always writes the summary, stops the
transcript, and zips the bundle -- even on a pre-flight failure. The existing
per-scenario gateway-cleanup try/finally stays nested inside. All scenario
logic, scoring rules, and comments are preserved.

Signed-off-by: Prashant S Khodade <pkhodade@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc-e2e): address CodeRabbit review on run-mxc-e2e.ps1

- Require -Scenario when -KeepRunning: the loop breaks after the first
  scenario, so a full-suite run would execute only one scenario yet still
  report the suite as PASS. Fail fast so a partial run can't be mislabeled
  complete.
- Start-Transcript now runs inside the guarded try block with a
  $transcriptStarted flag; Stop-Transcript is only called when it actually
  started, so a Start-Transcript failure still yields the results bundle.
- Wrap the -Scenario filter in @() so a single exact match stays an array
  (reliable .Count and a proper array for the scenario loop on PS 5.1).

Signed-off-by: Prashant S Khodade <pkhodade@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(examples): pass gateway config via OPENSHELL_GATEWAY_CONFIG for spaced paths

Start-Process -ArgumentList does not quote array elements, so launching the
gateway with a bare --config <path> token split on any space in the install
path (e.g. C:\Users\First Last\...), and clap rejected the fragment with
'unrecognized subcommand'. Every MXC example launcher that started the gateway
hit this when the kit was unzipped under a path containing a space.

Pass the config path through the OPENSHELL_GATEWAY_CONFIG env var (which the
gateway already reads via clap) and drop the --config token. Env vars carry
spaces safely.

Affected: run-ocsf-audit, run-mxc-e2e, run-demo, run-inference-test,
run-ollama-test. run-mtls-test was not affected (its launch passes no config
path). Root-caused and fix-verified on 7F203-MXC-003 from a spaced path.

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(run-mxc-e2e): improve scoring logic and enhance command execution handling

* fix(mxc): reconcile proxy support after rebase

Restore the proxy-enabled OCSF audit example removed by 13185f6e now that the host CONNECT proxy is present. Adapt the proxy lifecycle test to the target branch's DriverSandboxSpec policy delivery contract.

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc): grant sandbox access to proxy CA

- Share the per-sandbox public CA bundle with the AppContainer
- Add real wxc-exec HTTPS proxy coverage and document trust isolation

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc): reject unsupported network middleware

- Reject middleware-bearing MXC policies before sandbox lifecycle begins
- Guard host proxy startup and document the unsupported registry path
- Add mapper, lifecycle, and host proxy regression coverage

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc): reconcile host proxy with main

- remove obsolete inference routing from the host proxy adapter
- use the workspace AWS-LC provider in host-proxy tests
- adapt the forward-proxy test to ProxyIdentityMode

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(build): switch to bundled Z3 for Windows MSVC builds

* fix(mxc): reconcile host proxy after rebase

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Akber Raza <akberr@nvidia.com>
Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
Signed-off-by: Prashant S Khodade <pkhodade@nvidia.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Jamie King <jamiek@nvidia.com>
Co-authored-by: Shailendra Singh <shailendras@nvidia.com>
Co-authored-by: Prashant Khodade <pkhodade@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
2026-09-14 22:58:49 +00:00
krishicks 5b9daab935 fix(ci): restore mise run ci on macOS (#3294)
- Replace BSD-incompatible in-place sed calls with portable temp-file rewrites.
- Remove test-only shell interception and capture generated gateway config
  directly.
- Allow parity tests to use supplied supervisor binaries without resolving a
  Linux target.
- Normalize temporary-directory paths and use portable RPM config installation.
- Set a valid setuptools-scm version for Python protobuf generation in Jujutsu
  checkouts.

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-09-12 00:17:15 +00:00
Piotr Mlocek b92620e838 ci(rust): reject stale Cargo lockfiles (#3227)
* ci(rust): reject stale Cargo lockfiles

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(ci): clarify lockfile validation policy

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* refactor(ci): structure and test Cargo lockfile validation

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* chore(ci): remove lockfile validator regression tests

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(rust): complete locked validation and lint examples

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(rust): check lockfile diffs after validation

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(build): shorten lockfile validation notes

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(rust): skip lockfile check after failures

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

---------

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-09-11 08:52:41 +00:00
Jesse JaggarsandDrew Newberry 02b664bb0d refactor(config): normalize and enforce gateway schema v2 (#2814)
* refactor(config): normalize compute driver field names

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* refactor(config): introduce canonical gateway fields

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* refactor(config): enforce gateway schema version 2

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): preserve compute driver runtime guarantees

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): address schema v2 review regressions

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): complete schema v2 migration safeguards

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(config): expand schema v2 regression coverage

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(config): add schema v2 parity manifest

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): correct parity manifest inventory

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* docs(config): record schema v2 intentional changes

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* docs(config): disposition schema v2 parity gaps

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): add dual schema parity harness

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): establish compute lifecycle parity baseline

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): preserve gateway option compatibility

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): record gateway option parity

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* docs(config): close gateway-wide parity gaps

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(podman): apply configured pids limit

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): validate Podman option parity

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): add Kubernetes option parity harness

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): record Kubernetes option parity

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): disposition VM parity lanes

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): add external driver parity lane

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(e2e): preserve external driver pull policy

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): attest parity artifacts and launches

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): require clean parity build sources

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): bind parity runtime artifacts

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(e2e): use isolated supervisor tags

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(e2e): qualify parity image tags

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(e2e): serve parity supervisor locally

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): isolate parity podman services

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): harden parity evidence provenance

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): pin parity sandbox artifacts

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): attest parity runtime inputs

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): bind parity runtime evidence

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): record compute boundary parity

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): disposition cross-cutting parity lanes

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(packaging): preflight gateway config upgrades

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): preserve rebase integration guarantees

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(ci): isolate temporary git signing config

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): update remaining schema v2 consumers

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(ci): provide e2fs tools to VM tests

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): align preflight with gateway startup

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(vm): preserve rootfs tar configuration

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* chore(config): adopt duration unit constructors

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(packaging): preflight RPM gateway config

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): address driver review findings

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(e2e): require fresh semantic parity evidence

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(docker): update tests for renamed sandbox label

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(gateway): preserve selective driver coverage after rebase

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
2026-09-11 05:00:24 +00:00
Drew Newberry 38f2aef930 feat(gateway): support selective compute driver builds (#3118)
* feat(gateway): support selective compute driver builds

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(gateway): support selective Windows MXC builds

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-11 00:36:09 +00:00
Piotr Mlocek ddc8bba967 ci(windows): add Windows MSVC CI jobs (#2738)
* fix(ci): preserve Windows Rust build cache

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(ci): invalidate empty Windows caches

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* perf(ci): cache Windows builds with sccache

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(windows): restore target directory caching

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* perf(ci): use prebuilt Z3 on Windows

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* perf(ci): layer sccache on Windows target cache

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(windows): split PR checks from main validation

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(windows): separate checks builds and cache seeding

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(ci): simplify Windows build dependency

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(ci): rely on Windows job dependency status

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(ci): use valid opt-in Windows ARM runner

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(windows): keep ARM64 validation local

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(ci): install Clippy for Windows validation

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(windows): focus platform lint coverage

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(licenses): explain bzip2 allowance

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(windows): simplify workflow name

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(windows): allow async platform stub

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(network): make file fingerprints portable

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(windows): lint supported deliverables

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(windows): allow platform-gated lint

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* chore(ci): align Windows cache action with main

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(ci): align Windows validation with prerequisites

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(windows): pin Rust toolchain action

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(ci): use enterprise-approved Windows actions

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(windows): restore strict MSVC validation

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(windows): run Rust tests with nextest

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(ci): normalize nextest lock provenance

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(windows): add native arm64 validation

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(ci): lock nextest for Windows ARM64

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(windows): resolve duplicate MXC authentication method

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* test(conformance): use native absolute paths on Windows

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(windows): address MSVC review feedback

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(windows): simplify cache key names

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(ci): isolate Windows Rust toolchains for stable caches

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(ci): configure Rustup home in runner setup

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(windows): surface sccache server write diagnostics

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(windows): remove temporary cache diagnostics

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(windows): address review feedback

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(windows): reconcile merged driver capabilities

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(deps): preserve AWS-LC-only lockfile

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* chore(deps): allow z3 prebuilt TLS wrapper

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

---------

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-09-10 23:53:01 +00:00
Piotr Mlocek ce25acca5a fix(build): honor Cargo target directory when staging binaries (#3262)
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-09-10 23:19:13 +00:00
Piotr Mlocek a0814443f1 feat(docs): publish versioned release snapshots (#3149)
* feat(docs): add version availability labels

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(docs): use supported Python for sync

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* feat(docs): publish versioned docs from releases

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(docs): format dev version label

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* chore(docs): upgrade Fern CLI to 5.112.0

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(docs): make release publishing monotonic

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(docs): preserve snapshot release identity

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(fern): document versioned publishing

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* test(docs): cover explicit snapshot rollback

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* refactor(docs): use Fern refs for versions

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(docs): bundle components for ref versions

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* revert(docs): keep complete version copies

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(docs): publish latest and dev channels

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

---------

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-09-09 23:34:34 +00:00
John T. Myers f4dc6be4b2 refactor(inference): remove managed inference routes (#3195)
* refactor(inference): remove managed inference routes

Closes #3172

Remove the inference route control plane, inference.local data path, built-in router crate, and SDK surface. Move inference workloads to explicitly imported provider profiles and native endpoints, with migration cleanup and updated tests and documentation.

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(policy): preserve alternate upstream isolation

Restore the provider policy activation guard so legacy OpenAI and Anthropic providers configured for alternate base URLs do not grant egress to the built-in public vendor endpoints.

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-09-09 18:47:22 +00:00
Gaizka Menendez 8af79a7f4b fix(podman): resolve macOS Podman socket dynamically (#3135)
* docs(podman): document macOS socket path mismatch and dynamic lookup

On macOS, Homebrew-installed Podman does not create the default socket
path that the Podman driver probes. Document the OPENSHELL_PODMAN_SOCKET
override and the podman machine inspect lookup in both the compute
drivers reference and the debug-openshell-cluster skill.

Fixes #1690

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* fix(podman): resolve macOS Podman socket dynamically

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* chore: restore debug skill file

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* chore: drop legacy debug skill path

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* fix(podman): trim unrelated e2e changes

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* fix(e2e): harden shell array expansion

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* chore: remove unrelated skill note

* ci: retrigger checks

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

---------

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>
2026-09-09 14:10:59 +00:00
alangou 3693b32841 ci(trivy): add artifact and PR configuration scans (#3185)
* ci(trivy): add artifact and PR configuration scans

Signed-off-by: Adrien Langou <alangou@nvidia.com>

* fix(ci): harden Trivy gate detection and finding diff

Signed-off-by: Adrien Langou <alangou@nvidia.com>

* feat(ci): scan released artifacts in release pipelines

Signed-off-by: Adrien Langou <alangou@nvidia.com>

* fix(ci): harden and simplify Trivy scans

Signed-off-by: Adrien Langou <alangou@nvidia.com>

* fix(ci): consolidate Trivy reports and prevent collisions

Signed-off-by: Adrien Langou <alangou@nvidia.com>

---------

Signed-off-by: Adrien Langou <alangou@nvidia.com>
2026-09-09 13:58:06 +00:00
Piotr Mlocek d7cb6e456d fix(dev): inherit non-expiring sandbox JWT in local gateway scripts (#2636)
* fix(dev): inherit non-expiring sandbox JWT in local gateway scripts

The local gateway launcher scripts hardcode gateway_jwt.ttl_secs = 3600,
which overrides the non-expiring default introduced in #1721. Local
Docker, Podman, and VM sandboxes are still unrecoverable when the gateway
is down longer than that TTL: the on-disk token expires and only the
Kubernetes ServiceAccount path can rebootstrap, so the supervisor
crash-loops on policy fetch and the sandbox never leaves Provisioning.

Drop the override so local drivers inherit the default. gateway.sh also
serves the kubernetes driver, which is a shared deployment and must keep
a positive TTL, so it now emits ttl_secs only for that driver.

The e2e regression test added in #1721 does not catch this because the
e2e harness uses its own configs, which already set ttl_secs = 0.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(dev): expand sandbox JWT TTL in gateway config

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

---------

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-09-04 17:16:44 +00:00
alangou 64a858dade fix(ci): restore Codex Security scan execution (#3124)
* refactor(ci): resolve Codex Security range in Python

Signed-off-by: Adrien Langou <alangou@nvidia.com>

* fix(ci): allow unprivileged userns for Codex sandbox

Signed-off-by: Adrien Langou <alangou@nvidia.com>

---------

Signed-off-by: Adrien Langou <alangou@nvidia.com>
2026-09-03 07:38:51 +00:00
Drew Newberry 9ca19e6c80 refactor(compute): decouple gateway driver composition (#2823)
* refactor(compute): decouple gateway driver composition

Move first-party composition and VM process ownership into openshell-gateway, leaving openshell-server backend-independent. Update packaging and build references with the new crate, simplify the compiled-driver boundary, and keep the driver-free gateway path buildable with bundled Z3 tooling.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(telemetry): bound compute driver categories

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(core): keep runtime transport generic

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(compute): complete server driver decoupling

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(compute): preserve driver integrations after rebase

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>

* fix(compute): preserve docker tracing after decoupling

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>

* fix(compute): preserve driver behavior after extraction

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(compute): remove MXC policy side channel

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(compute): separate policy delivery from readiness

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>
2026-09-01 21:13:45 +00:00
Russell BryantandJohn Myers 5b925dd8af feat(build): add defaults-without-telemetry feature alias (#2843)
* feat(build): add defaults-without-telemetry feature alias

Cargo cannot subtract a single default feature, so compiling telemetry out
meant `--no-default-features` plus a hand-maintained keep-list of the crate's
other defaults. That keep-list was already wrong for operators: telemetry is
the only default on openshell-server and openshell-driver-vm, but
openshell-sandbox also defaults to `bundled-ca-roots`, so a bare
`--no-default-features` silently swapped the supervisor onto the platform
trust store.

Add a `defaults-without-telemetry` alias to each of the three telemetry-
carrying binary crates, enumerating every default except `telemetry`.
Telemetry-free builds become `--no-default-features --features
defaults-without-telemetry` and stay correct as the default set grows.

The alias is a keep-list, not a switch. Enabling it on top of the defaults
would otherwise produce a telemetry-on binary that reads as telemetry-free, so
each crate root carries a `compile_error!` for the `telemetry` +
`defaults-without-telemetry` combination.

Add `rust:verify:defaults-without-telemetry` to guard both properties: each
alias still equals its crate's defaults minus `telemetry`, and the
mutual-exclusion error is wired up. The additive-misuse check matches on the
`compile_error!` text rather than a nonzero exit code so it cannot pass
vacuously on hosts where openshell-driver-vm fails to build for unrelated
reasons. `rust:verify:telemetry-off` now builds through the alias.

Signed-off-by: Russell Bryant <rbryant@redhat.com>

* fix feature alias for openshell-server

Signed-off-by: Russell Bryant <rbryant@redhat.com>

* fix(ci): run Rust verification in Nix shell

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: Russell Bryant <rbryant@redhat.com>
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-09-01 19:55:42 +00:00
Simon Scatton a4f9c762ce fix(release): handle prerelease tag builds (#3094)
Signed-off-by: Simon Scatton <sscatton@nvidia.com>
2026-09-01 16:12:01 +00:00
Simon Scatton c8f13205e3 ci(release): publish prerelease artifacts (#3093)
Signed-off-by: Simon Scatton <sscatton@nvidia.com>
2026-09-01 15:31:18 +00:00
alangou f7180c0fd6 feat(ci): add Codex Security release qualification (#3087)
* feat(ci): add Codex Security release qualification

Scan cumulative release-train diffs through NVIDIA inference and publish findings to Code Scanning.

Signed-off-by: alangou <alangou@nvidia.com>

* fix(ci): disable package cache for security scan

Prevent cache poisoning in the tag-triggered Codex Security workflow.

Signed-off-by: alangou <alangou@nvidia.com>

* refactor(ci): simplify Codex Security reporting

Remove custom inference cost accounting so the workflow remains focused on scanning and SARIF publication.

Signed-off-by: alangou <alangou@nvidia.com>

---------

Signed-off-by: alangou <alangou@nvidia.com>
2026-09-01 13:57:14 +00:00
krishicks 197b41371d fix(dev): harden local cluster and gateway startup (#2993)
* fix(helm): refresh kubeconfig for existing k3d clusters

Docker can recreate the k3d load balancer on a new API port. Start existing
clusters and prefer fresh k3d entries so create does not retain a stale
endpoint.

Signed-off-by: Kris Hicks <khicks@nvidia.com>

* fix(dev): conditionally enable local OTLP export

Probe port 4317 before adding OTLP configuration for the VM, Docker,
and Podman gateway tasks. Document the startup behavior and troubleshooting
for local collector availability.

Signed-off-by: Kris Hicks <khicks@nvidia.com>

---------

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-08-27 17:55:01 +00:00
krishicks f68867b869 feat(gateway): identify gateways in exported traces (#2647)
* feat(gateway): add installation name configuration

Add a first-class operator-assigned gateway name with TOML, CLI, environment,
and Helm configuration surfaces. Local gateways default to openshell, while
Helm defaults to the chart fullname; operators sharing a collector across
namespaces or clusters can set a globally distinct name.

Signed-off-by: Kris Hicks <khicks@nvidia.com>

* feat(gateway): identify gateways in exported traces

Attach the configured gateway installation name and compute driver to the
gateway OpenTelemetry resource so operators can filter traces from multiple
installations that share a collector. Forward the gateway name and OTLP
endpoint to managed external drivers so their distinct service resources carry
the same installation identity.

Keep service.name stable per process type, omit blank resource values, and
leave per-span operation names and request attributes unchanged.

Refs #2507

Signed-off-by: Kris Hicks <khicks@nvidia.com>

---------

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-08-27 17:27:34 +00:00
alangou 37072ee81c feat(build): publish OCI SBOM and provenance attestations (#2836)
* feat(build): embed auditable Rust dependency metadata

Signed-off-by: Adrien Langou <alangou@nvidia.com>

* feat(build): publish OCI SBOM and provenance attestations

Signed-off-by: Adrien Langou <alangou@nvidia.com>

---------

Signed-off-by: Adrien Langou <alangou@nvidia.com>
2026-08-27 14:53:33 +00:00
krishicks d0dfb22baf feat(kubernetes): export driver traces over OTLP (#2958)
Mirror the VM, Podman, and Docker driver tracing setup for Kubernetes.
Export standalone driver spans through OTLP/gRPC as the distinct
openshell-driver-kubernetes service, preserve gateway trace context, record
lifecycle operations and gRPC failures, and flush spans on shutdown.

Kubernetes currently runs in-process when selected as a built-in gateway
driver. Use the temporary server-boundary shim shared with Podman and Docker
so traces retain the shape they will have when Kubernetes moves to a
separate process. Move the common ComputeDriver RPC tracing layer into
openshell-otel to keep all drivers aligned.

Propagate the active W3C context through the controller-reserved Sandbox
annotation and enable Agent Sandbox OTLP export in the local k3s workflow.
This connects asynchronous controller reconciliation spans to the originating
OpenShell create trace.

Expose gateway OTLP configuration through Helm and add an Aspire collector
to the local k3s workflow. Extend helm:k3s:forward with OTLP ingest and trace
UI forwarding for Kubernetes and local container gateway development.

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-08-26 21:22:51 +00:00
krishicks c399342649 feat(dev): unify local Kubernetes gateway workflow (#2914)
Make the local k3s gateway workflow match the Docker and Podman flows by
registering and selecting successful plaintext Skaffold deployments with the
OpenShell CLI. Derive the registration name from the worktree-specific k3d
cluster name so parallel worktrees retain independent gateway metadata.

Add helm:k3s:forward as the standard way to expose the Kubernetes gateway on
localhost:8090, and update the development and debugging guidance to use the
active registered gateway instead of one-off endpoint flags.

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-08-26 14:44:47 +00:00
alangou fb6610df39 feat(build): embed auditable Rust dependency metadata (#2734)
Signed-off-by: Adrien Langou <alangou@nvidia.com>
2026-08-25 14:36:15 +00:00
Simon Scatton 455883905a fix(python): remove CLI from wheel (#2321)
The Maturin-based wheel packaging was a historical remnant from when the local gateway launch path and OpenShell CLI were coupled in one binary. The gateway and CLI now ship as standalone artifacts, so the Python distribution should contain only the SDK.

Build a single platform-independent setuptools wheel, verify that it cannot contain native code or an openshell entry point, and simplify the release jobs and documentation for SDK-only PyPI installs.

Signed-off-by: Simon Scatton <sscatton@nvidia.com>
2026-08-25 13:20:42 +00:00
Philippe Martin 0a1f246587 feat(sandbox,podman): trust corporate CA for https:// proxies and intercepted TLS (#2512)
* feat(sandbox,podman): trust corporate CA for https:// proxies and intercepted TLS

The corporate proxy chaining only accepted plain http:// proxy URLs, so
operators whose forward proxy terminates TLS with a private corporate CA
had no way to reach it, and TLS-intercepting proxies (mitmproxy, squid
ssl-bump) that re-sign tunneled server certificates broke every upstream
handshake after CONNECT.

The supervisor now accepts https:// proxy URLs: it wraps the connection to
the proxy in TLS before the CONNECT handshake, verifying the proxy
certificate against the built-in Mozilla roots, the system CA bundle, and an
optional operator corporate CA bundle. The upstream dial returns a
Plain/Tls stream enum consumed generically by the relay paths.

The corporate CA is delivered as a driver-supplied command-line argument
(--upstream-proxy-ca-bundle), never an environment variable, matching the
hardened proxy-config model where a sandbox image cannot influence the
operator's egress boundary. It is folded into the sandbox combined trust
bundle (write_ca_files) and the L7 upstream verification store
(build_upstream_client_config) at startup, so intercepted upstream
handshakes succeed and sandbox workloads trust the re-signed certificates.
Configuration is fail-closed: a CA bundle set without a proxy, or an
unreadable or certificate-free file, is fatal rather than silently
weakening the trust boundary.

The shared parse_upstream_proxy_url validator accepts https:// (recording
the scheme so the driver and supervisor agree), keeping the explicit-port
requirement. The Podman driver gains a proxy_ca_bundle operator setting
(TOML, --sandbox-proxy-ca-bundle, OPENSHELL_SANDBOX_PROXY_CA_BUNDLE) that
bind-mounts the host PEM read-only into the sandbox (a CA certificate is not
secret) and points --upstream-proxy-ca-bundle at it, with a create-time
readability check.

The standalone dev gateway task passes OPENSHELL_SANDBOX_PROXY_CA_BUNDLE
through to the generated podman config, so a local gateway can be pointed at
a TLS-intercepting proxy without hand-editing the regenerated TOML.

Refs #1792

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(sandbox): reject CA bundles with valid PEM framing but invalid X.509 DER

The proxy CA bundle validation counted PEM blocks that base64-decoded
successfully, but did not verify the decoded bytes were accepted as
trust anchors by RootCertStore. A bundle with syntactically valid PEM
framing but invalid DER would pass the startup check while contributing
zero usable anchors, causing opaque TLS failures at runtime instead of
a fail-closed startup error.

Validate decoded certificates through RootCertStore::add_parsable_certificates
and reject the bundle unless at least one is accepted.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(sandbox,podman): allow proxy auth without insecure acknowledgement for https:// proxies

For an https:// proxy the Proxy-Authorization credential travels inside
the verified TLS session, so the proxy_auth_allow_insecure
acknowledgement is unnecessary. Previously both http:// and https://
proxies required it, producing a misleading cleartext-risk diagnostic
for a path that is already encrypted.

Skip the requirement when the proxy URL uses https://; the
acknowledgement is still tolerated if set. Updated in both the
supervisor and Podman driver validation paths, with docs and tests.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix: format

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(kubernetes): use truly unsupported scheme in proxy validation test

https:// is now a supported proxy scheme after a13c4dce, so the
unsupported-scheme test must use a genuinely unsupported scheme.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(e2e): sign the https proxy fixture listener cert with a CA

The corporate-proxy E2E fixture served a single `openssl req -x509`
certificate as its TLS listener identity. OpenSSL marks that certificate
`basicConstraints: critical, CA:TRUE`, and rustls refuses a CA
certificate presented as an end-entity certificate (CaUsedAsEndEntity).
The supervisor's TLS handshake with the proxy therefore failed, the
upstream dial errored, and the workload's CONNECT was dropped without a
response, so podman_corporate_proxy_trusts_ca_bundle_for_https_proxy
failed on the approved destination while policy denial still worked.

Generate a corporate CA and a separate listener leaf signed by it, serve
the leaf chain, and publish only the CA as the bundle the supervisor
trusts. This is what an intercepting proxy actually presents, and it
exercises the corporate-CA trust path rather than pinning the listener
certificate itself.

Refs #1792

Signed-off-by: Philippe Martin <phmartin@redhat.com>

---------

Signed-off-by: Philippe Martin <phmartin@redhat.com>
2026-08-24 21:10:25 +00:00
krishicks 6c38646c59 feat(dev): add dedicated gateway:podman task (#2880)
Previously, Podman could be selected through automatic driver detection or with
`mise run gateway -- --driver podman`, but it did not have a dedicated task
like the Docker and VM drivers.

This adds a gateway:podman task and moves the Podman-specific setup into its
own script. The generic gateway task now delegates Podman launches to that
script.

Additionally:

Unlike Docker, which rebuilds and bind-mounts the supervisor binary, Podman
uses a dev-tagged supervisor image that can become stale. The default Podman
supervisor image is therefore rebuilt on each launch.

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-08-21 19:11:19 +00:00
Drew Newberry ef296806f5 feat(sandbox): add canonical main process (#2726)
* feat(sandbox): add canonical main process

Closes #2710

Persist and supervise one canonical workload per sandbox, attach sandbox connect to its retained session, and make every unexpected main-process exit terminal.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(sandbox): simplify canonical main process contract

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): preserve legacy VM main compatibility

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): preserve main status across driver updates

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): satisfy macOS process lint

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): gate Linux exit acknowledgement publisher

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(sandbox): simplify main process plumbing

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(supervisor): make controlling tty ioctl portable

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(supervisor): initialize canonical process environment

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* perf(supervisor): optimize retained main session

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(sandbox): detach main session on ctrl-c

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): use explicit main detach keys

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(dev): atomically stage Docker supervisor

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(test): align Docker main environment assertion

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(sdk): expose canonical main process fields

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-08-20 22:18:23 +00:00
Derek Carr 59479f492a feat(k8s): add namespace-per-workspace support (RFC 0011 Phase 3) (#2656)
* feat(k8s): add namespace-per-workspace support (RFC 0011 Phase 3)

Implement three workspace namespace modes for the Kubernetes compute
driver: shared (default, preserves current single-namespace behavior),
managed (auto-creates/deletes namespaces per workspace), and operator
(pre-provisioned namespaces with dynamic discovery via label selector
or drop-in allowlist file).

Key changes:
- WorkspaceMode enum and namespace resolution in driver config
- Managed namespace lifecycle with ServiceAccount and OpenShift SCC
  annotation propagation
- Cluster-wide sandbox CR watchers for managed/operator modes
- NamespaceValidator (Exact/Prefix/Allowlist) for SA token auth
- Workspace-aware credential secret storage
- Helm ClusterRole for multi-namespace RBAC
- Gateway config, architecture, and reference docs

Signed-off-by: Derek Carr <decarr@redhat.com>

* test(k8s): add e2e tests for workspace namespace modes

Add end-to-end tests for managed and operator workspace modes
introduced in RFC 0011 Phase 3. The managed mode tests verify
namespace creation with correct labels, ServiceAccount provisioning,
sandbox CR placement, and namespace survival with remaining sandboxes.
The operator mode tests verify rejection of unlabeled and nonexistent
namespaces. The positive operator path (sandbox in labeled namespace)
is known to fail due to an RBAC gap and will be addressed separately.

Also fixes Helm 4 compatibility: move SPDX license headers inside
conditional guards in 8 chart templates to prevent empty comment-only
documents, and fix a trailing whitespace trimmer in clusterrole.yaml
that concatenated the license header with apiVersion.

Adds cleanup sweep in with-kube-gateway.sh to remove managed and
operator namespaces before Helm uninstall, and mise tasks for running
each mode independently.

Signed-off-by: Derek Carr <decarr@redhat.com>

* feat(k8s): add operator namespace label watcher

Spawn a background kube::runtime::watcher in the K8s driver that
watches namespaces matching the configured label selector and populates
the OperatorNamespaceAllowlist at runtime. The driver owns the
allowlist and exposes its Arc so the server can share the same set with
the SA token authenticator.

create_sandbox now gates pod creation on the allowlist in operator
mode — workspaces whose namespace is not yet labeled are rejected at
resource render time rather than silently proceeding. Workspace
lifecycle itself is unaffected; only sandbox (resource) creation is
gated.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(k8s): harden operator mode and address review findings

Close the fail-open gap in operator mode when only
operator_namespace_file is configured: the allowlist is now created
unconditionally in operator mode (fail-closed from startup).

Implement the namespace file watcher using the notify crate, following
the TLS hot-reload pattern (parent-directory watch, 1s debounce,
ConfigMap symlink-swap safe). The file format is a JSON array of
namespace name strings.

Additional fixes from the 10-reviewer audit:
- Change allowlist rejection from InvalidArgument to FailedPrecondition
  so callers know the request may succeed later once the namespace is
  provisioned.
- NamespaceValidator::Allowlist now holds the OperatorNamespaceAllowlist
  newtype instead of a raw Arc<RwLock<BTreeSet>>, eliminating silent
  denial on RwLock poison.
- Verify LABEL_MANAGED_BY and LABEL_GATEWAY_ID ownership before
  deleting a managed namespace.
- Replace fixed 5s sleep in operator e2e test with a 30s poll loop.
- Add Helm validation for workspaceMode values.
- Fix Helm README type column and description for operator fields.
- Add insert/remove methods to OperatorNamespaceAllowlist; label
  watcher now uses them instead of reaching through shared().
- Reject configs with both operator_namespace_label and
  operator_namespace_file set.

Signed-off-by: Derek Carr <decarr@redhat.com>

* feat(k8s): add workspace-level compute driver RPCs and harden RBAC

Decouple namespace lifecycle from sandbox lifecycle by adding
EnsureWorkspace/DeleteWorkspace RPCs to the ComputeDriver service.
Namespace creation now happens before credential storage and namespace
deletion happens on workspace delete, fixing credential storage in
managed workspace mode.

- Add EnsureWorkspace and DeleteWorkspace proto RPCs with
  implementations across all compute drivers (K8s managed delegates to
  ensure_namespace/delete_namespace_if_empty; others no-op)
- Wire ensure_workspace into provider create/update/refresh paths so
  the namespace exists before the credential driver writes secrets
- Wire delete_workspace into workspace deletion for cleanup
- Remove delete_namespace_if_empty from sandbox deletion path
- Scope ClusterRole secrets access to non-shared workspace modes
- Add TODO for TLS cert hot-reload in sandbox gRPC client
- Harden e2e tests with control-plane sandbox resolution assertions
- Fix docker image save --platform flag for OCI index manifests

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(k8s): address re-review findings and add test coverage

- Use server-side apply for TLS secret sync (fixes second sandbox
  creation failure when TLS is enabled)
- Scope gateway-ID label selector unconditionally across all workspace
  modes (fixes operator reads/watches/deletes seeing foreign sandboxes)
- Validate operator allowlist in EnsureWorkspace and DeleteWorkspace
  RPCs (prevents credential writes to namespaces outside the allowlist)
- Extend ClusterRole secrets patch+delete to all non-shared modes with
  credential driver enabled (fixes operator credential storage RBAC)
- Validate namespace ownership on 409 conflict in ensure_namespace
  (prevents adopting unowned namespaces in managed mode)
- Replace delete_namespace_if_empty with unconditional delete_namespace
  letting Kubernetes cascade cleanup (fixes stuck terminating CRs)
- Strengthen NetworkPolicy TODO to cover both managed and operator modes
- Extract selector and ownership logic into testable free functions
- Add unit tests for gateway-ID selectors and namespace ownership
- Add Helm ClusterRole RBAC tests for operator credential driver

Signed-off-by: Derek Carr <decarr@redhat.com>

* ci(k8s): add workspace managed and operator mode e2e to CI

Wire the existing e2e:kubernetes:workspace-managed and
e2e:kubernetes:workspace-operator mise tasks into the branch-e2e
workflow so they run alongside the other core Kubernetes e2e suites.
Both are gated by run_core_e2e and included in the Core E2E result
gate.

Signed-off-by: Derek Carr <decarr@redhat.com>

* test(k8s): add e2e tests for workspace namespace modes

Add 7 new e2e tests covering workspace namespace lifecycle, TLS secret
copying, ownership conflict detection, DNS-1123 validation, operator
namespace preservation, and dynamic label watcher behavior. Fix async
sandbox deletion race condition in existing tests by polling sandbox
list instead of asserting immediately after delete.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(k8s): grant secrets/patch unconditionally and backfill gateway-id labels

Address two review findings:

1. RBAC: server-side apply (PATCH) is used for TLS secret sync in
   multi-namespace modes, but the ClusterRole only granted patch when
   the kubernetes-secrets credential driver was enabled. Grant patch
   unconditionally for non-shared modes since TLS sync always needs it;
   keep delete gated on the credential driver.

2. Upgrade safety: the new gateway-id label selector would orphan
   legacy Sandbox CRs that predate its introduction. Add a startup
   backfill in shared mode that patches any managed Sandbox CR missing
   the gateway-id label before the driver begins serving requests.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(k8s): address workspace namespace review findings

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(k8s): address follow-up review findings

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(k8s): preserve workspace lookup after rebase

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(helm): allow managed secret creation

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(k8s): stop pods in workspace namespace

Signed-off-by: Derek Carr <decarr@redhat.com>

* test(k8s): scope pod deletion check to v1alpha1

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(k8s): address workspace namespace review findings

Signed-off-by: Derek Carr <decarr@redhat.com>

---------

Signed-off-by: Derek Carr <decarr@redhat.com>
2026-08-14 21:22:41 +00:00
krishicks ae40cf6744 fix(gateway): respect OPENSHELL_BIND_ADDRESS in dev task (#2756)
This is convenient when you want to run a local gateway pointed at a remote
compute driver so that the supervisor can reach across the network to the
gateway which is listening on 0.0.0.0.

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-08-14 21:04:16 +00:00
Drew Newberry f12f3ef8d5 fix(macos): restore Homebrew sandbox callbacks (#2739)
* fix(macos): restore Docker gateway callbacks

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(gateway): reuse reachable primary callback listener

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-08-14 18:38:20 +00:00
Max DubrinskyandDrew Newberry 35fb27ef14 feat(sdk): add TypeScript SDK (@nvidia/openshell-sdk) (#2122)
* feat(sdk): add TypeScript SDK (@nvidia/openshell-sdk)

First native, per-language SDK for the OpenShell gateway: a thin, idiomatic
TypeScript client over proto-generated gRPC stubs (connect-es), no FFI. Covers
the v0.1 surface — sandbox lifecycle (create/get/list/delete + waitReady/
waitDeleted), health, and streamed exec.

- sdk/typescript/: package, client/transport/errors, protoc + protoc-gen-es
  codegen (gen/ gitignored, absorbed into dist/ at build), committed lockfile.
- tasks/typescript.toml: sdk:ts install/proto/typecheck/build/ci/publish;
  sdk:ts:typecheck wired into `check`; sdk-typescript job in branch-checks
  (typecheck, build, and a --dry-run publish that validates the release path).
- Enforce SPDX headers on .ts/.tsx/.mts/.cts (skip node_modules and gen/);
  back-fill docs/_components/jsx.d.ts and fern/components/CustomFooter.tsx.
- release.py gains an npm version format; release-tag.yml publishes to
  GitHub Packages on tag, stamping the version (0.0.0 placeholder in git);
  prerelease builds publish under the `next` dist-tag, not `latest`.

Ships as @nvidia/openshell-sdk on GitHub Packages pre-GA; public npm
(@openshell/sdk) follows at GA with an unchanged public API.

Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com>

* chore(sdk): adopt TypeScript 6, tidy @types/node range

- typescript ^5.7.2 -> ^6.0.3 (6.0 is now `latest`; the old caret capped at 5.x)
- @types/node ^24.0.0 -> ^24 (same range, tidier)

No source changes; codegen, typecheck, and build pass on 6.0.3. Verified the
emitted d.ts still type-check for downstream consumers on TypeScript 5.0.4
through 5.9.3, so this does not raise the SDK's consumer TS floor.

Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com>

* refactor(sdk): group operations under a composable SandboxClient

Reshape the client from flat methods (createSandbox, listSandboxes, exec) to a
scoped SandboxClient reached as `client.sandbox.create/get/list/delete/exec`
(+ waitReady/waitDeleted), mirroring the CLI's noun-verb model and the Python
SDK's SandboxClient.

SandboxClient is also usable standalone via SandboxClient.connect();
OpenShellClient composes it over a single shared transport, so future
service/provider clients reuse one connection. health() stays top-level as a
gateway call. No behavior change; types are unchanged.

Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com>

* chore(sdk): generate TypeScript SDK stubs with buf

Replace the protoc gen.sh with `buf generate` + buf.gen.yaml. `buf`
(@bufbuild/buf) is a package devDependency and self-compiles the protos, so
the TS SDK no longer depends on the mise-pinned protoc; it drives the same
connect-es plugin. Generation stays limited to the client-surface closure
(openshell/sandbox/datamodel) via the input paths.

Output is byte-identical to the previous protoc + protoc-gen-es pipeline. Lays
the groundwork for a shared buf.yaml (lint/breaking/LSP) as a follow-up.

Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com>

* build(proto): add repo-level buf module with lint

Declare proto/ as a single buf v2 module in a root buf.yaml so buf
generate, lint, breaking, and the editor LSP resolve imports the same
way. Lint uses STANDARD with six documented exceptions for deviations
the current protos intentionally make: the flat proto/ layout with
nested packages (DIRECTORY_SAME_PACKAGE, PACKAGE_DIRECTORY_MATCH) and
the established API shape with unsuffixed services and reused
request/response messages (RPC_REQUEST_RESPONSE_UNIQUE,
RPC_REQUEST_STANDARD_NAME, RPC_RESPONSE_STANDARD_NAME, SERVICE_SUFFIX).
Every other STANDARD rule now enforces on future protos. Breaking uses
FILE.

Code generation stays package-scoped in sdk/typescript/buf.gen.yaml
since it binds to that package's connect-es plugin and output dir;
its inputs are unchanged and regeneration is byte-identical.

Wire the check in via a proto:lint mise task that runs buf from the
SDK devDependencies. It is a dependency of both sdk:ts:ci (so the
TypeScript SDK CI job enforces it) and the top-level lint aggregate
(so local pre-commit covers it).

Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com>

* chore(sdk): publish as unscoped openshell-sdk on public npm

Rename the package from @nvidia/openshell-sdk to the unscoped
openshell-sdk and target public npm (registry.npmjs.org) instead of
GitHub Packages. GitHub Packages requires a scope matching the owning
org, and the @openshell scope is blocked by an unrelated existing
package, so an unscoped name on public npm is the lowest-friction
distribution path and needs no org approval.

Rework the release-tag publish job to auth against registry.npmjs.org
with NPM_TOKEN (the job now only needs packages: read to pull the CI
image). Update the README install instructions and usage imports.

Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com>

* chore(sdk): publish @nvidia/openshell-sdk to GitHub Packages

Revert the unscoped-name switch. GitHub Packages only accepts scoped
names matching the owning org, so shipping there first (which needs no
external npm org or NPM_TOKEN, just the repo's GITHUB_TOKEN) requires
the @nvidia scope. Keeping the @nvidia/openshell-sdk name also lets a
later public-npm release use the same install specifier, so adding
public npm becomes a second publish step rather than a rename.

Restore the GitHub Packages publish auth in the release-tag job and the
scoped install instructions in the README (keeping the buf codegen
note).

Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com>

* feat(sdk-ts): add streaming exec, forward, ssh, provider, and config methods

Grow SandboxClient to the surface the first two consumers need. execStream
yields stdout/stderr chunks as they arrive and exec now drains it, keeping its
buffered ExecResult and signature unchanged. execInteractive is the TTY + stdin
transport primitive (start-first framing, output/write/resize/close/done, no
terminal glue). forward binds a local TCP listener that tunnels each accepted
connection into the sandbox for the process lifetime, minting and revoking a
per-socket SSH session token around a forwardTcp bidi. Adds createSshSession /
revokeSshSession, attach/detach/listProviders, and getConfig / setPolicy /
setSetting (sandbox-scoped, network-policy-only, with an optional wait poll).

Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com>

* build(sdk-ts): add Biome and Vitest tooling

The TypeScript SDK had no formatter or linter and no test runner. Add Biome
(format + lint, generated src/gen excluded) enforcing 2-space indent, single
quotes, semicolons, and a 120-column width, and reformat the existing
hand-written sources accordingly. Add Vitest for unit tests. Wire sdk:ts:format,
sdk:ts:lint, and sdk:ts:test mise tasks into the fmt/lint aggregates, the root
test suite, and sdk:ts:ci so they run in CI.

Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com>

* test(sdk-ts): cover the sandbox surface with in-memory transport tests

Exercise SandboxClient against an in-memory OpenShell service built with
createRouterTransport: request assembly and id resolution, u64/int64 rendered as
strings, enum lowercasing, fromConnect code mapping, the exec/execStream drain
plus a backward-compat check on exec, execInteractive start-first ordering and
done resolution, and a forward() byte relay against a loopback echo with close()
teardown.

Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com>

* docs(sdk-ts): document the new surface and connect/upload/download boundaries

Document execStream, execInteractive, forward, ssh sessions, providers, and
config/policy in the SDK README, and record the intentional boundaries:
interactive connect / PTY ownership, upload/download (no file-transfer RPC), and
detached forwards stay out of scope. Note the Biome/Vitest dev commands.

Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com>

* feat(sdk-ts): support mTLS client authentication

Add clientCert and clientKey to ConnectOptions so the SDK can
authenticate to the default local gateway, which uses mTLS user
authentication. Without a client certificate and key the SDK could
verify the server but never authenticate the caller, so it could not
connect to the standard Docker, VM, Homebrew, or Linux-package gateway.

Validate the pair as both-or-neither and pass cert and key through to
the Node TLS options for https gateways. The h2c path is unchanged.

Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com>

* chore(sdk-ts): drop the demo script and its tsx dependency

Remove src/demo.ts, the demo npm script, the tsx devDependency, and the
tsconfig build exclude for the demo. The demo was never part of the
published package, and dropping it also removes the only place that
logged part of an SSH session token.

Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com>

* fix(sdk-ts)!: harden exec streaming, waits, SSH, and forwarding

Address review feedback on the sandbox surface.

- Make the streamed command exit code observable from idiomatic
  for-await: the terminal exit is now an in-band ExecStreamEvent
  ({ type: 'exit', exitCode }) rather than the async generator return
  value, which for-await discards. A stream that ends without an exit
  event now throws instead of reporting success.
- Bound waitReady, waitDeleted, and the setPolicy wait by their timeout:
  each poll RPC carries a per-iteration deadline and the waits accept an
  AbortSignal, so a stalled call can no longer leave a wait pending
  forever. Add waitTimeoutSecs to SetPolicyOptions.
- Validate the CreateSshSession response against the proto charset and
  range contract before returning it or using its token, since the
  values feed an OpenSSH ProxyCommand.
- Respect socket backpressure when relaying forwarded responses: pause
  reading the gRPC stream when the local socket buffer is full and
  resume on drain so memory stays bounded.
- Expose create-time sandbox policy: add policy and an advanced rawSpec
  passthrough to SandboxSpec so the safety boundary is expressible at
  creation and new spec fields do not require an SDK change.

BREAKING CHANGE: execStream and the interactive exec output now yield a
terminal { type: 'exit', exitCode } event; consumers iterating the
stream must handle that arm. The exit code is no longer the async
generator return value.

Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com>

* feat(sdk-ts): export error contract, enum unions, and caller cancellation

Address Tier-1 review feedback on the TypeScript SDK public surface (PR #2122).

- Errors: export SdkError and SdkErrorCode so callers can use instanceof and
  exhaustively switch on .code. fromConnect preserves the originating
  ConnectError as .cause and its status as .connectCode, maps Aborted to a new
  'aborted' code for optimistic-concurrency conflicts, and maps Canceled and
  DeadlineExceeded to 'canceled'. errorCode() behavior is unchanged.
- Enums: replace the string-typed phase, status, scope, and policySource fields
  with lowercase literal unions (SandboxPhaseName, HealthStatus,
  SettingScopeName, PolicySourceName) backed by exhaustive Record maps. The
  unions are a hand-maintained mirror of the generated proto enums; a new drift
  test pins each literal to its generated member name.
- Cancellation: accept an optional AbortSignal on exec, execInteractive, and
  forward, threaded into both sandbox resolution and the streaming RPC. forward
  tears down its local listener on abort.

Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com>

* feat(sdk-ts): add raw escape hatch for uncurated gateway RPCs

The curated sub-clients reduce proto messages to ergonomic subsets (for
example get() drops created_at_ms, the full spec, conditions, runtime
endpoints, and current_policy_version), and not every gateway RPC has a
typed helper yet. Rather than ship methods that exist but throw, expose a
generated client for the full surface.

OpenShellClient.raw and SandboxClient.raw are generated clients covering
every gateway RPC, returning the verbatim wire messages so proto
distinctions the curated types smooth over are preserved. .transport
exposes the shared connection for building extra clients over one socket.
Generated request/response types are published at the new
@nvidia/openshell-sdk/raw subpath. Curated methods stay the default;
raw is the always-available floor.

Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com>

* fix(sdk-ts): address review feedback on exec, forward, and auth transport

Settle exec's `done` promise before yielding the exit event so a consumer
that breaks on exit no longer leaves it pending forever, and give it a lone
rejection handler plus a finally-settle so a stream error or early abandon
can never surface as an unhandled rejection or a hang.

Attach an 'error' listener to each accepted forward socket synchronously,
before forwardConnection awaits CreateSshSession; a peer reset in that
window previously emitted an unhandled 'error' and crashed the process.

Reject ambiguous or unsafe transport configs at buildTransport: oidcToken
and edgeToken together (silently OIDC-only), and any auth token sent over
plaintext http:// to a non-loopback host unless allowInsecureAuth is set.

Wrap versionPin so a non-u64 expectedResourceVersion raises
SdkError('invalid_config') instead of a raw BigInt SyntaxError, and raise
the Node engine floor to >=20.3 for AbortSignal.any().

Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com>

* fix(sdk-ts): address review feedback

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(sdk-ts): defer published sdk guide

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
2026-08-13 20:50:38 +00:00
krishicks 245fe27589 fix(dev): separate Podman Machine loopback listeners (#2725)
On macOS, bind the standalone Podman gateway to IPv6 loopback while
registering localhost as the TLS endpoint. This keeps IPv4 loopback
available for the callback-only listener, matching the e2e fix in
commit 4cb77a9.

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-08-12 22:45:15 +00:00
2f96c53b8c feat(gateway,cli): windows compilation support (#2496)
* chore(windows): gate Unix-only workspace code for MSVC

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* feat(windows): stub unsupported compute drivers

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* ci(windows): add MSVC mise build lane

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* docs(windows): document MSVC build-only design

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* docs(agent): add Windows MSVC build skill

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* feat(windows): add Windows build support

Signed-off-by: Akber Raza <akberr@nvidia.com>

* refactor(windows): consolidate Windows-specific dependencies and improve build logic

Signed-off-by: Akber Raza <akberr@nvidia.com>

* feat(windows): add libclang path resolution and update cargo commands with bundled Z3 features

Signed-off-by: Akber Raza <akberr@nvidia.com>

* chore(tooling): lock Windows tool artifacts

Signed-off-by: Giedrius Burachas <gburachas@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* feat(windows): enhance libclang path resolution to support architecture-specific subdirectories

Signed-off-by: Akber Raza <akberr@nvidia.com>

* Fix Windows dependency gating after sync merge

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(z3): update Z3 header path requirements in Windows build documentation and scripts

Signed-off-by: Akber Raza <akberr@nvidia.com>

* docs(windows): relocate Windows MSVC build design to architecture/

Why: windows-msvc-build-design.mdx is a design document ("design decisions for
the native Windows MSVC build lane"), but it lived in the published, user-facing
docs/reference/ tree. Per AGENTS.md (Documentation) and architecture/README.md
("rfc/ vs architecture/"), design content belongs in architecture/ (or rfc/),
not in published reference. It also shared Fern sidebar "position: 6" with the
MXC compute-driver design page, colliding in the Reference nav ordering.

What:
- Move docs/reference/windows-msvc-build-design.mdx ->
  architecture/windows-msvc-build.md.
- Strip the Fern publish frontmatter and add a plain H1, matching the other
  architecture docs.
- Register it in the architecture doc index in architecture/README.md.
- Repoint the inbound references (build-openshell-mxc-windows skill + reference,
  implement-openshell-mxc-driver skill) to the new path.

With both design pages moved out of docs/reference/, the duplicate position-6
sidebar collision is resolved.

Signed-off-by: Akber Raza <akberr@nvidia.com>

* remove openshell-supervisor-network from unsupported driver package test exclusion list

Signed-off-by: Akber Raza <akberr@nvidia.com>

# Conflicts:
#	tasks/scripts/windows-msvc.ps1

* fix(interceptors): gate unix-only imports so the crate builds on Windows

openshell-gateway-interceptors failed to compile on Windows (E0432: no UnixStream in tokio::net), breaking any Windows build of openshell-server (which depends on it unconditionally). The connect_unix_endpoint fn was already #[cfg(unix)]-gated, but the imports it uses (UnixStream, TokioIo, Uri, service_fn) were left ungated. Gate those four imports with #[cfg(unix)] too. No behavior change on unix; Windows now compiles (no errors, no unused-import warnings).

Signed-off-by: Akber Raza <akberr@nvidia.com>

* feat(windows): add native ARM64 test support

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mise): skip Skaffold on Windows

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(windows): harden ARM64 toolchain discovery

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(windows): scope ARM64 toolchain preflight

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(windows): restore compatibility after GitHub sync

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(windows): avoid rate-limited Z3 source lookup

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mise): skip Helm checks on Windows

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(windows): support repository pre-commit checks

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(windows): stabilize native MSVC validation

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(windows): harden shared Z3 source cache

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>

* fix(windows): avoid leaking MSVC flags into clang-cl

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(windows): complete ARM64 migration audit

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(windows): restore ARM64 Ninja discovery

Signed-off-by: Akber Raza <akberr@nvidia.com>

* refactor(windows): separate platform crate roots

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(windows): restore proto include cfg gating

Signed-off-by: Akber Raza <akberr@nvidia.com>

* refactor: address lint errors

* fix(windows): add preflight check for proxy auth file path

* docs(windows): update GitHub checkout guidance

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(windows): restore CI after dependency updates

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mise): repair Windows sccache lock entry

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(windows): reconcile validation after rebase

Signed-off-by: Akber Raza <akberr@nvidia.com>

* refactor(server): exclude unsupported drivers on Windows

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* refactor(server): isolate platform driver config

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(windows): repair unsupported driver contract test

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(sandbox): remove stale dependencies

Signed-off-by: Akber Raza <akberr@nvidia.com>

* ci(windows): pin x64 workflow actions

Signed-off-by: Akber Raza <akberr@nvidia.com>

* ci(windows): align x64 Rust toolchain

Signed-off-by: Akber Raza <akberr@nvidia.com>

* ci(windows): align ARM64 workflow setup

Signed-off-by: Akber Raza <akberr@nvidia.com>

* refactor(windows): exclude unsupported runtime crates

Signed-off-by: Akber Raza <akberr@nvidia.com>

* refactor(windows): exclude unsupported crates at workspace boundary

Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com>

* refactor(server): gate builtin driver config by platform

Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com>

* fix(sandbox): restore crate documentation

Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com>

* ci(windows): make build workflow manual

Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com>

* ci(windows): temporarily enable pull request builds

Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com>

* ci(windows): cache Rust dependencies

Signed-off-by: Akber Raza <akberr@nvidia.com>

* refactor(windows): remove unnecessary platform changes

Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com>

* ci(windows): make build workflow manual

Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com>

* fix(ci): synchronize mise lockfile

Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com>

* fix(ci): normalize mise provenance metadata

Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com>

* refactor(python): isolate Windows atomic replace retry

Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com>

* fix(python): type Windows permission test errors

Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com>

---------

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>
Signed-off-by: Giedrius Burachas <gburachas@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Signed-off-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com>
Co-authored-by: Shailendra Singh <shailendras@nvidia.com>
Co-authored-by: Giedrius Burachas <gburachas@nvidia.com>
Co-authored-by: Jamie King <jamiek@nvidia.com>
Co-authored-by: Piotr Mlocek <pmlocek@nvidia.com>
Co-authored-by: Piotr Mlocek <1116309+pimlock@users.noreply.github.com>
2026-08-11 21:00:36 +00:00
Emilien MacchiandMrunal Patel 3e191558b3 feat(build): add glibc-static supervisor libc variant (#2682)
The supervisor binary runs inside sandbox images whose libc and glibc
version are unknown at build time, so it must be statically linked. Add
SUPERVISOR_LIBC to select between the default musl variant and a new
glibc-static variant that builds the GNU target with +crt-static.

glibc-static has no cross-compile path: zig cc accepts -static for
*-linux-gnu targets and emits a dynamically linked binary anyway. The
staging script therefore refuses a cross-arch request for that variant
rather than silently degrading linkage, and requires a native
per-architecture build.

Add verify-static-binary.sh, run after every supervisor build in both the
staging script and CI so linkage cannot regress unnoticed for either
variant. It inspects via readelf (or greadelf/llvm-readelf) and fails closed
rather than trusting the tool's exit status: every inspection must produce no
diagnostics, the input must be an executable ELF (ET_EXEC, or ET_DYN with
DF_1_PIE) whose PT_LOAD segments all lie within the file, whose dynamic table
agrees with PT_DYNAMIC, and which carries no PT_INTERP and no DT_NEEDED. That
rejects a dynamically linked, truncated, corrupt, non-ELF, or shared-object
input that naive parsing would misread as static. Hosts without any inspector
(e.g. macOS, which ships no binutils) skip with a warning; Linux, including
CI, requires one and fails closed.

No image or release workflow builds the glibc-static variant, so add a
dedicated supervisor-static-validate workflow that builds it on both
architectures and runs the verifier. rust-native-build.yml uses self-hosted
runners, which reject pull_request-triggered jobs, so it validates in the merge
queue and on pushes to main that touch the build inputs, plus a nightly
schedule, so the GNU + crt-static build branch cannot regress unnoticed.

The default is unchanged, so image, release, and CI behavior is identical.
Selecting glibc-static statically links LGPL glibc into a redistributed
binary, which is why it is opt-in.

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
Co-authored-by: Mrunal Patel <mrunalp@gmail.com>
2026-08-11 17:43:09 +00:00
Evan LezarandDrew Newberry d220d89468 feat(compute): negotiate gateway callback listeners (#2492)
* feat(compute): query gateway listener requirements

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* feat(compute): add Podman listener requirements

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(docker): use default gateway bind address

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(gateway): avoid wildcard primary listener

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(podman): validate callback listener discovery

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(server): support split dual-stack listeners

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(podman): support legacy rootless listener discovery

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(e2e): accept loopback plaintext rejection

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* docs(agent): add callback listener diagnostics

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(server): restrict compute callback listeners

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(podman): validate local callback port

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(server): clarify callback listener contract

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(podman): require pasta for local callbacks

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* docs(gateway): document RPM listener default

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* refactor(server): keep listener provenance diagnostic-only

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(compute): preserve callback listener isolation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(e2e): remove Podman callback relay

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(packaging): preserve Podman callback loopback

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci(e2e): run VM smoke on nested-virt runner

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci(e2e): gate VM smoke on usable KVM

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci(e2e): probe KVM through VM driver

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci(e2e): tolerate hosted KVM denial

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(server): close traced futures before assertions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* revert: remove tracing test stabilization

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Evan Lezar <elezar@nvidia.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
2026-07-31 16:41:06 +00:00
Drew Newberry 1cbfc0d510 test(e2e): add reusable QEMU infrastructure for E2E tests (#2471)
* test(vm): add composable QEMU test guests

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(vm): describe test VM directory structure

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(vm): replace shell catalog functions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(vm): add Fedora release guest support

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(vm): enable rootless Podman socket

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(test-guest): add OCI-backed image caching

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* perf(test-guest): accelerate cached guest startup

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(test-guest): address review feedback

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(test-guest): verify OCI cache provenance

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(test-guest): harden cached guest reuse

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(test-guest): refresh runtime setup state

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(test-guest): support E2E runner inputs

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(test-guest): harden runner and OCI reuse

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(test-guest): prepare Podman E2E artifacts

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(test-guest): address review findings

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* revert(test-guest): remove recent Podman artifact changes

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(test-guest): canonicalize scp source paths

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(test-guest): provision artifacts with Ansible

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(test-guest): populate missing caches on startup

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-07-29 23:41:29 +00:00
Evan Lezar eb380d71a3 ci(e2e): probe VM gateway readiness (#2544)
Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-07-29 18:33:16 +00:00
Philippe Martin 77e5c32217 feat(sandbox,gateway): route sandbox egress through corporate HTTP proxy (#2245)
* feat(sandbox,gateway): route sandbox egress through corporate HTTP proxy

- Chain sandbox egress through a corporate HTTP proxy so outbound
  traffic from within the sandbox respects the host proxy settings
- Forward sandbox proxy environment variables to the generated Podman
  config so the proxy is applied consistently to Podman-managed workloads

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(sandbox,podman): make corporate proxy routing operator-owned

The operator-configured corporate egress proxy was injected under the
conventional HTTPS_PROXY/HTTP_PROXY/NO_PROXY names as defaults beneath
sandbox spec/template environment, so a sandbox creator could redirect
egress at an arbitrary proxy or disable proxying with NO_PROXY=*.

Route the boundary through reserved, supervisor-only variables
(OPENSHELL_UPSTREAM_HTTPS_PROXY/HTTP_PROXY/NO_PROXY) written in the
Podman driver's required-variable tier. Any sandbox-supplied value under
a reserved name is stripped before the operator value is applied, so the
supervisor never observes a reserved proxy variable the operator did not
set. The supervisor now reads only the reserved names and ignores the
conventional proxy variables the sandbox controls.

Add the reserved proxy variables to the supervisor-only child-environment
denylist so the corporate proxy URL and any embedded credentials are not
inherited by the sandbox workload, which reaches egress through the local
policy proxy and never needs them.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* feat(sandbox,podman): deliver corporate proxy credentials via secret file

Proxy credentials were embedded inline in the proxy URL, so they were
stored in gateway.toml and exposed in container metadata via
'podman inspect'.

Reject inline 'user:pass@' credentials in https_proxy/http_proxy at
startup (parsed with the url crate rather than hand-splitting), and add a
proxy_auth_file option pointing at a 'user:pass' file. The driver stages
that file as a per-sandbox root-only Podman secret, mounts it at a fixed
path, and exports only the path in the reserved
OPENSHELL_UPSTREAM_PROXY_AUTH_FILE variable, so the credential never
appears in config, environment, or container metadata. Reading the file
fails closed on a missing, empty, or control-character-bearing value.

The supervisor reads the credential from the mounted file and builds the
Proxy-Authorization: Basic header, rejecting control characters, and no
longer derives credentials from URL userinfo. The auth-file path is added
to the child-environment strip list, and the generated gateway.toml is
written owner-only (mode 0600).

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(sandbox,podman): fail closed on invalid upstream proxy configuration

The reserved OPENSHELL_UPSTREAM_* variables are an operator-owned egress
boundary, but the supervisor treated present-but-invalid values as unset:
an unsupported or malformed proxy URL was ignored with a warning, an
unreadable auth file proceeded without credentials, and a malformed
credential silently became unauthenticated. Any of these could quietly
downgrade the corporate proxy boundary to direct dialing or
unauthenticated proxy access.

Make every configured-but-invalid proxy or auth setting fatal to
supervisor proxy startup, emit an OCSF ConfigStateChange failure event
before refusing, and share URL validation semantics between the Podman
driver and the supervisor through a single validator in
openshell-core (parse_upstream_proxy_url), so a value accepted at
sandbox-create time can never be rejected in-container or vice versa.

Inline user:pass@ URL credentials are now fatal in the supervisor too
(previously warn-and-strip), matching the driver. Unset or empty
variables still mean no proxy; only present-but-invalid values fail.
Error paths never include credential content.

Addresses the fail-closed review item on #2245.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(sandbox,podman): finish the fail-closed upstream proxy credential contract

The upstream proxy URL already had a single shared validator, but the
credential did not: the Podman driver rejected only CR/LF/NUL while the
supervisor rejected every control character, so a credential accepted at
sandbox-create time (e.g. one containing a tab) could still be rejected
in-container. Present-but-whitespace reserved OPENSHELL_UPSTREAM_*
values were also silently treated as unset, quietly downgrading the
operator's egress boundary to direct dialing.

Add parse_upstream_proxy_credential to openshell-core as the single
source of truth for the documented user:pass credential form (non-empty
user, no control characters, trimmed) and use it in both the Podman
driver's secret staging and the supervisor's Proxy-Authorization header
construction. Error variants carry no payload so credential content can
never leak into messages.

Make a present-but-empty reserved variable fatal to supervisor proxy
startup instead of meaning "unset"; only fully unset variables disable
the proxy. The driver correspondingly rejects an empty no_proxy at
config time so it can never inject a value the supervisor refuses.

Addresses the remaining fail-closed credential/config review item on

Signed-off-by: Philippe Martin <phmartin@redhat.com>
#2245.

* fix(sandbox,podman): close remaining fail-open upstream proxy config paths

Two configuration paths could still silently run without the proxy
boundary the operator believed was in effect.

A no_proxy bypass list configured without any https_proxy/http_proxy
was accepted by both the driver and the supervisor and simply meant
"dial everything directly". Reject it on both sides, exactly like the
existing proxy_auth_file-without-proxy rule: an operator who wrote a
bypass list assumed proxying was active, so accepting it hides a
fail-open state.

The gateway.sh dev script guarded proxy settings with [[ -n "${VAR:-}"
]], which conflates unset with explicitly-empty and dropped the latter
before the gateway's validation could see it. Use ${VAR+x} instead so
a set-but-empty variable is written into gateway.toml and rejected at
startup by validate_proxy_config rather than silently discarded.

Addresses the remaining fail-open configuration review item on #2245.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(sandbox,podman): reject upstream proxy URLs with path, query, or fragment

parse_upstream_proxy_url accepted URLs like http://proxy.corp.com:8080/some/path
and silently discarded everything after host:port, a lenience inherited
from the original supervisor parser. A forward proxy is addressed by
host:port only, so extra components indicate a misconfiguration (for
example a pasted endpoint URL) and silently truncating them violates
the present-but-invalid-is-fatal contract enforced everywhere else in
this configuration surface.

Reject a path, query, or fragment in the shared validator with a new
UnexpectedComponent error. A bare trailing slash remains accepted
because the url crate normalizes an absent http path to "/", making the
two indistinguishable. Both the Podman driver (gateway startup) and the
supervisor (sandbox startup) inherit the rule through the shared
parser, keeping their semantics identical by construction.

Addresses the proxy URL component review item on #2245.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(sandbox,podman): remove plain-HTTP upstream proxy support

The http_proxy path tunneled plain-HTTP requests through the corporate
proxy with CONNECT to port 80 and then sent origin-form requests down
the tunnel. Conventional enterprise forward proxies expect plain HTTP
as absolute-form requests sent directly over the proxy connection, and
commonly refuse CONNECT to port 80, so the setting looked supported but
failed against typical deployments. Tunneling also blinds the proxy to
the one protocol it could inspect.

Narrow the feature to TLS (CONNECT) egress only, which is the
conventional and already-correct case: plain-HTTP requests now always
dial the destination directly, and only client CONNECT tunnels chain
through the corporate proxy. Remove the http_proxy config field, the
--sandbox-http-proxy / OPENSHELL_SANDBOX_HTTP_PROXY driver surface, the
reserved OPENSHELL_UPSTREAM_HTTP_PROXY variable, and the UpstreamScheme
plumbing. The feature never shipped, so this is a clean removal; a
stray http_proxy key in gateway.toml still fails loudly through the
config's deny_unknown_fields.

Removing the plain-HTTP proxy branch also removes its host-gateway
special case; the architecture doc now documents the real host-gateway
behavior (add driver-injected host aliases to the reserved NO_PROXY
list) instead of an invariant the HTTPS path never implemented.

Plain-HTTP forwarding through a corporate proxy can return later as
absolute-form forwarding behind its own design review.

Addresses the plain-HTTP forwarding review item on #2245.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(sandbox,podman): escape generated TOML and require explicit proxy URL form

Escape backslashes, quotes, and control characters when gateway.sh writes
proxy values into gateway.toml, so a hostile or unusual environment value
cannot corrupt the config or inject extra keys.

Restrict the upstream proxy URL grammar to the documented http://host:port
form: a scheme-less value is no longer normalized to http:// and a missing
port is no longer silently defaulted to 80. Docs, README, and CLI help now
state the explicit-form requirement consistently.

Also fix a test-only call of handle_tcp_connection that was missing the
upstream_proxy argument added in an earlier commit.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(sandbox,podman): preserve tunneled bytes read with the CONNECT response

The CONNECT handshake reads the corporate proxy's response in chunks, so
the read that completes the header block can also contain the first
tunneled payload bytes. Those bytes were discarded, silently corrupting
the start of the tunnel for server-speaks-first destinations or proxies
that coalesce writes.

connect_via now returns a PrefixedStream that replays any bytes received
past the response terminator before reading from the socket again; writes
pass through unchanged. Direct dials wrap the stream with an empty prefix
so downstream relay and TLS paths keep a single stream type, and
tls_connect_upstream is generalized to any AsyncRead + AsyncWrite stream.

Adds regression coverage for a combined response/payload read and for
prefix replay ordering.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(sandbox,podman): gate cleartext proxy Basic auth behind an explicit opt-in

Proxy-Authorization: Basic is base64 over the plain-TCP connection to the
http:// corporate proxy, so anyone on the network path between the sandbox
host and the proxy can recover the credential. Sending it is now an
explicit operator decision instead of an implicit side effect of
configuring proxy_auth_file.

Add a proxy_auth_allow_insecure driver setting (CLI
--sandbox-proxy-auth-allow-insecure, env
OPENSHELL_SANDBOX_PROXY_AUTH_ALLOW_INSECURE), delivered to the supervisor
as the reserved OPENSHELL_UPSTREAM_PROXY_AUTH_ALLOW_INSECURE variable.
Fail-closed pairing on both sides: an auth file without the
acknowledgement is rejected at gateway startup and at supervisor startup,
as is the acknowledgement without an auth file or any value other than
'true'. gateway.sh writes the key only as a TOML boolean; a non-boolean
value is emitted as a quoted string so the gateway rejects it at startup
instead of risking injection.

Documents the exposure prominently in the gateway config reference, the
driver README, and the sandbox architecture doc.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(sandbox,podman): honor port qualifiers and resolved addresses in NO_PROXY

Two divergences from the documented NO_PROXY contract:

- A port-qualified entry (internal.corp:8443) was silently stripped to its
  hostname and bypassed the proxy for every port, excluding traffic the
  operator never listed. Entries now keep the optional :port qualifier
  (also on IP and CIDR entries) and only apply to that destination port; a
  trailing qualifier that is not a valid port stays part of the pattern
  instead of widening the entry.

- IP and CIDR entries only matched IP-literal hosts, so a bypass like
  10.0.0.0/8 never applied to hostnames resolving into that range. NO_PROXY
  evaluation now sees the validated resolved addresses: an IP/CIDR entry
  matching through resolution authorizes a direct dial of only the
  addresses it contains, so a bypass scoped to an internal range cannot
  widen into a direct dial of addresses outside it. Hostname-level matches
  (loopback, wildcard, domain entries, IP-literal hosts) keep authorizing
  all validated addresses.

proxy_for is replaced by decision(host, port, resolved) returning either
the proxy endpoint or the permitted direct-dial subset, and dial_upstream
restricts the direct connect to that subset.

Adds regression coverage for port-scoped bypasses on domain, IP, and CIDR
entries, invalid port qualifiers, resolved-address matching, and
split-resolution subset dialing.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(sandbox,podman): bind proxied CONNECT tunnels to validated addresses

The CONNECT request sent to the corporate proxy carried the destination
hostname, so the proxy resolved the name itself and the addresses that had
passed SSRF and allowed_ips validation were discarded. Split-horizon DNS
or rebinding at the proxy could then reach internal or otherwise
unapproved destinations through a tunnel the supervisor logged as
validated, and IP-range policy could not be enforced at all on proxied
dials.

CONNECT now targets a validated resolved address by default: the proxy
performs no DNS resolution and the tunnel stays bound to the answer the
supervisor checked. The hostname still travels inside the tunnel (TLS SNI,
application Host), so destination servers behave normally. In
split-horizon networks, operators point the gateway host at the corporate
resolver so internal names validate to their internal addresses.

For proxies whose ACLs filter on hostnames and reject IP CONNECT targets,
a new proxy_connect_by_hostname opt-in (CLI
--sandbox-proxy-connect-by-hostname, env
OPENSHELL_SANDBOX_PROXY_CONNECT_BY_HOSTNAME, reserved
OPENSHELL_UPSTREAM_PROXY_CONNECT_BY_HOSTNAME) restores hostname CONNECT,
documented as re-opening proxy-side resolution and making the proxy's ACLs
the effective egress control. Fail-closed pairing on both sides: the
opt-in without a proxy, or any value other than 'true', is fatal.

Adds regression coverage for the IP CONNECT request line (including IPv6
bracketing and hostname non-leakage), the hostname opt-in, and the
config pairing rules in driver and supervisor.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(sandbox,podman): reject empty port after bracketed IPv6 proxy host

http://[fd00::1]: passed the explicit-port check because the bracketed
branch only tested for a colon after the bracket, then fell back to port
80 — violating the fail-closed http://host:port contract and potentially
sending configured Basic credentials to an unintended service. Require a
non-empty suffix after ]:, matching the unbracketed branch, and cover the
case in the shared parser tests.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(sandbox,podman): fall back across validated addresses in proxied CONNECT

The direct path hands TcpStream::connect the whole validated address list
and it falls back across them, but the validated-IP CONNECT path attempted
only the first address, so a dual-stack destination could fail through the
corporate proxy even when a later validated address was reachable.

connect_via_validated tries each validated address in order under one
aggregate CONNECT_HANDSHAKE_TIMEOUT budget, returning the first success;
when every attempt fails the error names the attempt count and carries the
last failure. An empty address list is rejected up front.

Adds regressions for first-fails/second-succeeds fallback (asserting both
CONNECT request lines), the aggregate all-addresses failure message, and
the empty-list rejection.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(sandbox,podman): strip new reserved proxy vars and complete docs/tests

Follow-ups from review:

- Add OPENSHELL_UPSTREAM_PROXY_AUTH_ALLOW_INSECURE and
  OPENSHELL_UPSTREAM_PROXY_CONNECT_BY_HOSTNAME to the supervisor-only
  strip list so workload child processes never inherit them, matching the
  documented contract for the other reserved proxy variables, and cover
  both in the supervisor-only variable test.
- Document all five OPENSHELL_SANDBOX_* proxy variables in the
  mise run gateway help text, marked Podman-only and stating the
  auth-file/acknowledgement pairing, and complete the gateway-key list in
  the Podman README.
- Add NO_PROXY composition coverage for bracketed IPv6 entries with port
  qualifiers, bare IPv6 entries, and IPv6 CIDR matching against an
  IPv6-literal host and against a hostname's resolved addresses,
  including the port-qualified CIDR form.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(sandbox,podman): cap each proxied CONNECT attempt within the shared budget

A proxy that accepted the first CONNECT request but never responded
consumed the entire aggregate handshake timeout, so later validated
addresses were never tried and the hang defeated the multi-address
fallback.

Each attempt is now time-boxed to its fair share of the time remaining
before the shared deadline (remaining / attempts_left): a hanging attempt
is cut off with enough budget left for every remaining address, while
time a fast failure does not use rolls over to later attempts and the
total never exceeds CONNECT_HANDSHAKE_TIMEOUT. A timed-out attempt is
recorded like any other failure, and the aggregate error distinguishes
all-attempted from budget-exhausted runs.

Adds a first-hangs/second-succeeds regression driven through a
test-visible budget parameter so it runs in about a second instead of a
real 30s window.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(sandbox,podman): deliver corporate proxy config on the supervisor argv

The proxy settings were injected as reserved OPENSHELL_UPSTREAM_* container
environment variables. The driver only wrote names the operator configured,
but container runtimes layer the spec environment over ENV values baked
into the sandbox image, so an image could supply NO_PROXY=*, enable
hostname CONNECT, or point an unconfigured deployment at an
attacker-controlled proxy whenever the operator left a field unset.

The settings now travel as supervisor command-line arguments
(--upstream-proxy, --upstream-no-proxy, --upstream-proxy-auth-file,
--upstream-proxy-auth-allow-insecure,
--upstream-proxy-connect-by-hostname) built by the driver from operator
config. The driver sets the container entrypoint and command explicitly,
so neither sandbox spec/template environment nor image ENV can influence
argv, and an omitted flag genuinely means unconfigured — in every
supervisor topology, since the supervisor no longer consults its
environment for these settings at all. Credentials stay on the root-only
secret mount; only the mount path appears on argv.

The reserved environment names, their strip-list entries, and the
env-based validation surface are removed. UpstreamProxyConfig::from_args
replaces from_env, reusing the same shared fail-closed validation and
pairing rules keyed by the CLI flag names.

Driver tests now assert the argv contract, including that sandbox-supplied
environment cannot add, remove, or redirect proxy flags; supervisor tests
cover from_args mapping and its pairing rules.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* docs(sandbox): align proxy comments with the argv transport

The argv migration left comments describing the configuration as reserved
environment variables ("reserved value", "present-but-empty variable",
"reserved upstream proxy variables"). Rephrase them as driver-supplied
arguments and operator settings so the documented trust boundary matches
the implementation. Comment-only change.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(sandbox): parse bracketed IPv6 authorities in client CONNECT targets

parse_target split the CONNECT authority at the first colon, so an
IPv6-literal target like [2001:db8::1]:443 always failed port parsing and
IPv6-literal clients could never reach policy evaluation; a regression
test even locked in that failure. Parse the RFC 3986 bracketed form and
return the host bracket-free, matching what DNS resolution, SSRF
validation, NO_PROXY matching, and the upstream CONNECT builder expect.
Unclosed brackets, a missing or empty port after the bracket, and
non-numeric ports are rejected; unbracketed behavior is unchanged.

Replaces the failure-locking test with success coverage for bracketed
targets and adds malformed-bracket rejection cases.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* test(podman): cover proxy-auth secret cleanup across lifecycle failures

The per-sandbox proxy-auth credential secret is staged before the
container is created and removed on cleanup, but no test proved the
cleanup paths actually issue the secret removal. Add Podman-stub tests
that drive create_sandbox to a container-create failure and to a
start failure, and delete_sandbox for an out-of-band deletion, asserting
each path issues the DELETE for the per-sandbox proxy-auth secret so a
credential can never outlive the sandbox that owned it.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* docs: list corporate proxy keys in the Podman compute-driver overview

The Fern Podman driver section enumerated its gateway.toml keys but
omitted the corporate egress proxy settings. Add https_proxy, no_proxy,
proxy_auth_file, proxy_auth_allow_insecure, and proxy_connect_by_hostname
with a pointer to the gateway configuration reference for the full
contract. No navigation change: the reference folder already includes the
gateway configuration page.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* test(sandbox): cover the SSRF-to-TLS composition across the proxy tunnel

Existing tests exercised validated-IP CONNECT and the upstream-TLS helper
independently, but not the full boundary. Add an end-to-end regression
that stands up a fake corporate proxy tunneling to a fake TLS server and
drives the real path: connect_via_validated CONNECTs to the validated
address, the proxy splices the tunnel, and tls_connect_upstream verifies
the upstream certificate against the original hostname carried in SNI.

It asserts the CONNECT authority is the validated IP and never the
hostname, that verification succeeds for the matching hostname, and that a
mismatched hostname is rejected — proving a rebinding or split-horizon
substitution behind the proxy cannot pass certificate verification.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(sandbox,podman): bound proxy-auth reads, reject port 0, fix stale comment

Three review findings:

- CWE-400: the proxy-auth credential file was read with an unbounded
  read_to_string on both the driver (sandbox-create) and supervisor
  (startup) paths, so a huge file or a special file such as /dev/zero
  could exhaust memory. Add a shared bounded reader in openshell-core that
  rejects non-regular files, caps the size at 4 KiB, and reads at most that
  many bytes; the driver runs it via spawn_blocking. Covers oversized,
  special-file, and missing-path cases on both sides.

- Reject an upstream proxy URL with port 0: it passed the explicit-port
  check and startup validation but is not a connectable TCP port, so every
  proxied dial would fail later. Add a typed ZeroPort error with
  shared-validator and Podman-config tests.

- Reword a driver-config comment that still described a 'reserved
  variable' to match the argv transport.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(sandbox): open proxy-auth file non-blocking to reject FIFOs promptly

read_upstream_proxy_credential_file opened the path with a blocking
File::open before the regular-file check, so a configured FIFO with no
writer would block open() indefinitely — hanging sandbox creation on the
driver and supervisor startup. Open with O_NONBLOCK on Unix so the open
returns immediately, then reject the non-regular file as before;
O_NONBLOCK has no effect on the later read of a regular file. Adds a
mkfifo regression asserting the reader returns promptly with a
non-regular-file error instead of hanging.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* test(podman): cover corporate proxy egress across driver and supervisor

The existing corporate-proxy tests construct config structs or call CONNECT
helpers directly, so none of them detect a break in the wiring between
layers: gateway TOML deserialization, the Podman argv and secret-mount
semantics, supervisor CLI parsing, or policy denial before proxy contact.

Add a Podman e2e that drives the whole chain against a fake authenticated
forward proxy and asserts that an approved TLS request traverses it with a
validated-IP CONNECT, a policy-denied destination is refused with 403
without ever reaching the proxy, credentials arrive through the mounted
per-sandbox secret, and deleting the sandbox removes that secret.

SupportContainer is a new harness fixture. Unlike ContainerHttpServer it
probes readiness with a TCP connect rather than an HTTP GET, so it can host
a forward proxy and TLS servers, and it exposes container logs and network
IP for assertions.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* test(podman): restart the gateway on the proxy-config panic path

The panic cleanup for the temporary corporate proxy configuration
restored the gateway TOML but left the gateway process running with the
temporary configuration still loaded, which could poison later test
binaries in the same run. Nothing restarted it: the only ManagedGateway
is the short-lived one inside restart_gateway, and its Drop only calls
start, which does not reload config for an already-running gateway.

Restore and synchronously stop/start the gateway in Drop, and set
restored only after the normal restore and restart both succeed so a
failed restart no longer suppresses the fallback.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(supervisor-network): derive the loopback proxy bypass from resolved IPs

The automatic bypass treated the host string "localhost" as proof that
the destination was loopback and returned every resolved address. A
sandbox controls its own /etc/hosts and resolve_socket_addrs consults it
before DNS, so a workload could map localhost to any policy-allowed
address and dial it directly, escaping the operator proxy and the
inspection and audit boundary it exists to provide.

Check the resolved addresses instead: the name bypasses only when the
resolution is non-empty and every address is loopback. A mixed answer is
not partially honored, and an IP literal is still authoritative for
itself. A spoofed localhost falls through to the entries below, so an
explicit operator NO_PROXY entry is still honored.

This matches the trust model detect_trusted_host_gateway already applies
to the same hosts file, which validates the mapped address rather than
the alias.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(podman): clean up proxy-auth secret when container already deleted

The delete_sandbox early-return path cleaned up the token secret but
skipped the proxy-auth secret, leaking it on disk. Also update tests
for recent API changes (Optional socket_path, workspace field,
list-based container lookup).

Signed-off-by: Philippe Martin <philippe@openshell.dev>
Signed-off-by: Philippe Martin <phmartin@redhat.com>

---------

Signed-off-by: Philippe Martin <phmartin@redhat.com>
Signed-off-by: Philippe Martin <philippe@openshell.dev>
2026-07-24 16:51:52 +00:00
Jim Meyer 745512e325 fix(build): raise open-file limit for host musl cross-compile on macOS (#2307)
Closes #2112

The host `cargo zigbuild` for `*-unknown-linux-musl` opens ~333 `.rlib`
files at once during the static link, exceeding macOS's default soft
limit of 256 and failing with `ProcessFdQuotaExceeded`. This blocked the
docker/podman `mise run gateway` paths for macOS contributors; only the
VM driver path guarded against it.

Extract the VM path's `ensure_build_nofile_limit` guard into a shared
`tasks/scripts/build-env.sh` and call it from the host-staging chokepoint
(`stage-prebuilt-binaries.sh`), fixing docker, podman, and all
docker:*/multiarch host cross-compiles at once. De-duplicate the VM
script to source the shared helper. The guard is a no-op on Linux and
when cargo-zigbuild is absent, so CI and Linux dev are unaffected. The
limit is read from `OPENSHELL_BUILD_NOFILE_LIMIT` (default 8192),
honoring the legacy `OPENSHELL_VM_BUILD_NOFILE_LIMIT` for back-compat.

Also correct the stale comment in gateway-docker.sh (the cross-compile
runs on the host, not inside Linux containers) and document the guard in
architecture/build.md. Adds tasks/scripts/test-build-env.sh, wired into
`mise run test` via `test:build-env`.

Signed-off-by: Jim Meyer <jim@meyer4hire.com>
2026-07-20 21:11:20 +00:00
Ignas Baranauskas a72711697d chore: remove deprecated --keep flag from docs, scripts, and e2e tests (#2126)
* docs: remove deprecated --keep flag from tutorials and examples

The --keep flag is deprecated, hidden, and a no-op since sandboxes
are kept by default. Remove references from tutorial docs and example
READMEs that explain it as a real feature.

- Remove --keep from sandbox create commands
- Remove --keep explanation text
- Clarify that sandboxes are kept by default

Signed-off-by: Ignas Baranauskas <ibaranau@redhat.com>

* chore: remove deprecated --keep usage from scripts and e2e tests

The --keep flag is a deprecated no-op since sandboxes are kept by
default. Stop passing it in internal scripts, e2e test scripts,
and example demo scripts.

Signed-off-by: Ignas Baranauskas <ibaranau@redhat.com>

---------

Signed-off-by: Ignas Baranauskas <ibaranau@redhat.com>
2026-07-07 14:36:25 +02:00
Evan Lezar a226806026 test(e2e): run gpu workloads from manifest (#1709)
* test(e2e): add workload manifest build flow

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(e2e): add gpu workload validation tests

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* ci(e2e): build gpu workloads before gpu e2e

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-06-30 09:54:12 +02:00
Taylor Mutch 7e0cce405d fix(build): use zig archive tools for cross builds (#2014)
Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
2026-06-26 15:35:54 -05:00