Commit Graph
174 Commits
Author SHA1 Message Date
pkhodade-NV 9ceec26d42 fix(mxc): redact injected secrets from gateway diagnostics (#3853)
Redact injected environment values and the per-sandbox proxy password in
captured MXC output and decoded relay launch-failure diagnostics before
logging or publishing sandbox failure status.

Match the original text and redact the union of overlapping occurrences.
Preserve control-channel payloads and avoid allocating for unmatched text.
Add regression coverage and document exact-match and length limits.

Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
2026-10-01 10:02:55 -07:00
pkhodade-NV ab64e84bfc fix(core): enforce owner-only Windows ACLs on sensitive files and dirs (#3495)
* fix(core): enforce owner-only Windows ACLs on sensitive files and dirs

set_dir_owner_only/set_file_owner_only were unconditional no-ops on
Windows, so the CLI's mTLS client private key, OIDC/edge tokens, cached
SSH keys, and the gateway's key-encryption key relied entirely on
inherited NTFS ACLs with no OpenShell-applied restriction. Apply an
owner-only DACL via SetEntriesInAclW/SetNamedSecurityInfoW with
PROTECTED_DACL_SECURITY_INFORMATION to strip inherited ACEs, matching
the 0700/0600 guarantee already provided on Unix. is_file_permissions_too_open
now also works on Windows instead of being Unix-only, closing the
detection gap alongside the prevention gap.

Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
(cherry picked from commit 71560e947f85819efbcddf70ddda94befab62b0b)

* fix(core): treat a NULL DACL as too open in is_file_permissions_too_open

has_foreign_trustee conflated a NULL DACL with an unreadable/invalid
ACL and returned Some(false) (not too open) for both. Per the Win32
contract, a NULL DACL means the object grants full access to everyone
-- the most permissive state possible -- so it must be flagged as too
open. Split the null and invalid-ACL branches: null now returns
Some(true), invalid ACL keeps the existing unreadable-ACL fallback
(None, which the caller maps to false via unwrap_or). Adds a
regression test that constructs a real NULL DACL via a
SetNamedSecurityInfoW helper confined to the windows_acl module,
consistent with the existing unsafe-FFI confinement in that module.

Found by CodeRabbit review on MR !113.

Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
(cherry picked from commit 46e635a4ef1d6937cdb088f46aa85baee3d6ad28)

* fix(core): close three false-negative gaps in the Windows ACL audit

restrict_to_current_user() updated only the DACL, leaving a foreign
owner's implicit WRITE_DAC right intact -- they could later replace
the DACL we just set. Query OWNER_SECURITY_INFORMATION and take
ownership in the same SetNamedSecurityInfoW call; if the caller can't
(a genuinely foreign-owned object), the call now fails instead of
silently leaving the object insecure.

is_file_permissions_too_open() mapped every Win32 inspection failure
(missing READ_CONTROL, an invalid ACL, a token-query failure) to
"not too open" via unwrap_or(false). Fail closed instead: an
inspection failure is a security false-negative risk, not a green
light.

has_foreign_trustee()'s ACE loop only recognized plain
ACCESS_ALLOWED_ACE_TYPE and treated every other type as non-granting.
Windows also defines access-allowed object, callback, and
callback-object ACE variants that can grant rights to a foreign
trustee; this audit doesn't parse their wider layouts, so their mere
presence is now conservatively flagged as too open instead of
silently skipped.

Also updates architecture/gateway.md, which still described the
SQLite file-tightening behavior only in terms of Unix mode 0o600, to
distinguish it from the owner-only DACL behavior on Windows.

Addresses review comments on PR #3495.

Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>

* fix(core): conditional owner claim and audit owner in Windows ACL helpers

restrict_to_current_user: query the current owner before calling
SetNamedSecurityInfoW. Include OWNER_SECURITY_INFORMATION only when the
path has a foreign owner -- requesting it unconditionally fails with
ACCESS_DENIED (0x80070005) on standard credentials even when the current
user is already the owner, because WRITE_OWNER is not implied by object
ownership. A foreign-owned path still triggers an ownership claim and
fails hard if the claim is denied, preserving the security contract.

has_foreign_trustee: request OWNER_SECURITY_INFORMATION alongside
DACL_SECURITY_INFORMATION and reject paths with a foreign owner
immediately, before inspecting the DACL. A foreign owner has implicit
WRITE_DAC rights and can replace any DACL we set, so a clean DACL is not
sufficient evidence of safety on a foreign-owned object.

architecture/gateway.md: clarify that the Windows path-hardening behavior
sets mode 0o600 on Unix and applies a protected owner-only DACL on
Windows, with conditional ownership claim and fail-hard semantics for
foreign-owned objects.

Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>

---------

Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
2026-09-23 10:13:32 -07:00
Prekshi Vyas 0b8f3821e2 fix(network): normalize Windows binary paths (NVBug 6782969) (#3482)
* fix(network): normalize Windows policy binary paths

Match Windows executable identities using a stable case-insensitive, separator-normalized representation across policy data, L4 input, and L7 relay evaluation. Preserve exact matching on other platforms and keep the original path for hashing and filesystem access.

NVBug 6782969

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* fix(network): harden Windows binary matching

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* fix(ci): scope Windows relay test imports

* fix(network): harden Windows binary path matching

---------

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>
2026-09-22 13:47:25 -07:00
Prekshi Vyas 80b35dbfe0 fix(mxc): repair Windows inference demos (NVBug 6782874) (#3473)
* fix(mxc): repair Windows inference demos

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* fix(mxc): address inference demo review feedback

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

---------

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>
2026-09-21 17:28:52 -07:00
Drew Newberry 49b4f0eb7f feat(mxc): add UI policy, credentials, relay lifecycle, and proxy auth
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-17 12:06:06 -07:00
Seth Jennings dbe36eaf85 fix(security): harden Vault credential transport (#3329)
Reject non-loopback plaintext Vault endpoints, disable redirects, and support private CA bundles without weakening hostname verification. Update Helm configuration, documentation, operator skills, and regression coverage for OSSR-002.

Signed-off-by: Seth Jennings <sjenning@redhat.com>
2026-09-15 19:56:09 +00:00
krishicks 481ce566e1 fix(ocsf): correct HTTP activity context (#3316)
Previously, metadata events omitted both HTTP request and response objects,
early proxy rejections used HTTP Activity without request context, and
unsupported-scheme events did not expose enough safe HTTP context to satisfy
the OCSF 1.8 schema.

Now, metadata events include a method-only request and their actual HTTP
response codes without recording the metadata URL. Unsupported-scheme events
also include a method-only request plus the generated 400 response. Authority
mismatches and credential-resolution denials use HTTP Activity with their
generated 403 or 500 responses, and HTTP activity IDs are derived from the
request method.

Additionally, HttpActivityBuilder now enforces the OCSF request-or-response
constraint at compile time.

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-09-15 18:47:55 +00:00
Mrunal Patel 39cf4823f7 feat(api): add structured gateway errors and SDK decoding (#3313)
* feat(api): expose structured gateway errors across SDKs

Refs #3051. Add standard validation, conflict, and retry details; preserve raw transport status in Rust, Go, TypeScript, and Python; document status and recovery guidance.

This is the structured-error foundation only. Mutation result shapes, allow_missing, durable request deduplication, and exec retry semantics remain follow-up work.

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(python): preserve wrapped RPC cleanup handling

Inspect the original gRPC call when handling missing sandboxes during deletion waits and managed cleanup. Add intercepted cleanup regressions and clarify the error-wrapper migration contract.

Addresses the cleanup review on #3313; part of #3051.

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

---------

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
2026-09-15 17:51:11 +00:00
Evan Lezar c195e23267 test(conformance): remove plan-driven continuity tests (#3342)
Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-09-15 14:57:07 +00:00
Shiju fd3fd9cf74 feat(sandbox): explain failed calls to external tool servers (#3207)
Show configured tool server addresses and their last observed connection
results together in sandbox status. Keep sandbox lifecycle readiness
separate so an external connection failure does not mark the sandbox unready.

Expose direct endpoint records through the CLI and SDKs, with plain-language
failure explanations and gateway acceptance times. Keep observation tracking,
runtime reporting, and gateway validation in dedicated endpoint status modules.

Preserve bounded reporting, request attribution, retry ordering, and
configuration and supervisor authority checks. Clear obsolete observations
while retaining the configured addresses, and document the distinction
between an observed HTTP response, current availability, and tool success.

Signed-off-by: Shiju <shiju@nvidia.com>
2026-09-15 13:46:38 +00:00
26f2f96393 feat(mxc): add Windows host proxy for MXC sandbox network egress (#3163)
* Implement Windows host proxy integration and update dependencies for OpenShell

* Update README and gateway config to clarify egress proxy address handling and allocation

* Refactor ProxyIdentityMode to return Result for static_binary and add tests for binary path and SHA256 hash

* Enhance platform_hosts_path for Windows to use SystemRoot and improve error handling for hosts file reading

* Refactor FileFingerprint to use Option for mtime and ctime, simplifying metadata handling

* Add conditional compilation for Windows host module

* add unit tests for OPA policy evaluation and identity handling

* remove openshell-supervisor-network from unsupported driver package test exclusion list

* feat(mxc): enable host proxy TLS state generation

Generate per-sandbox TLS state for the MXC host proxy so HTTPS L7 enforcement can use the same MITM path as Linux. Grant generated CA material to the MXC process and inject standard trust env vars, while matching Linux behavior by disabling TLS termination on CA setup failure and relying on proxy fail-closed handling.

* fix(docs): remove outdated notes on governed egress from docs

* fix(tests): update TLS environment variable paths to use temporary directory

* fix(examples): make run-mxc-e2e harness correct and orphan-free

The MXC e2e harness never actually exercised the fs scenarios: it started
the gateway once and patched agent_command per scenario AFTERWARDS, so the
running gateway kept launching the default demo agent (not shipped in the
kit) and every fs scenario failed with CreateProcessW error:2. It also
scored on the `sandbox create` exit code (non-zero due to the harmless
interactive attach), wrote sandbox records to the persistent gateway DB
(leaving orphans that collided on later runs), and its deny scenarios never
proved denial.

Changes:
- Start a FRESH gateway per scenario so each scenario's agent_command is
  actually loaded (root cause of CreateProcessW error:2).
- Score by on-disk artifact / expected outcome, not `sandbox create` exit.
- Real deny assertions: a control write to a granted path must succeed
  (proves the agent ran) while the denied write must be absent. fs-empty
  probes an ungranted out-of-share path (share_dir is mapped rw by design).
- Run the gateway on an ephemeral in-memory DB (sqlite::memory:) so the
  harness never writes to the persistent store and cannot leave orphan
  sandbox records; also use unique per-run sandbox names + pre-delete.
- Fix the process_container probe: use a real cwd + absolute cmd.exe
  (canonical wxc-exec does not expand %TEMP% -> 0x8007010B).
- Fix summary counts (@() so a single FAIL is counted and exit is non-zero).

Verified PASS=4 FAIL=0 on 7F203-MXC-003 (no BaseContainer velocity keys)
using a canonical wxc-exec build (AppContainer fallback).

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(e2e): probe timeout is milliseconds (10ms->30000ms)

MXC process.timeout is wall-clock ms (wire.rs). The 10 value meant 10ms,
which the base-container tier (7F203-MXC-001/.181) enforced strictly and
timed the probe out. AppContainer path (.18/-003) happened to slip under
it. Bump to 30000ms so the process_container preflight probe is reliable
across both tiers.

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc): use native paths in real runtime probes

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc): make processcontainer work with mxc-latest-released wxc-exec

Three fixes to support the release wxc-exec binary (BaseContainer dispatcher)
in addition to mxc-fixes-env-vars:

1. Seed process env from host (driver.rs)
   ProcessContainer starts with a completely blank environment -- no PATH,
   SystemRoot, or anything.  Seed the process env from the gateway host
   environment so the agent binary can locate DLLs and run.  Skip internal
   Windows drive-letter variables (keys starting with '=') which cause
   CreateProcessW to return ERROR_ENVVAR_NOT_FOUND.  User agent_env entries
   and TLS CA vars are applied as overrides on top of the host env.

2. Remove TLS readonly_paths grant (driver.rs)
   The release wxc-exec (BaseContainer dispatcher) requires write-DAC
   permission on every path in readonly_paths to set up AppContainer ACLs.
   Adding the proxy's temp TLS directory caused a DACL error and exit -1.
   The CA cert paths remain available to the agent via TLS env vars.

3. Remove allowedHosts from network JSON (mxc.rs)
   The release wxc-exec rejects network.allowedHosts / network.blockedHosts
   on Windows with "not yet supported".  Removed the loopback exemption
   attempt (127.0.0.1, ::1, localhost) from the network section.
   Intra-container loopback works natively in the release binary without
   it -- the spawner can connect to the server at 127.0.0.1:22000 directly.

Additional changes:
- mxc-ws-agent.rs: add relay-debug.txt error capture and relay-ready.txt
  marker for reliable timing of host client connections.
- mxc-ws-gateway.toml: debug = true for JSON config dump during diagnosis.
- run-ws-agent-test.ps1: default port changed to 17670 (gateway default);
  relay-ready.txt polling before ws-echo to avoid connecting before the
  spawner has established the proxy bridge.

Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(e2e): address CodeRabbit review on run-mxc-e2e.ps1 (MR !46)

Four robustness/correctness fixes from CodeRabbit:

1. Start-Gw: kill the spawned gateway before the "did not start within 30s"
   throw. If the process is alive but never binds the port, $gw is not yet
   assigned in the caller, so the finally block cannot reap it -> orphan
   gateway holding the port for the next run.

2. create-fail scoring: a non-zero `sandbox create` exit alone is not proof
   of a policy rejection (gateway-registration/transport/fixture errors also
   exit non-zero and would false-pass). PASS now requires a genuine
   rejection signal (network / invalid_argument / network_policies) AND that
   it is not an infrastructure failure; other non-zero exits go to FAIL with
   output captured.

3. deny scenarios (ControlTarget path): snapshot the deny target AFTER
   Wait-File lands the control artifact, so a late denied write (enforcement
   regression racing the control write) can no longer be recorded as PASS.

4. -KeepRunning: break out of the scenario loop after the first scenario so
   a later scenario does not start a second gateway on the same port
   (previously a reliable port collision instead of a usable debug mode).

Re-verified PASS=4 FAIL=0 on both boxes (7F203-MXC-001 base-container and
7F203-MXC-003 AppContainer fallback); network-policy-rejected correctly
scores as "policy rejection".

Signed-off-by: Akber Raza <akberr@nvidia.com>

* feat(mxc-e2e): collect run-mxc-e2e output into a results bundle

Mirror the sibling run-*.ps1 scripts by collecting every run's logs into a
timestamped results-e2e-<stamp>\ folder and zipping it. The bundle contains the
console transcript, per-scenario gateway stdout/stderr, the exact TOML rendered
for each scenario, the policy fixture used, and a summary.txt with the verdict
table.

Per-scenario gateway logs now land in gateway.<scenario>.log/.err.log inside the
bundle instead of a single fixed gateway.e2e.log in the script directory.

Wrap pre-flight, mode setup, scenario definitions, and the scenario loop in a
single try/catch/finally so the finally always writes the summary, stops the
transcript, and zips the bundle -- even on a pre-flight failure. The existing
per-scenario gateway-cleanup try/finally stays nested inside. All scenario
logic, scoring rules, and comments are preserved.

Signed-off-by: Prashant S Khodade <pkhodade@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc-e2e): address CodeRabbit review on run-mxc-e2e.ps1

- Require -Scenario when -KeepRunning: the loop breaks after the first
  scenario, so a full-suite run would execute only one scenario yet still
  report the suite as PASS. Fail fast so a partial run can't be mislabeled
  complete.
- Start-Transcript now runs inside the guarded try block with a
  $transcriptStarted flag; Stop-Transcript is only called when it actually
  started, so a Start-Transcript failure still yields the results bundle.
- Wrap the -Scenario filter in @() so a single exact match stays an array
  (reliable .Count and a proper array for the scenario loop on PS 5.1).

Signed-off-by: Prashant S Khodade <pkhodade@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(examples): pass gateway config via OPENSHELL_GATEWAY_CONFIG for spaced paths

Start-Process -ArgumentList does not quote array elements, so launching the
gateway with a bare --config <path> token split on any space in the install
path (e.g. C:\Users\First Last\...), and clap rejected the fragment with
'unrecognized subcommand'. Every MXC example launcher that started the gateway
hit this when the kit was unzipped under a path containing a space.

Pass the config path through the OPENSHELL_GATEWAY_CONFIG env var (which the
gateway already reads via clap) and drop the --config token. Env vars carry
spaces safely.

Affected: run-ocsf-audit, run-mxc-e2e, run-demo, run-inference-test,
run-ollama-test. run-mtls-test was not affected (its launch passes no config
path). Root-caused and fix-verified on 7F203-MXC-003 from a spaced path.

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(run-mxc-e2e): improve scoring logic and enhance command execution handling

* fix(mxc): reconcile proxy support after rebase

Restore the proxy-enabled OCSF audit example removed by 13185f6e now that the host CONNECT proxy is present. Adapt the proxy lifecycle test to the target branch's DriverSandboxSpec policy delivery contract.

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc): grant sandbox access to proxy CA

- Share the per-sandbox public CA bundle with the AppContainer
- Add real wxc-exec HTTPS proxy coverage and document trust isolation

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc): reject unsupported network middleware

- Reject middleware-bearing MXC policies before sandbox lifecycle begins
- Guard host proxy startup and document the unsupported registry path
- Add mapper, lifecycle, and host proxy regression coverage

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc): reconcile host proxy with main

- remove obsolete inference routing from the host proxy adapter
- use the workspace AWS-LC provider in host-proxy tests
- adapt the forward-proxy test to ProxyIdentityMode

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(build): switch to bundled Z3 for Windows MSVC builds

* fix(mxc): reconcile host proxy after rebase

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Akber Raza <akberr@nvidia.com>
Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
Signed-off-by: Prashant S Khodade <pkhodade@nvidia.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Jamie King <jamiek@nvidia.com>
Co-authored-by: Shailendra Singh <shailendras@nvidia.com>
Co-authored-by: Prashant Khodade <pkhodade@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
2026-09-14 22:58:49 +00:00
Max Dubrinsky 8d19308c08 fix(bootstrap): emit RFC 5280 extensions on generated gateway PKI (#3286)
generate_pki minted a CA with no key usage and server and client leaves
with no Authority Key Identifier. RFC 5280 requires both, and verifiers
that enforce it reject the chain: OpenSSL X509_STRICT fails with
"Missing Authority Key Identifier", and Python 3.13 turned that flag on
by default in ssl.create_default_context(). rustls and BoringSSL do not
enforce it, so gRPC clients kept working while an HTTPS client built on
Python 3.13 (for example a platform proxying to an exposed sandbox
service) could not complete a handshake with a pkiInitJob-provisioned
gateway at all. cert-manager PKI was unaffected.

Set keyCertSign and cRLSign on the CA and use_authority_key_identifier
on both leaves, matching what the sandbox L7 CA already does. Add a
test that parses the bundle and asserts the extensions, including that
each leaf AKI matches the CA SKI.

Verified: openssl verify -x509_strict accepts both leaves, and a strict
Python 3.13 client completes an mTLS handshake against a server using
the new bundle where the previous bundle reproduces the failure.

Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com>
2026-09-14 16:49:44 +00:00
Jesse JaggarsandDrew Newberry 02b664bb0d refactor(config): normalize and enforce gateway schema v2 (#2814)
* refactor(config): normalize compute driver field names

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* refactor(config): introduce canonical gateway fields

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* refactor(config): enforce gateway schema version 2

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): preserve compute driver runtime guarantees

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): address schema v2 review regressions

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): complete schema v2 migration safeguards

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(config): expand schema v2 regression coverage

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(config): add schema v2 parity manifest

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): correct parity manifest inventory

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* docs(config): record schema v2 intentional changes

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* docs(config): disposition schema v2 parity gaps

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): add dual schema parity harness

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): establish compute lifecycle parity baseline

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): preserve gateway option compatibility

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): record gateway option parity

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* docs(config): close gateway-wide parity gaps

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(podman): apply configured pids limit

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): validate Podman option parity

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): add Kubernetes option parity harness

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): record Kubernetes option parity

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): disposition VM parity lanes

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): add external driver parity lane

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(e2e): preserve external driver pull policy

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): attest parity artifacts and launches

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): require clean parity build sources

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): bind parity runtime artifacts

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(e2e): use isolated supervisor tags

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(e2e): qualify parity image tags

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(e2e): serve parity supervisor locally

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): isolate parity podman services

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): harden parity evidence provenance

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): pin parity sandbox artifacts

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): attest parity runtime inputs

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): bind parity runtime evidence

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): record compute boundary parity

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): disposition cross-cutting parity lanes

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(packaging): preflight gateway config upgrades

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): preserve rebase integration guarantees

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(ci): isolate temporary git signing config

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): update remaining schema v2 consumers

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(ci): provide e2fs tools to VM tests

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): align preflight with gateway startup

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(vm): preserve rootfs tar configuration

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* chore(config): adopt duration unit constructors

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(packaging): preflight RPM gateway config

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): address driver review findings

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(e2e): require fresh semantic parity evidence

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(docker): update tests for renamed sandbox label

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(gateway): preserve selective driver coverage after rebase

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
2026-09-11 05:00:24 +00:00
ae57979b03 feat(mxc): add Windows ETW-to-OCSF audit trail (#3015)
* feat(mxc): ETW->OCSF audit consumer + Windows OCSF JSONL parity (cp6 P1)

Add a Windows MXC ETW->OCSF audit trail in openshell-driver-mxc: a real-time
Sandboxing-provider ETW consumer that decodes events (TDH), attributes each to
an OpenShell sandbox_id, and maps them to OCSF (lifecycle 6002, config 5019,
process 1007, finding 2004).

cp6 Phase 1 - durable OCSF JSONL audit-file parity with Linux:
- openshell-ocsf: add emit_ocsf_event_routed (populates the event-bridge
  thread-local AND stamps sandbox_id+message in one dispatch) plus public
  set/clear_current_event; OS-aware device (Device::windows/for_current_os) so
  device.os.name reflects the host instead of a hardcoded Linux stub.
- etw_consumer: emit via the routed emit (previously fired a bare info! that
  never populated the bridge, so the structured event was dropped).
- openshell-server: install OcsfJsonlLayer over a synchronous daily-rotated
  appender (durable under force-kill), gated by OPENSHELL_OCSF_JSON, path via
  %PROGRAMDATA%\OpenShell\logs (override OPENSHELL_OCSF_LOG_DIR).
- device.hostname now resolves to the real gateway machine name.

Box-proven on 7F203-MXC-001: JSONL lines == shorthand OCSF rows, all valid
OCSF JSON, per-sandbox attribution intact, disabled state writes nothing.

Signed-off-by: Akber Raza <akberr@nvidia.com>

* feat(mxc): map remaining Sandboxing ETW events to OCSF

Close the last three ETW->OCSF gaps so the audit trail covers the full
set of events the Sandboxing provider emits (12/12):

- ProcessLaunched -> Process Activity [1007] "Launch" (confirmed start;
  carries the real processId/threadId, the twin of CreateProcessInSandbox
  which only has the request + command line).
- SandboxProxyConfigured -> Device Config State Change [5019] (the one
  network-plane setup event; surfaces proxyPort, "no proxy" when 0).
- SandboxConsoleReferencePlumbed -> Device Config State Change [5019]
  (console-handle plumbing).

map_config_state now handles the full config/hardening/setup family and
carries proxyPort/hasConsoleReference/creationFlags as unmapped fields.
Verified on 7F203-MXC-001: 11/12 event types emit OCSF without a proxy
(SandboxProxyConfigured requires proxy config to fire).

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc): seed ETW attribution under registry lock + Device tests

Address CodeRabbit review on !31:

- Prevent stale ETW attribution on a delete/launch race: register the
  wxc-exec pid while holding the registry lock, and bail if the sandbox
  entry is already gone. Previously the attribution key could be seeded
  after `delete` had removed the sandbox, leaving a stale key that could
  misroute later Sandboxing ETW events to a dead sandbox_id. Lock order
  (registry -> attribution) matches the delete path, so no deadlock.
- Add unit tests for the new Device::windows and Device::for_current_os
  constructors to harden Windows/Linux OCSF device parity.

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc-etw): buffer+replay racing events and harden attribution keys

Addresses two ETW->OCSF attribution review items (Shailendra #1, #2).

#2 early-event loss: ETW delivers the sandbox create/config burst the instant wxc-exec starts, which can beat the driver's register_launch (now under the registry lock post-Ready). process_event previously dropped anything unresolved, losing the racing burst. Add a bounded, time-bounded pending buffer (PENDING_MAX=4096, PENDING_TTL=5s): unresolved events are held and replayed once attribution lands, aged-out ones dropped. Consumer switched to a timed recv_timeout(200ms) so the buffer is re-driven after each event and on a tick. Emit path factored into shared emit_resolved().

#1 attribution collisions: a Windows PID is recycled after exit and a command line is commonly identical across sandboxes. register_launch now rebinds by_pid on reuse and clears the stale last_pid_sid hint (warns if the PID still pointed at a different, leaked sandbox); command line is held in by_cmd only while unique and demoted to a new ambiguous_cmds set on a second owner, so a duplicate command refuses to resolve rather than misroute.

Unit tests: buffer replay (direct + cross-link), buffer bound, PID-reuse rebind, duplicate-cmd non-resolution. Box-verified on 7F203-MXC-001 (5 sandboxes, identical cmd -> 5 isolated sandbox_ids, 50/50 OCSF/JSONL, BuffersLost=0).

Signed-off-by: Akber Raza <akberr@nvidia.com>

* docs(mxc-etw): note cmd_line is captured raw with no privacy filtering

Review item #3 (Shailendra): add a PRIVACY NOTE on map_process_launch stating cmd_line is copied verbatim into OCSF process.cmd_line with no redaction, so secrets/PII on a command line land unredacted in the durable audit trail (deliberate audit-fidelity trade-off; treat the log as sensitive). Redaction is owned by an upstream privacy layer, not this path; no general audit-output PII scrubber exists today (openshell_core::secrets [CREDENTIAL] redaction is scoped to the proxy HTTP-target logging, a separate egress path).

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc-etw): open ETW trace on caller thread so start_session reports real status

Review item #4 (Shailendra): start_session previously returned Ok(EtwSession) as soon as the pump thread was spawned, but OpenTraceW ran later inside that thread; if it failed we still handed back a live-looking session and logged 'consumer started' (silent failure = false audit coverage).

Split the two Win32 calls instead of adding a channel handshake (avoids any lost-wakeup/hang risk): the quick, synchronous OpenTraceW now runs on the caller thread (open_trace), and only the blocking ProcessTrace runs on the pump thread (run_trace). start_session returns Err if OpenTraceW fails (reclaiming the boxed Sender so the consumer disconnects, stopping the session, joining the consumer) and returns Ok/logs 'started' only once capture is genuinely open. Opened handle + LoggerName buffer + boxed Sender are carried to the pump via a Send OpenedTrace so they outlive ProcessTrace.

Box-verified on 7F203-MXC-001: consumer started=True, failed-to-start=False, 50 OCSF rows / 50 JSONL, BuffersLost=0 (no regression to capture/emit).

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc-etw): guard pending-event replay against PID recycling

CodeRabbit flagged that drain_resolved() re-resolved buffered events
against the live by_pid map, so if Windows recycled a wxc-exec PID within
PENDING_TTL a stale event from the dead sandbox could be emitted under the
new owner.

Stamp each by_pid registration with its Instant and add resolve_replay(),
used only on the buffered/replay path. It (a) never falls back to the
recycle-/ambiguity-prone by_cmd or last_pid_sid keys, and (b) trusts a PID
match only when the registration is not newer than the buffered event by
more than REPLAY_PID_GRACE (2s) - a recycled PID's registration lands well
outside that window, so the stale event ages out instead of misattributing.
The legitimate #2 seed race (registration lands ~immediately) still replays.

Adds unit tests for the recycle-refusal, in-grace acceptance, and
weak-fallback exclusion.

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc-etw): surface unexpected ProcessTrace termination (review #4)

start_session already returns Err on OpenTraceW failure (runs on the
caller thread since e41a7701), closing the first half of Shailendra's #4.
This closes the second half: ProcessTrace's result was discarded, so if
capture died mid-run the backend had no way to know.

Add a shared CaptureHealth (stopped/stopping/exit_code) between the pump
thread and EtwSession. run_trace now records ProcessTrace's WIN32_ERROR
and, when the pump returns without a deliberate stop, logs at ERROR that
MXC OCSF capture is no longer running. EtwSession::stop() sets `stopping`
before teardown so a normal shutdown isn't misreported, and
EtwSession::is_capture_alive() exposes the state for status/diagnostics.

Box-verified on 7F203-MXC-001: 5 sandboxes, 50 attributed OCSF rows,
JSONL parity 50/50, BuffersLost=0, clean start/stop (no false failure).

Signed-off-by: Akber Raza <akberr@nvidia.com>

* feat(mxc-ocsf): add ETW->OCSF audit-trail example kit; fix proxy-configured message

Add a runnable OCSF audit-trail example under examples/ (run-ocsf-audit.ps1,
mxc-ocsf-audit.toml, ocsf-audit.yaml, README) that spins up sandboxes with the
in-process ETW consumer and egress proxy on, emitting a full OCSF JSONL audit
trail across all four classes (6002/5019/1007/2004).

Fix SandboxProxyConfigured mapping to log "MXC sandbox proxy configured" instead
of a misleading "(no proxy)" when the provider reports proxyPort=0; the event's
presence already indicates proxy configuration. Verified on-box: 26 events, all
mapped ETW event types present.

Signed-off-by: Akber Raza <akberr@nvidia.com>

* feat(mxc-ocsf): clearer audit report + client-safe run-ocsf-audit.ps1

Improve the ETW to OCSF audit-trail example output and make it safe to ship.

Report:
- Add an event-type coverage count ("N of M expected event types fired");
  the denominator auto-adjusts (8 with proxy on, 7 with -NoProxy).
- Split the checklist into expected event types vs anomaly findings
  (ActivityError/FallbackError), which are reported separately and not
  counted toward coverage (a clean run may emit none).
- Verdict is now coverage-based (all expected types must fire) instead of
  the looser "at least 3 OCSF classes".
- Call out the absolute path to the durable OCSF JSONL log prominently.

Client-safety:
- Default -ShareOut to empty (no auto-copy); pass -ShareOut a UNC path to
  opt in. Removes a hardcoded internal share path from a published example.
- Drop internal-team wording ("Hand that zip back for evaluation", "BUNDLE:")
  in favor of neutral "Results bundle:".
- Update README-ocsf-audit.txt to match the opt-in -ShareOut behavior.

Verified on both MXC boxes: 7F203-MXC-001 (base-container) -> PASS, 8 of 8
event types, 26 OCSF events across 4 classes; 7F203-MXC-003 (AppContainer
fallback) -> reduced set as expected, clean output.

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc): configure OCSF audit workloads per sandbox

- remove unsupported gateway-scoped workload fields from the shipped MXC audit example.
- build the command, working directory, and filesystem grant from each run's ShareDir
- pass the workload through --driver-config-json.
- preserve the host CONNECT proxy configuration and conditional audit coverage for the future host_connect_proxy merge
- require the workload output when determining the audit verdict.

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc): omit command arguments from OCSF audit logs

- record only the executable basename for MXC CreateProcessInSandbox audit events
- leave process.cmd_line unset so workload arguments cannot reach shorthand or JSONL logs
- cover tokens, passwords, signed URLs, and PII with a secret-leak regression test
- update the audit example, architecture guidance, and published logging documentation
- preserve ETW attribution and future host_connect_proxy enforcement behavior

Signed-off-by: Akber Raza <akberr@nvidia.com>

* feat(etw): enhance ETW session management with distinct naming for concurrent gateways

* fix(etw): bound the audit queue during overload

- replace the unbounded ETW callback channel with count- and byte-bounded buffering

- keep the ETW callback non-blocking and count records rejected during overload

- emit immediate, rate-limited warnings that identify resulting audit coverage gaps

- make the audit example fail when queue overload causes dropped ETW records

- cover stalled consumers, oversized events, and warning throttling with unit tests

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(etw): harden sandbox audit attribution

- remove command-line and persistent per-PID fallback keys from live and replay resolution

- retire the driver-owned wxc-exec PID before publishing child completion

- retain established identity, activity, and correlation-vector links only for the five-second late-event window

- prevent buffered records from crossing rapid PID retirement and reuse boundaries

- add resolver and lifecycle coverage and document the attribution trust boundary

Signed-off-by: Akber Raza <akberr@nvidia.com>

* chore(mxc): address rebase follow-ups

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(mxc): align OCSF audit example with driver config

- remove unsupported egress proxy settings

- stop requiring the unavailable proxy audit event

- update example documentation for supported event coverage

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(etw): redact command-line secrets in DecodedEtwEvent summary

* fix(etw): enhance PID resolution and event attribution logic for ETW records

* fix(ocsf): restrict gateway-local JSONL sink to Windows/MXC path with opt-in configuration

* address rebase issues

* fix(mxc): address ETW audit review feedback

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(mxc): fail closed across ambiguous PID reuse

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(mxc): bind ETW attribution to process generation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Akber Raza <akberr@nvidia.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Jamie King <jamiek@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
2026-09-11 02:04:26 +00:00
Drew Newberry 38f2aef930 feat(gateway): support selective compute driver builds (#3118)
* feat(gateway): support selective compute driver builds

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(gateway): support selective Windows MXC builds

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-11 00:36:09 +00:00
Drew Newberry 33bbda3d33 refactor(persistence): adopt continuation-token pagination (#3249)
* refactor(persistence): adopt continuation-token pagination

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(pagination): address continuation review findings

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(tui): recover completed list refreshes

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(pagination): address review scalability findings

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(pagination): repair branch validation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(go): use page size in template example

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-11 00:02:07 +00:00
Piotr Mlocek ddc8bba967 ci(windows): add Windows MSVC CI jobs (#2738)
* fix(ci): preserve Windows Rust build cache

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(ci): invalidate empty Windows caches

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* perf(ci): cache Windows builds with sccache

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(windows): restore target directory caching

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* perf(ci): use prebuilt Z3 on Windows

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* perf(ci): layer sccache on Windows target cache

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(windows): split PR checks from main validation

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(windows): separate checks builds and cache seeding

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(ci): simplify Windows build dependency

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(ci): rely on Windows job dependency status

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(ci): use valid opt-in Windows ARM runner

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(windows): keep ARM64 validation local

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(ci): install Clippy for Windows validation

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(windows): focus platform lint coverage

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(licenses): explain bzip2 allowance

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(windows): simplify workflow name

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(windows): allow async platform stub

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(network): make file fingerprints portable

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(windows): lint supported deliverables

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(windows): allow platform-gated lint

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* chore(ci): align Windows cache action with main

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(ci): align Windows validation with prerequisites

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(windows): pin Rust toolchain action

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(ci): use enterprise-approved Windows actions

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(windows): restore strict MSVC validation

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(windows): run Rust tests with nextest

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(ci): normalize nextest lock provenance

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(windows): add native arm64 validation

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(ci): lock nextest for Windows ARM64

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(windows): resolve duplicate MXC authentication method

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* test(conformance): use native absolute paths on Windows

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(windows): address MSVC review feedback

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(windows): simplify cache key names

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(ci): isolate Windows Rust toolchains for stable caches

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(ci): configure Rustup home in runner setup

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(windows): surface sccache server write diagnostics

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(windows): remove temporary cache diagnostics

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(windows): address review feedback

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(windows): reconcile merged driver capabilities

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(deps): preserve AWS-LC-only lockfile

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* chore(deps): allow z3 prebuilt TLS wrapper

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

---------

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-09-10 23:53:01 +00:00
Artem Lytvyn 25021ee31d fix(docker): reclaim sandbox token files on out-of-band removal (#3220)
Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>
2026-09-10 18:23:16 +00:00
Varsha 0357daee31 refactor(proto): isolate gateway storage messages (#3169)
* refactor(proto): isolate gateway storage messages

Move persistence-only protobufs into a server-private versioned package, remove them from generated public SDKs, and gate durable/public schema compatibility with legacy database fixtures.

Closes #3053

Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>

* docs(gateway): sync protobuf schema inventory

Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>

---------

Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>
2026-09-10 16:36:07 +00:00
krishicks 67374efdf8 fix(ocsf): emit schema-valid event identities (#3247)
Previously, every event from a sandbox reused the sandbox ID as its event
ID. Consumers deduplicating security records could mistake separate events
for the same record, and missing device types or empty image objects could
prevent schema validation.

Give each event its own ID, retain the sandbox association separately, and
classify the environment as Other/Sandbox while keeping the OS separate.
Omit unknown container details instead of emitting empty objects.

Security tooling can now distinguish events from the same sandbox and read
their identity consistently after serialization.

Refs #1055

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-09-10 16:35:24 +00:00
John T. Myers f4dc6be4b2 refactor(inference): remove managed inference routes (#3195)
* refactor(inference): remove managed inference routes

Closes #3172

Remove the inference route control plane, inference.local data path, built-in router crate, and SDK surface. Move inference workloads to explicitly imported provider profiles and native endpoints, with migration cleanup and updated tests and documentation.

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(policy): preserve alternate upstream isolation

Restore the provider policy activation guard so legacy OpenAI and Anthropic providers configured for alternate base URLs do not grant egress to the built-in public vendor endpoints.

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-09-09 18:47:22 +00:00
Simon Scatton 48c449d8c8 chore(deps): replace ring with AWS-LC (#3243)
* chore(deps): replace ring with AWS-LC

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* fix(lint): address warnings after dependency upgrades

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* fix(tls): limit provider initialization to reqwest clients

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

---------

Signed-off-by: Simon Scatton <sscatton@nvidia.com>
2026-09-09 17:57:42 +00:00
Evie Howard 6e6b3c8905 refactor(cli): remove local Dockerfile image builds (#3214)
Signed-off-by: Evie Howard <evhoward@redhat.com>
2026-09-09 13:28:17 +00:00
Evan Lezar e4369adcd0 chore(deps): replace serde_yml with noyalib (#3031)
* chore(deps): replace serde_yml with noyalib

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(providers): annotate generic YAML test values

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-09-07 13:22:27 +00:00
krishicks 17171cd933 refactor(otel): unify compute driver tracing (#2995)
Centralize compute-driver RPC descriptors, stream instrumentation, provider
routing, and standalone installation in openshell-otel. Use typed RPC
constants so gateway and in-process driver paths cannot panic on unknown
operation strings or repeat runtime method parsing.

Emit semantic-convention rpc.service and rpc.method attributes, preserve
trace context and resource identity across deployment modes, and route both
RPC boundary and backend crate spans to each selected driver provider. Leave
consumer-dropped watch spans unset while recording observed terminal status,
and avoid reboxing untraced external-driver streams.

Derive each driver tracing identity from Cargo package and crate metadata and
attach its descriptor to the compute-driver registration, keeping provider
selection and target routing tied to the registered implementation. Share
tracing setup and round-trip test support across Docker, Podman, Kubernetes,
and VM, and update the gateway tracing documentation.

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-09-03 15:26:14 +00:00
Evan Lezar 487b26574d test(conformance): add plan-driven continuity verification (#3107)
* refactor(test-guest): compose Ansible provisioner roles

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(conformance): add plan-driven sandbox continuity

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(test-guest): add gateway continuity actions

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(test-guest): add RPM gateway reinstall action

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(test-guest): add RPM gateway upgrade action

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(test-guest): install latest-release RPM baseline

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(test-guest): add gateway upgrade-restart plan

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* ci(conformance): run Fedora gateway upgrade plan

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-09-03 12:05:59 +00:00
Evan LezarandPiotr Mlocek 8e73f1db99 fix(deps): remediate h2 advisory (#3085)
* fix(deps): remediate h2 advisory

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(deps): update h2 to 0.4.19

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

---------

Signed-off-by: Evan Lezar <elezar@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Co-authored-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-09-02 21:03:31 +00:00
Philippe Martin 857af42a16 feat(vm): support corporate HTTP forward proxy egress for microVM sandboxes (#3090)
* feat(vm): support corporate HTTP forward proxy egress for microVM sandboxes

The corporate forward proxy machinery from #1792 is driver-agnostic and
already merged: openshell-supervisor-network implements CONNECT chaining,
NO_PROXY matching, credentials, https:// proxies and corporate CA trust, and
openshell-sandbox exposes it as six argv-only flags. Podman gained the driver
half in #2245/#2512 and Kubernetes in #2633; the VM driver had none of it, so
VM sandboxes on proxy-only networks could not reach any destination requiring
the proxy even when policy allowed it.

The blocking piece was not proxy logic but delivery: the VM guest init script
runs as PID 1 and execs a fixed supervisor command line, and libkrun's
krun_set_exec receives an empty argv, so there was no channel for driver-owned
supervisor arguments. The supervisor's proxy flags deliberately have no
environment fallback, and build_guest_environment merges user-supplied
environment, so the guest env is not a safe transport either.

Add a driver-authored argument file, mirroring the existing init.d manifest:
the driver writes /opt/openshell/supervisor-args into the overlay upperdir on
every launch and the guest reads it verbatim, one argument per line, appending
it to every supervisor exec. It is written even when empty, which is what makes
the channel unforgeable -- the upperdir always shadows the read-only image
layer, so an image can neither supply its own arguments nor disable the
operator's by omitting the file. Because both launch backends exec the same
init script, this covers libkrun and QEMU without touching either.

A microVM has no bind mounts or container secrets, so the credential and CA
bundle are staged into the per-sandbox overlay the way the gateway JWT already
is: credential root-only at 0600, CA at 0644, both rewritten every launch so a
removed setting clears prior material, and both deleted with the sandbox state
directory. This places the credential at rest in the overlay image on the
gateway host, which differs from the Podman secret model and is documented as
an explicit security consideration.

Validation is fail-closed and shared: a new
openshell_core::driver_utils::validate_upstream_proxy_settings holds the
pairing rules the Podman driver established, and both the gateway and the
driver call it so an invalid table names the offending key instead of
surfacing as an opaque driver-readiness timeout.

Guest egress leaves through gvproxy, so a proxy on the gateway host's loopback
is reachable only through host.openshell.internal; the guest to gateway
callback is unaffected.

Closes #3088

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(vm): bound the proxy CA read and scope the host-loopback recipe

Two review findings on the corporate forward proxy support for microVM
sandboxes.

The driver read the operator's proxy_ca_bundle with an unbounded fs::read and
accepted it on a substring match for the PEM BEGIN CERTIFICATE marker. A
special file such as /dev/zero therefore grew driver memory without bound on
every authorized sandbox create, and a PEM block holding invalid DER passed
the host check but contributes no trust anchor in the guest, so every
supervisor would fail after boot with an error attributed to the sandbox
rather than to the setting.

Move the read into openshell-core as read_upstream_proxy_ca_bundle_file: it
reuses the credential reader's bounded-read path (non-regular files rejected
on fstat, size capped, read bounded even if the file grows), then requires at
least one anchor that RootCertStore::add_parsable_certificates accepts. The
supervisor's own reader now delegates to it, so host acceptance and guest
acceptance are the same function and cannot drift.

The published host-loopback recipe was written for libkrun only. gvproxy NATs
host.openshell.internal to the gateway host's 127.0.0.1, but GPU sandboxes run
on the QEMU/TAP backend where that name resolves to the TAP host address and
the driver's own nftables input chain accepts only the gateway port from the
guest — no proxy on the gateway host is reachable there at any bind address,
so an operator following the generic recipe lost all proxy-required egress
while configuration validation succeeded.

Scope the recipe to libkrun in every reference and reject a gateway-host proxy
URL when a launch plan resolves to QEMU, naming the reason, instead of booting
a sandbox whose policy-approved CONNECTs all time out.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(vm): match the QEMU proxy preflight to the selected TAP host

The gateway-host proxy guard added for the QEMU/TAP backend classified the
wrong set of addresses in both directions.

It ran at the top of configure_qemu_launch_plan, before the subnet allocation
that settles plan.host_ip, so it could not compare against the address the
guest actually reaches the host on. An operator pointing https_proxy at the
sandbox's own TAP host address, such as 10.0.128.1, passed the check, and the
driver's nftables input chain — which accepts only the gateway port from the
guest — then dropped every policy-approved CONNECT, which is exactly the
silent timeout the guard exists to prevent.

In the other direction it rejected 192.168.127.254 unconditionally. That
address is special only to libkrun/gvproxy; on QEMU/TAP it is an ordinary
address that may be routable through the guest's masqueraded egress, so the
guard refused a working configuration.

Run the check after the launch plan's network allocation, on both the
freshly-allocated and already-complete paths, and compare IP literals with
that sandbox's selected TAP host. Loopback literals, localhost, and the
documented host aliases that write_host_gateway_aliases seeds to the TAP host
still classify as the gateway host, and the failure names the address. The
gvproxy host-loopback constant returns to being a documentation anchor.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

---------

Signed-off-by: Philippe Martin <phmartin@redhat.com>
2026-09-02 14:56:08 +00:00
Drew Newberry 9ca19e6c80 refactor(compute): decouple gateway driver composition (#2823)
* refactor(compute): decouple gateway driver composition

Move first-party composition and VM process ownership into openshell-gateway, leaving openshell-server backend-independent. Update packaging and build references with the new crate, simplify the compiled-driver boundary, and keep the driver-free gateway path buildable with bundled Z3 tooling.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(telemetry): bound compute driver categories

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(core): keep runtime transport generic

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(compute): complete server driver decoupling

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(compute): preserve driver integrations after rebase

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>

* fix(compute): preserve docker tracing after decoupling

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>

* fix(compute): preserve driver behavior after extraction

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(compute): remove MXC policy side channel

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(compute): separate policy delivery from readiness

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>
2026-09-01 21:13:45 +00:00
Mrunal Patel b4afcd8a43 fix(cli): suppress ANSI color when stdout is not a terminal (#3026)
* fix(cli): suppress ANSI color when stdout is not a terminal

The CLI colorized output unconditionally. owo-colors is built without
its `supports-colors` feature, so `.green()` and friends emitted escape
sequences regardless of destination, and nothing in the CLI read
NO_COLOR. Piping any command through grep or awk matched against bytes
the caller could not see; `forward list` was the case that surfaced it,
where an escape sits immediately before the STATUS word and defeats a
pattern anchored on whitespace.

Add a `color` module holding a process-wide switch resolved once in
run_async, before any output. Command modules import its `Colorize`
trait in place of `OwoColorize`; the method names match, so the ~450
call sites are unchanged, but each consults the switch when it renders
and delegates to owo-colors so the escape bytes stay identical. The two
traits collide by design: importing both in one module is an ambiguity
error, which keeps unconditional coloring from returning.

owo-colors is not the only styled path, and the rest each carry their
own default, so the switch governs them too:

  - tracing_subscriber formats with ANSI on, does no terminal detection,
    and writes to stdout, so `openshell -v ... | ...` leaked escapes the
    same way the tables did. It now takes the setting via with_ansi.
  - indicatif and dialoguer both style through console, which has its
    own detection but cannot learn about --color. Overriding console's
    global switch covers every progress bar and prompt rather than the
    specific ones constructed today. Both the stdout and stderr switches
    are set, since prompts and progress bars draw to stderr.
  - miette renders errors through its own handler, likewise unaware of
    --color, so init installs one built from the setting.

Resolution order: `--color always|never`, then NO_COLOR, then
CLICOLOR_FORCE, then whether stdout is a terminal. The decision is made
against stdout even for stderr text, since stdout is what gets parsed;
`--color always` restores styling when redirecting.

Padding is unaffected — the format spec is forwarded to the inner
Display, so widths measure text rather than text plus escapes.

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(cli): resolve color per output stream

Review feedback on #3026.

Resolving one answer from stdout and handing it to every library meant a
redirected stream inherited the other stream's terminal check. Running
`openshell ... 2> build.log` from a terminal wrote escapes into the log,
because console's stderr switch and miette's handler were both given
stdout's answer. That is worse than the behavior before this branch,
where both libraries did their own per-stream detection.

Resolve `auto` separately for stdout and stderr and hand each library
the answer for the stream it writes to: tracing and console's stdout
switch get stdout, miette and console's stderr switch get stderr. The
owo-colors wrapper is the exception, since its call sites are split
across println! and eprintln! and a Painted value cannot tell which
macro will consume it; it styles only when both streams accept escapes,
erring toward plain text rather than risking a redirected stream.

Existing tests could not catch this: Command::output gives both streams
pipes, so a per-stream decision and a single stdout-derived one look
identical. Add a test that puts stdout on a pty and stderr on a pipe,
which fails when stderr is handed stdout's answer.

Replace CLICOLOR_FORCE with FORCE_COLOR. The clicolors spec does not say
how to treat `0`, and implementations that special-case it disagree with
force-color.org, which keys on presence and non-emptiness only. Using
FORCE_COLOR gives it the same rule as NO_COLOR: set and non-empty means
yes, whatever the value. Nothing depended on CLICOLOR_FORCE, which was
introduced earlier on this branch and never released.

Carry the whole style in an owo_colors::Style rather than dispatching a
local enum through a six-arm match, and merge styles when chaining so
`x.green().bold()` emits one `\x1b[32;1m...\x1b[0m` instead of nesting
two wrappers. No call site styles already-styled text, so merging is
safe; the emitted bytes are shorter and there is a single reset.

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

---------

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
2026-09-01 21:11:31 +00:00
Evan Lezar bb70461878 test(e2e): run conformance in gateway lanes (#2925)
* test(e2e): isolate VM-specific smoke assertions

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(e2e): add portable CLI conformance baseline

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* feat(conformance): add standalone CLI runner

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(e2e): run conformance in gateway lanes

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-09-01 12:36:51 +00:00
Evan Lezar 8a13bc1298 chore(deps): remove legacy rustls webpki path (#3013)
Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-09-01 09:49:01 +00:00
dependabot[bot] 23351771ca chore(deps): bump quinn-proto from 0.11.14 to 0.11.17 (#2986)
Bumps [quinn-proto](https://github.com/quinn-rs/quinn) from 0.11.14 to 0.11.17.
- [Release notes](https://github.com/quinn-rs/quinn/releases)
- [Commits](https://github.com/quinn-rs/quinn/compare/quinn-proto-0.11.14...quinn-proto-0.11.17)

---
updated-dependencies:
- dependency-name: quinn-proto
  dependency-version: 0.11.17
  dependency-type: indirect
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-27 13:33:36 +00:00
bcd517bbe0 feat(driver-mxc): native Windows MXC compute driver + server wiring (#2721)
* feat(driver): add MXC compute driver for Windows isolation sessions

Introduces the openshell-driver-mxc crate implementing ComputeDriver
backed by Microsoft MXC isolation sessions (Windows only). Wires the
new driver into the server's build_compute_runtime dispatch and adds
the Mxc variant to ComputeDriverKind.

Also adds a local protobuf-src stub (tools/protobuf-src-local) to
unblock Windows builds that lack MSYS2/MinGW, and pins the zig
Windows x64 toolchain in mise.lock.

(cherry picked from commit 4f7012224efb18fbfeb47aa87e0cfd3f036f32f0)
Signed-off-by: Jamie King <jamiek@nvidia.com>

* wip(mxc): checkpoint hung-agent work (recon, policy_map embed, A1 wiring, demo artifacts)

Safety checkpoint of uncommitted work from the background agent run that stalled mid-Step-7. Includes: mxc-driver-recon.md (Step 0.5), policy_map.rs (~876L embedded mapper), A1 policy-threading edits across driver.rs/policy.rs/mxc.rs/compute/mod.rs, and examples/ (demo.yaml + mxc-gateway.toml). Not yet verified to compile end-to-end; to be reorganized into the skill's Step 11 commit sequence.

(cherry picked from commit 38e42c03870be3d10e984a54f17a3b61122ff510)
Signed-off-by: Jamie King <jamiek@nvidia.com>

* test(mxc): fix lifecycle and policy unit-test compile drift

- Bring futures::StreamExt into scope for the watch-stream `.next()` call in
  driver::lifecycle_tests so the negative policy proof test compiles.
- Bind a local `mapper` and drop the unused/deprecated NetworkBinary in the
  embedded-mapper network-policy rejection test.

Signed-off-by: Jamie King <jamiek@nvidia.com>
(cherry picked from commit 039b0baf98735ca672dae52be8d3af2417dc0c1a)
Signed-off-by: Jamie King <jamiek@nvidia.com>

* fix(mxc): downgrade missing sandbox_token to debug log

The gateway mints `sandbox_token` only when a sandbox-JWT issuer is
configured. There is no in-sandbox supervisor on MXC (supervisor-removal
design — D1/D4), so no component ever consumes the token; requiring it
on the driver side blocks the demo's `--disable-tls` smoke gateway with a
spurious `invalid_argument`. Log the absence and proceed instead.

Signed-off-by: Jamie King <jamiek@nvidia.com>
(cherry picked from commit cea209797d0edcb1d152251e748900b0a63cca62)
Signed-off-by: Jamie King <jamiek@nvidia.com>

* fix(mxc): keep sandbox Ready after a successful one-shot agent exec

monitor_exec demoted Ready->Error on exit 0 (reason ExecCompleted), so the positive demo (write hello.txt + exit) landed in Error phase. Keep Ready=True (reason AgentCompleted) on success; only non-zero exits go to ExecFailed. Tighten the positive lifecycle test to assert the terminal condition stays Ready=True/AgentCompleted. Verified live via gateway mock round-trip: phase now Provisioning->Ready with no demotion.

(cherry picked from commit 54ab030f03ca0f83d0050d8dd843b633651684ad)
Signed-off-by: Jamie King <jamiek@nvidia.com>

* feat(mxc): add processContainer backend for default-deny enforcement

Add a backend selector to the MXC driver (isolation_session default | process_container). process_container drives a one-shot AppContainer that is genuinely default-deny: a write to any ungranted path is denied by the OS, unlike isolation_session which is grant-only and cannot deny. The lifecycle forks on the flag - isolation_session keeps provision/start/exec, process_container runs a single ephemeral container via run_oneshot.

Also: run-demo.ps1 gains -Backend and hardens the CLI register/create calls; docs corrected to state isolation_session does NOT deny out-of-policy writes and that the negative proof requires process_container.

Verified end-to-end on a real demo box (gateway -> CLI -> driver -> MXC): in-policy write succeeds, out-of-policy write denied (PermissionDenied), OVERALL: PASS.

(cherry picked from commit c6cde3860bbe1b8edb3147d3e840f6bf0ece32d8)
Signed-off-by: Jamie King <jamiek@nvidia.com>

* refactor(driver-mxc): embed policy mapper as a module; remove standalone crate

Adopt the proto-based mapper (map_to_mxc) as the single source of truth,
embedded in openshell-driver-mxc as a Windows-gated `policy_map` module.
Rewire EmbeddedPolicyMapper to call it directly on the typed SandboxPolicy,
deleting the serde_yaml proto->YAML bridge. Move the CLI to a windows-gated
example and the parity tests into the crate; delete openshell-policy-mapper.

- gate policy_map + seam Windows-only (MXC is Windows-only)
- drop serde_yaml; add dev-deps openshell-policy, clap, anyhow
- normalize mapped paths to Windows form in the seam, in one place
- docs: add driver-mxc to AGENTS.md table; correct design doc section 17 test lane

Signed-off-by: Giedrius Burachas <gburachas@nvidia.com>
(cherry picked from commit f22f9c7a25b9a651c5c5cc73f62fb01c4d6c1a8d)
Signed-off-by: Jamie King <jamiek@nvidia.com>

* feat(driver-mxc): implement lossless split_policy for proxy-delegated egress

Signed-off-by: Giedrius Burachas <gburachas@nvidia.com>
(cherry picked from commit 96d6afa0e2dc6a1d54edd12c34a0ceb0a30dadd0)
Signed-off-by: Jamie King <jamiek@nvidia.com>

* feat(driver-mxc): implement Pattern-C governed-egress split through the policy seam

- split_policy: SocketAddr proxy_redirect (replaces bare port), processcontainer
  containment guard naming MXC M1, version preserved in the trimmed proxy_policy,
  delegation reported as an info loss item
- seam: MappedConfig carries trimmed_policy + proxy_addr; MapCtx.egress selects
  the split path; coarse path unchanged when egress is disabled
- driver: [openshell.drivers.mxc] egress_proxy / egress_proxy_addr config,
  validated at create (isolation_session rejected until M1); lifecycle threads
  the redirect into provision and stores the trimmed policy per sandbox,
  emitting an EgressRedirect platform event
- mxc: optional MxcNetwork block (defaultPolicy=block + proxy) in provision and
  one-shot configs; mock records configs for test assertions
- tests: lossless-invariant suite over all example policies (validate +
  serialize round-trip), split lifecycle proof, M1 rejection; example gains
  --split --proxy-addr writing mxc-config.json / trimmed-policy.yaml /
  loss-report.json

Signed-off-by: Giedrius Burachas <gburachas@nvidia.com>
(cherry picked from commit 34d54ad9f25dc6034c3ba15668555ff0d22cddd8)
Signed-off-by: Jamie King <jamiek@nvidia.com>

* fix(driver-mxc): emit MXC network.proxy as {localhost: port}

Verified against the real wxc-exec 0.6.0-alpha via --dry-run: MXC accepts
only the {localhost: N} proxy shape (the form the design doc specifies)
and rejects {host, port} with a parse error. Schema 0.6.0-alpha can
express only a loopback port, so non-127.0.0.1 redirect addresses are now
rejected: split_policy emits an error loss (no proxy block) and the driver
refuses egress_proxy_addr values off 127.0.0.1. Per-sandbox attribution
must use per-sandbox ports until the schema widens.

Signed-off-by: Giedrius Burachas <gburachas@nvidia.com>
(cherry picked from commit edde8d5434571fd5398409204fcf6862672c0793)
Signed-off-by: Jamie King <jamiek@nvidia.com>

* fix(driver-mxc): serialize isolation_session stop/deprovision as unit variants

Empirical contract finding from the real test lane (build 26300.8553,
wxc-exec 2026-06-10): the stop and deprovision experimental blocks are
unit variants in the wxc-exec schema and must serialize as null; sending
{} is rejected with malformed_request (invalid type: map, expected unit),
while provision/start accept maps. The production invoker, the real-lane
test, the probe script, and the e2e runner all sent {} - the driver could
provision and run an agent but never stop or delete an isolation-session
sandbox against this build. Pinned by a unit test.

Signed-off-by: Giedrius Burachas <gburachas@nvidia.com>
(cherry picked from commit 0df39ca0b22ebc21eb965b2b567a5b2cac26af32)
Signed-off-by: Jamie King <jamiek@nvidia.com>

* test(driver-mxc): add Tier-0 mapper coverage matrix with schema drift guard

Three-quadrant, table-driven matrix (38 tests): mappable fields assert
exact MXC output; every OpenShell field MXC cannot express asserts a loss
item with the expected severity (and seam rejection on error); an empty
policy asserts the restrictive default-deny posture for every MXC knob
OpenShell does not control. The handled_fields_inventory drift guard
serializes a fully-populated policy and compares its YAML keys against
the mapper-handled field lists, so a new openshell-policy field fails the
suite until consciously mapped, delegated, or reported as loss.

Re-exports the policy seam types for integration tests; adds serde_yml,
base64, serde_json as dev-dependencies.

Signed-off-by: Giedrius Burachas <gburachas@nvidia.com>
(cherry picked from commit 91807f984a3b16846e35d6ca0d5ec41057aafa3a)
Signed-off-by: Jamie King <jamiek@nvidia.com>

* feat(driver-mxc): inject agent_env into sandbox process.env

Add MxcComputeConfig.agent_env: each entry is either KEY=VALUE (verbatim) or a bare KEY resolved from the gateway host environment at launch, keeping secrets (e.g. inference API keys) out of the config file. Wire it into the agent process so gateway-launched agents can authenticate to cloud endpoints (process.env was previously hardcoded empty). Unit-tested via resolve_agent_env_passthrough_and_host_lookup.

Also add a gateway-driven cloud-inference (T1) test harness: mxc-inference.toml (agent_env + curl agent), inference.yaml policy, and run-inference-test.ps1 which starts the gateway, creates an isolation_session sandbox, runs an authenticated Nemotron call, and bundles redacted results. Documented agent_env in mxc-gateway.toml. Validated end-to-end on the test box (chat HTTP 200 + completion via the gateway).

(cherry picked from commit 94d9e827b8b77e0af9dc654943ea8e7d8400cc9f)
Signed-off-by: Jamie King <jamiek@nvidia.com>

* test(driver-mxc): avoid unsafe env mutation in resolve_agent_env test

Replace std::env::{set,remove}_var (unsafe + racy under parallel test
execution in edition 2024) with a read-only PATH lookup. Preserves all
three behaviors under test and drops the #[allow(unsafe_code)].

(cherry picked from commit ac5766eeab2db4e8cc6fcd8d8a97809edaf3df30)
Signed-off-by: Jamie King <jamiek@nvidia.com>

* fix(mxc): adapt MXC driver to current GitHub OpenShell API

The MXC driver crate was authored on GitLab against an earlier proto/core
API. Adapt it to the API on GitHub main:

- build_capabilities_response no longer takes supports_interactive_session
- DriverSandboxSpec.gpu (bool) is now resource_requirements; detect GPU via
  effective_driver_gpu_count(driver_gpu_requirements(..))
- DriverSandbox gained a `workspace` field
- SandboxPolicy gained `network_middlewares`: pass it through the proxy split,
  emit a loss item on the coarse MXC path, and account for it in the mapper
  drift-guard test

Verified: cargo check + 75 mock-based tests pass (lib 27, examples 10,
policy_mapper_matrix 38).

Signed-off-by: Jamie King <jamiek@nvidia.com>

* feat(server): wire the MXC compute driver into the gateway on Windows

Register openshell-driver-mxc as the Windows-only in-process compute
backend so compute_driver = "mxc" resolves to a working runtime:

- ComputeRuntime::new_mxc, adapted to the current 11-arg from_driver
- mxc_policy_sink A1 side channel, staged in create_sandbox before dispatch
- mxc_config_from_context loader and the Mxc dispatch arm (Windows
  constructs; other targets return an explicit "Windows-only" error)
- Windows-gated openshell-driver-mxc dependency
- Mxc arms for the telemetry, config-file required-fields, and CLI
  reserved-builtin matches to keep them exhaustive/correct

Verified with cargo check --workspace --features openshell-prover/bundled-z3
on x86_64-pc-windows-msvc, stacked on PR #2496.

Signed-off-by: Jamie King <jamiek@nvidia.com>

* fix(driver-mxc): implement GetGatewayListenerRequirements for #2496 base

Signed-off-by: Jamie King <jamiek@nvidia.com>

* test(driver-mxc): add probe-gated real wxc-exec test lane (no mocks)

- tests/wxc_exec_real.rs: ignored-by-default integration tests against a
  real wxc-exec. Six --dry-run contract tests run wherever the binary
  exists (they caught the network.proxy shape mismatch); enforcement
  tests (processcontainer default-deny positive/negative, isolation
  session lifecycle round trip with a deprovision drop-guard) probe the
  backend and SKIP with a recorded reason where it is not live.
- examples/probe-mxc-host.ps1: classifies a host (OS build, --probe,
  per-backend trial) and emits a JSON capability verdict.
- examples/run-mxc-e2e.ps1 + e2e-policies/: scenario runner generalizing
  run-demo.ps1 (fs-rw, fs-readonly, fs-default-deny-empty,
  network-policy-rejected) with PASS/FAIL/SKIP gating and a stale
  OPENSHELL_MXC_MOCK_WXC guard in real mode.
- tasks/windows.toml: windows:test:mxc-real:x64, windows:e2e:mxc,
  windows:e2e:mxc:mock.

Signed-off-by: Giedrius Burachas <gburachas@nvidia.com>
(cherry picked from commit 49afafe892caded59c4df50a9b652011ad97f41c)
Signed-off-by: Jamie King <jamiek@nvidia.com>

* fix(driver-mxc): address PR review feedback

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>

* docs: defer public MXC documentation

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>

* fix(driver-mxc): build Windows capabilities response

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>

* fix(server): gate in-tree tracing on Windows

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>

* test(windows): fix cross-platform test assumptions

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>

* test(windows): verify process and PEM portably

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>

* test(windows): keep lifecycle command in policy

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>

---------

Signed-off-by: Jamie King <jamiek@nvidia.com>
Signed-off-by: Giedrius Burachas <gburachas@nvidia.com>
Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>
Co-authored-by: Prashant K <pkhodade@nvidia.com>
Co-authored-by: Giedrius Burachas <gburachas@nvidia.com>
Co-authored-by: Shailendra Singh <shailendras@nvidia.com>
Co-authored-by: Drew Newberry <385+drew@users.noreply.github.com>
2026-08-27 07:24:04 +00:00
krishicks d0dfb22baf feat(kubernetes): export driver traces over OTLP (#2958)
Mirror the VM, Podman, and Docker driver tracing setup for Kubernetes.
Export standalone driver spans through OTLP/gRPC as the distinct
openshell-driver-kubernetes service, preserve gateway trace context, record
lifecycle operations and gRPC failures, and flush spans on shutdown.

Kubernetes currently runs in-process when selected as a built-in gateway
driver. Use the temporary server-boundary shim shared with Podman and Docker
so traces retain the shape they will have when Kubernetes moves to a
separate process. Move the common ComputeDriver RPC tracing layer into
openshell-otel to keep all drivers aligned.

Propagate the active W3C context through the controller-reserved Sandbox
annotation and enable Agent Sandbox OTLP export in the local k3s workflow.
This connects asynchronous controller reconciliation spans to the originating
OpenShell create trace.

Expose gateway OTLP configuration through Helm and add an Aspire collector
to the local k3s workflow. Extend helm:k3s:forward with OTLP ingest and trace
UI forwarding for Kubernetes and local container gateway development.

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-08-26 21:22:51 +00:00
Evan Lezar 4e992093f8 fix(docker): trace standalone driver over OTLP (#2923)
Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-08-26 06:27:21 +00:00
krishicks e457974a52 feat(docker): export driver traces over OTLP (#2851)
Mirror the VM and Podman driver tracing setup for Docker. Export Docker
driver spans through OTLP/gRPC as the distinct openshell-driver-docker
service, preserve gateway trace context, record lifecycle and asynchronous
provisioning operations, and report gRPC failures.

Docker currently runs in-process when selected as a built-in gateway driver.
Add the same temporary server-boundary shim used by Podman so traces retain
the shape they will have when Docker moves to a separate process. Generalize
the gateway provider selection for both in-process drivers and share the OTLP
collector fixture across Docker, Podman, and VM tracing tests.

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-08-24 18:39:21 +00:00
Simon Scatton 905e99aa2a feat(build): add Nix-native Linux toolchains (#2875)
* feat(nix): add glibc 2.28 development shell

* feat(nix): add musl development shell

* feat(build): use mold in musl development shell

* feat(flake): add nix remote cache

* feat(build): use mold in default development shell
2026-08-24 12:16:00 +00:00
grs 40d1b48666 feat(provider): support for SPIFFE backed token exchange (#1970)
* feat(provider): add ability to request token exchange instead of client credentials as OAuth grant_type

Signed-off-by: Gordon Sim <gsim@redhat.com>

* test(proxy): add further tests for token exchange

Signed-off-by: Gordon Sim <gsim@redhat.com>

* test(provider): add runnable example for token exchange

Signed-off-by: Gordon Sim <gsim@redhat.com>

* test(e2e): cover Podman token exchange grants

Signed-off-by: Gordon Sim <gsim@redhat.com>

* refactor(oauth): extract duplicated functionality from server and supervisor

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(provider): evict nearest-to-expiry entry from intermediate token cache

Signed-off-by: Gordon Sim <gsim@redhat.com>

* doc(supervisor): add podman example for token exchange

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(provider): withhold token-exchange subject credentials

Signed-off-by: Gordon Sim <gsim@redhat.com>

---------

Signed-off-by: Gordon Sim <gsim@redhat.com>
2026-08-24 05:42:30 +00:00
Drew NewberryandEvan Lezar 40f822906c feat(compute): add standalone first-party drivers (#2822)
* feat(compute): add standalone first-party drivers

Build Docker, Podman, Kubernetes, and VM drivers as external binaries and
exercise each through the public compute-driver API. Keep the external E2E
setup complete at introduction, including VM image selection, Kubernetes
post-renderer isolation, supervisor reuse, and scoped Podman coverage.

External Kubernetes endpoints support shared and managed workspace modes.
Operator mode remains restricted to the in-process driver because gateway
authentication and the driver must share a dynamic namespace allowlist.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci(e2e): run managed and external drivers independently

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(compute): cover external driver socket contract

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* ci(e2e): install bundled Z3 build dependency

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(server): gate in-tree driver tracing

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>
Co-authored-by: Evan Lezar <elezar@nvidia.com>
2026-08-21 02:49:55 +00:00
Drew Newberry ef296806f5 feat(sandbox): add canonical main process (#2726)
* feat(sandbox): add canonical main process

Closes #2710

Persist and supervise one canonical workload per sandbox, attach sandbox connect to its retained session, and make every unexpected main-process exit terminal.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(sandbox): simplify canonical main process contract

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): preserve legacy VM main compatibility

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): preserve main status across driver updates

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): satisfy macOS process lint

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): gate Linux exit acknowledgement publisher

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(sandbox): simplify main process plumbing

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(supervisor): make controlling tty ioctl portable

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(supervisor): initialize canonical process environment

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* perf(supervisor): optimize retained main session

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(sandbox): detach main session on ctrl-c

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): use explicit main detach keys

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(dev): atomically stage Docker supervisor

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(test): align Docker main environment assertion

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(sdk): expose canonical main process fields

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-08-20 22:18:23 +00:00
John T. Myers b2ea81822b feat(network): enable Docker and Podman policy DNS and transparent TCP (#2723)
* feat(network): enable Docker transparent TCP egress

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(e2e): cover Docker transparent TCP egress

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* feat(network): correlate transparent TCP audit events

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(examples): add transparent TCP Redis demo

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(examples): demonstrate blocked TCP connections

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(examples): focus Redis demo audit output

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): close transparent TCP policy bypasses

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(sandbox): reject unsupported TCP policy reloads

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(ci): satisfy Linux transparent TCP lints

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* feat(podman): enable transparent TCP egress

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(podman): permit policy DNS port binding

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(e2e): use qualified transparent TCP hostname

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(podman): preserve exact policy DNS names

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(podman): route policy DNS over TCP

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(dns): serve multiple TCP queries per connection

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(network): explain native DNS and TCP egress

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(sandbox): reconcile runtime reload with upstream

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): harden transparent DNS capture

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(podman): preserve resolver behavior for native tcp

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(network): clarify native tcp runtime constraints

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): remove unused transparent tcp pin

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): admit redirected transparent tcp

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): restore podman transparent networking

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(podman): permit alpine busybox binaries

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(podman): use portable alpine keepalive

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(podman): build musl networking fixture

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(podman): isolate musl DNS probe

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(podman): keep privileged port capability dropped

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): preserve transparent TCP port 53

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): report synthetic pool pressure by family

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(podman): bind tcp fixtures before readiness

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(podman): grant fixture low-port bind

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-08-20 19:26:53 +00:00
krishicks 2c0adf4868 feat(podman): export driver traces over OTLP (#2782)
Mirror the VM driver tracing setup for Podman. Export standalone driver
spans through OTLP/gRPC as the distinct openshell-driver-podman service,
propagate W3C context across ComputeDriver RPCs, record bounded RPC names
and failures, and flush buffered spans during graceful shutdown.

Podman still runs in-process when selected as a built-in gateway driver.
Add a temporary tracing shim that partitions gateway and Podman spans by
target into separate tracer providers while preserving their shared trace
and parentage. The shim also emits the same ComputeDriver server boundary
that the tonic layer emits out of process, keeping the observable trace
shape stable when Podman is eventually extracted.

Trace container create preparation, image and storage setup, lifecycle
operations, and cleanup. Document the service boundary and cover it with
isolated and repeated tracing tests.

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-08-20 18:43:16 +00:00
John T. Myers 0c6a3443ec feat(network): add policy DNS correlation store (#2713)
* feat(network): add policy DNS correlation foundation

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(network): describe dormant policy DNS boundary

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): harden policy DNS publication

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): audit policy DNS failures

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): preserve DNS answer order

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-08-20 18:17:57 +00:00
John T. Myers 4d7f402ce2 feat(policy): establish direct TCP egress foundation (#2711)
* feat(policy): accept explicit tcp endpoint protocol

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* refactor(network): snapshot authoritative egress decisions

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(policy): document explicit tcp protocol

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(policy): defer transparent TCP release guidance

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* chore(go): regenerate sandbox protobuf bindings

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): complete tcp egress foundation

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(policy): document explicit tcp contract

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): fail closed on authorization errors

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(podman): fence delayed exit events before restart

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(policy): validate network endpoint destinations

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(providers): opt in tcp credential fixture

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(policy): require dns host for transparent tcp

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-08-20 16:43:09 +00:00
Giuseppe Scrivano d51a653f9c feat(driver-podman): add userns config (#2562)
* refactor(driver): extract shared supervisor binary helpers

Move supervisor binary extraction, caching, and validation helpers from
the Docker driver into openshell-core::driver_utils so both Docker and
Podman drivers can reuse them.

Moved helpers: extract_first_tar_entry, write_cache_binary_atomic,
supervisor_cache_path, temp_extract_container_name, and
validate_linux_elf_binary.

The shared extract_first_tar_entry gains entry-type and empty-payload
checks that the Docker-local version lacked.  supervisor_cache_path
takes a driver_subdir parameter so each driver caches under its own
namespace (docker-supervisor vs podman-supervisor).

Signed-off-by: Giuseppe Scrivano <gscrivan@redhat.com>

* feat(driver-podman): add userns config

Add a `userns` option to the Podman compute driver that maps to
Podman's user namespace modes.  The mode string is split on the first
colon into the API's `nsmode` and `value` fields so parameterized
values like `auto:size=65536` and `keep-id:uid=1000,gid=1000` are
forwarded correctly.  When the mode is `auto`, the container spec
also sets `idmappings.AutoUserNs = true` as required by the API.

An allowlist validates the mode at startup: `auto` and `keep-id`
accept optional parameters; `host`, `private`, and `nomap` reject
them; everything else is an error.

Podman image volumes use overlay mounts internally and the kernel
does not support idmapped mounts on overlay (`mount_setattr` returns
EINVAL).  When userns is configured (any mode except `host`), the
driver extracts the supervisor binary from the image to a host-side
cache and bind-mounts it instead of using an image volume.

Configurable via TOML `userns = "auto"`, CLI `--userns`, or
environment variable `OPENSHELL_PODMAN_USERNS`.

Signed-off-by: Giuseppe Scrivano <gscrivan@redhat.com>

---------

Signed-off-by: Giuseppe Scrivano <gscrivan@redhat.com>
2026-08-15 17:10:47 +00:00
Piotr Mlocek 44bf0df485 feat(middleware): inspect WebSocket text messages (#2477)
* feat(middleware): inspect websocket text messages

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): address websocket review feedback

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(network): bound websocket message assembly

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(network): harden websocket upgrade lifecycle

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* refactor(middleware): unify in-process and remote transports

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* feat(middleware): support regex websocket redaction

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): bound persistent streaming sessions

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): accept websocket sequence gaps

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* refactor(middleware): refine websocket introspection contract

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): clarify websocket preflight lifecycle

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): clarify websocket coverage semantics

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(network): type websocket frame failures

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(network): return 503 when middleware admission is exhausted

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): align streaming API contract

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): clarify WebSocket event result scope

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(rfc): simplify middleware revision history

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* feat(examples): add WebSocket content guard support

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): unify binding payload limits

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* refactor(middleware): align payload limit terminology

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): address websocket review feedback

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(network): address websocket review findings

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* test(network): allow Linux handler setup in preflight regression

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(network): harden websocket relay finalization

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(network): inspect compressed websocket messages

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* test(network): stabilize compressed websocket regressions

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(go-sdk): regenerate middleware protobuf binding

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): clarify websocket skip lifecycle

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): address WebSocket review feedback

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

---------

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-08-14 21:51:42 +00:00
Derek Carr 59479f492a feat(k8s): add namespace-per-workspace support (RFC 0011 Phase 3) (#2656)
* feat(k8s): add namespace-per-workspace support (RFC 0011 Phase 3)

Implement three workspace namespace modes for the Kubernetes compute
driver: shared (default, preserves current single-namespace behavior),
managed (auto-creates/deletes namespaces per workspace), and operator
(pre-provisioned namespaces with dynamic discovery via label selector
or drop-in allowlist file).

Key changes:
- WorkspaceMode enum and namespace resolution in driver config
- Managed namespace lifecycle with ServiceAccount and OpenShift SCC
  annotation propagation
- Cluster-wide sandbox CR watchers for managed/operator modes
- NamespaceValidator (Exact/Prefix/Allowlist) for SA token auth
- Workspace-aware credential secret storage
- Helm ClusterRole for multi-namespace RBAC
- Gateway config, architecture, and reference docs

Signed-off-by: Derek Carr <decarr@redhat.com>

* test(k8s): add e2e tests for workspace namespace modes

Add end-to-end tests for managed and operator workspace modes
introduced in RFC 0011 Phase 3. The managed mode tests verify
namespace creation with correct labels, ServiceAccount provisioning,
sandbox CR placement, and namespace survival with remaining sandboxes.
The operator mode tests verify rejection of unlabeled and nonexistent
namespaces. The positive operator path (sandbox in labeled namespace)
is known to fail due to an RBAC gap and will be addressed separately.

Also fixes Helm 4 compatibility: move SPDX license headers inside
conditional guards in 8 chart templates to prevent empty comment-only
documents, and fix a trailing whitespace trimmer in clusterrole.yaml
that concatenated the license header with apiVersion.

Adds cleanup sweep in with-kube-gateway.sh to remove managed and
operator namespaces before Helm uninstall, and mise tasks for running
each mode independently.

Signed-off-by: Derek Carr <decarr@redhat.com>

* feat(k8s): add operator namespace label watcher

Spawn a background kube::runtime::watcher in the K8s driver that
watches namespaces matching the configured label selector and populates
the OperatorNamespaceAllowlist at runtime. The driver owns the
allowlist and exposes its Arc so the server can share the same set with
the SA token authenticator.

create_sandbox now gates pod creation on the allowlist in operator
mode — workspaces whose namespace is not yet labeled are rejected at
resource render time rather than silently proceeding. Workspace
lifecycle itself is unaffected; only sandbox (resource) creation is
gated.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(k8s): harden operator mode and address review findings

Close the fail-open gap in operator mode when only
operator_namespace_file is configured: the allowlist is now created
unconditionally in operator mode (fail-closed from startup).

Implement the namespace file watcher using the notify crate, following
the TLS hot-reload pattern (parent-directory watch, 1s debounce,
ConfigMap symlink-swap safe). The file format is a JSON array of
namespace name strings.

Additional fixes from the 10-reviewer audit:
- Change allowlist rejection from InvalidArgument to FailedPrecondition
  so callers know the request may succeed later once the namespace is
  provisioned.
- NamespaceValidator::Allowlist now holds the OperatorNamespaceAllowlist
  newtype instead of a raw Arc<RwLock<BTreeSet>>, eliminating silent
  denial on RwLock poison.
- Verify LABEL_MANAGED_BY and LABEL_GATEWAY_ID ownership before
  deleting a managed namespace.
- Replace fixed 5s sleep in operator e2e test with a 30s poll loop.
- Add Helm validation for workspaceMode values.
- Fix Helm README type column and description for operator fields.
- Add insert/remove methods to OperatorNamespaceAllowlist; label
  watcher now uses them instead of reaching through shared().
- Reject configs with both operator_namespace_label and
  operator_namespace_file set.

Signed-off-by: Derek Carr <decarr@redhat.com>

* feat(k8s): add workspace-level compute driver RPCs and harden RBAC

Decouple namespace lifecycle from sandbox lifecycle by adding
EnsureWorkspace/DeleteWorkspace RPCs to the ComputeDriver service.
Namespace creation now happens before credential storage and namespace
deletion happens on workspace delete, fixing credential storage in
managed workspace mode.

- Add EnsureWorkspace and DeleteWorkspace proto RPCs with
  implementations across all compute drivers (K8s managed delegates to
  ensure_namespace/delete_namespace_if_empty; others no-op)
- Wire ensure_workspace into provider create/update/refresh paths so
  the namespace exists before the credential driver writes secrets
- Wire delete_workspace into workspace deletion for cleanup
- Remove delete_namespace_if_empty from sandbox deletion path
- Scope ClusterRole secrets access to non-shared workspace modes
- Add TODO for TLS cert hot-reload in sandbox gRPC client
- Harden e2e tests with control-plane sandbox resolution assertions
- Fix docker image save --platform flag for OCI index manifests

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(k8s): address re-review findings and add test coverage

- Use server-side apply for TLS secret sync (fixes second sandbox
  creation failure when TLS is enabled)
- Scope gateway-ID label selector unconditionally across all workspace
  modes (fixes operator reads/watches/deletes seeing foreign sandboxes)
- Validate operator allowlist in EnsureWorkspace and DeleteWorkspace
  RPCs (prevents credential writes to namespaces outside the allowlist)
- Extend ClusterRole secrets patch+delete to all non-shared modes with
  credential driver enabled (fixes operator credential storage RBAC)
- Validate namespace ownership on 409 conflict in ensure_namespace
  (prevents adopting unowned namespaces in managed mode)
- Replace delete_namespace_if_empty with unconditional delete_namespace
  letting Kubernetes cascade cleanup (fixes stuck terminating CRs)
- Strengthen NetworkPolicy TODO to cover both managed and operator modes
- Extract selector and ownership logic into testable free functions
- Add unit tests for gateway-ID selectors and namespace ownership
- Add Helm ClusterRole RBAC tests for operator credential driver

Signed-off-by: Derek Carr <decarr@redhat.com>

* ci(k8s): add workspace managed and operator mode e2e to CI

Wire the existing e2e:kubernetes:workspace-managed and
e2e:kubernetes:workspace-operator mise tasks into the branch-e2e
workflow so they run alongside the other core Kubernetes e2e suites.
Both are gated by run_core_e2e and included in the Core E2E result
gate.

Signed-off-by: Derek Carr <decarr@redhat.com>

* test(k8s): add e2e tests for workspace namespace modes

Add 7 new e2e tests covering workspace namespace lifecycle, TLS secret
copying, ownership conflict detection, DNS-1123 validation, operator
namespace preservation, and dynamic label watcher behavior. Fix async
sandbox deletion race condition in existing tests by polling sandbox
list instead of asserting immediately after delete.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(k8s): grant secrets/patch unconditionally and backfill gateway-id labels

Address two review findings:

1. RBAC: server-side apply (PATCH) is used for TLS secret sync in
   multi-namespace modes, but the ClusterRole only granted patch when
   the kubernetes-secrets credential driver was enabled. Grant patch
   unconditionally for non-shared modes since TLS sync always needs it;
   keep delete gated on the credential driver.

2. Upgrade safety: the new gateway-id label selector would orphan
   legacy Sandbox CRs that predate its introduction. Add a startup
   backfill in shared mode that patches any managed Sandbox CR missing
   the gateway-id label before the driver begins serving requests.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(k8s): address workspace namespace review findings

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(k8s): address follow-up review findings

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(k8s): preserve workspace lookup after rebase

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(helm): allow managed secret creation

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(k8s): stop pods in workspace namespace

Signed-off-by: Derek Carr <decarr@redhat.com>

* test(k8s): scope pod deletion check to v1alpha1

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(k8s): address workspace namespace review findings

Signed-off-by: Derek Carr <decarr@redhat.com>

---------

Signed-off-by: Derek Carr <decarr@redhat.com>
2026-08-14 21:22:41 +00:00
Piotr Mlocek bdabb54cb3 fix(security): authenticate extension services (#2638)
* fix(security): authenticate extension services

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* feat(extension-core): verify gateway JWTs

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* refactor(extension-core): keep inbound verification external

This should become an extension SDK package.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(security): harden the extension authentication contract

Follow-up hardening on the alpha extension authentication mechanism.

Claim contract:
- Extension tokens carry an explicit `typ` of `openshell-ext+jwt`. They
  share a signing key with sandbox-to-gateway admission tokens and were
  otherwise separated by audience alone, so a verifier that neglects to
  check `aud` could accept a gateway credential. The header is a second,
  independent discriminator.
- Publish OIDC-shaped discovery at `/.well-known/openid-configuration`
  so a service configured with only the gateway URL can learn the exact
  expected issuer and the JWKS location. It is shaped, not compliant:
  `issuer` is the gateway identity, not the serving URL.

Audience agreement:
- `MiddlewareManifest` and `InterceptorManifest` gain `expected_audience`.
  The audience is otherwise configured independently on each side of the
  boundary, where a mismatch surfaces only as an opaque authentication
  failure on every call. OpenShell now compares the two and fails at
  startup. An empty field keeps the check off for existing services.

Compatibility:
- Add `allow_insecure_transport` per registration. Enabling gateway JWT
  signing previously made any plaintext endpoint a hard startup failure,
  including the endpoint form used in our own documentation. The opt-out
  attaches no credential, is refused by the gateway if a supervisor asks
  for one, and warns at every startup.
- Make the transport requirement kind-aware. A middleware endpoint must
  be reachable from every sandbox supervisor, so only interceptors may
  use a gateway-local Unix socket.

Credential lifecycle:
- Replace the process-global slot map with a supervisor-owned
  `ExtensionCredentialStore` shared explicitly across the gateway
  connections the supervisor opens, removing test-order coupling.
- Rotate only when a credential is missing or has passed four fifths of
  its lifetime. Configuration polling ran every ten seconds against
  fifteen-minute credentials, so each poll re-ran gateway effective-policy
  resolution and re-minted the gateway token.
- Bound credential minting per sandbox, since each request resolves the
  caller's effective policy.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs: record alpha extension authentication in RFC appendices

Restore the RFC 0009 and 0010 bodies to their accepted text and move
every extension-authentication update into appendices instead. An RFC
records a decision at a point in time; superseding detail belongs
alongside it rather than rewritten into it.

RFC 0009's appendix carries the shared contract: claims, authorization,
key distribution, the `allow_insecure_transport` replacement for the
body's `allow_insecure`, and residual risks. RFC 0010's records only
what differs for interceptors and links to it. The existing
protocol-extensions appendix, which parked the phase 2 transport
question, now points forward to what was built.

Also document the audience handshake, the discovery endpoint, the
`typ` requirement, and `jti` replay guidance in the extensibility and
gateway configuration pages, and correct the middleware transport
guidance: middleware endpoints must be reachable from sandbox
supervisors, so Unix sockets are not an option there.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* refactor(extension-core): abstract extension server trust

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(core): update middleware manifest example

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(extension-auth): preserve unsigned gateway compatibility

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(extension-auth): reject cross-domain token replay

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

---------

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-08-14 20:34:10 +00:00
LR90 d0c6dc3fd8 feat(kubernetes): support corporate upstream proxy (#2633)
* feat(kubernetes): support corporate upstream proxy

Signed-off-by: loveRhythm1990 <qiuweimin@126.com>

* fix(kubernetes): reject proxy_auth_secret_key values Kubernetes cannot create

Gateway validation accepted proxy_auth_secret_key values that Kubernetes
rejects when creating the Secret (keys longer than 253 bytes, or the
reserved "."/".." names), turning an invalid deployment setting into
repeated sandbox Pod-provisioning failures instead of a startup error.
Reject them in validate_upstream_proxy_config so they fail closed at
gateway startup.

Signed-off-by: loveRhythm1990 <qiuweimin@126.com>

* docs(skill): add corporate upstream proxy checks to debug-openshell-cluster

Add a Kubernetes corporate upstream proxy troubleshooting section covering
rendered [openshell.drivers.kubernetes] configuration, credential Secret
volume events, supervisor arguments and mounts confined to the network-
supervising container, and proxy reachability.

Signed-off-by: loveRhythm1990 <qiuweimin@126.com>

---------

Signed-off-by: loveRhythm1990 <qiuweimin@126.com>
2026-08-14 16:16:43 +00:00