Commit Graph
174 Commits
Author SHA1 Message Date
pkhodade-NV 9ceec26d42 fix(mxc): redact injected secrets from gateway diagnostics (#3853)
Redact injected environment values and the per-sandbox proxy password in
captured MXC output and decoded relay launch-failure diagnostics before
logging or publishing sandbox failure status.

Match the original text and redact the union of overlapping occurrences.
Preserve control-channel payloads and avoid allocating for unmatched text.
Add regression coverage and document exact-match and length limits.

Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
2026-10-01 10:02:55 -07:00
Prekshi Vyas bd49d45cb4 docs(mxc): document Windows host preparation (NVBug 6842834) (#3892)
Document standalone MXC executable placement and condition system-drive ACL preparation on the AppContainer + DACL tier and probe recommendation. Explain the persistent metadata-only grant and link upstream verification and rollback guidance.

NVBug: 6842834

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>
Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
2026-09-29 15:30:25 -07:00
pkhodade-NVandShailendra Singh 9e397e8a43 fix(mxc): reject cpu/memory limits instead of silently discarding them (#3548)
* fix(mxc): reject cpu/memory limits instead of silently discarding them

CreateSandbox accepted --cpu/--memory and reached Ready with no Job
Object enforcement and no diagnostic, leaving the SDD's T11
host-exhaustion mitigation silently unmet. MXC's schema does not
expose CPU rate control or memory limiting outside the WSLC backend,
so reject requests carrying cpu/memory limits synchronously at
CreateSandbox, matching the existing fail-closed GPU rejection.

NVBug 6782894

Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>

* fix(server): validate sandbox resource quantities

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>

* docs(skill): clarify MXC resource limit behavior

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>

* chore(go): regenerate protobuf bindings

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>

* fix(go): align generated protobuf comments

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>

---------

Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Co-authored-by: Shailendra Singh <shailendras@nvidia.com>
2026-09-23 11:16:54 -07:00
pkhodade-NVandShailendra Singh 7305209ac7 fix(mxc): honor the generic sandbox create -- <COMMAND> syntax (#3553)
* fix(mxc): honor the generic sandbox create -- <COMMAND> syntax

sandbox_config only ever read the command from the MXC-specific
driver_config (--driver-config-json), never DriverSandboxSpec.command
-- the portable field the CLI's documented, driver-agnostic
`sandbox create -- <COMMAND>` syntax actually populates, and the same
field every other compute driver honors. A caller following that
syntax got "driver_config.command must contain a non-empty
executable", a message that reads as if no command was supplied at
all.

Add spec.command as a fallback source when no driver_config is
present, and reword the empty-command error to name both ways to
supply one.

NVBug 6782884

Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>

* docs(mxc): document generic sandbox commands

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>

---------

Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Co-authored-by: Shailendra Singh <shailendras@nvidia.com>
2026-09-22 13:49:45 -07:00
Prekshi Vyas 0b8f3821e2 fix(network): normalize Windows binary paths (NVBug 6782969) (#3482)
* fix(network): normalize Windows policy binary paths

Match Windows executable identities using a stable case-insensitive, separator-normalized representation across policy data, L4 input, and L7 relay evaluation. Preserve exact matching on other platforms and keep the original path for hashing and filesystem access.

NVBug 6782969

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* fix(network): harden Windows binary matching

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* fix(ci): scope Windows relay test imports

* fix(network): harden Windows binary path matching

---------

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>
2026-09-22 13:47:25 -07:00
pkhodade-NV 27a36b4004 fix(mxc): reject a non-absolute wxc_exec_path at gateway startup (#3497)
* fix(mxc): reject a non-absolute wxc_exec_path at gateway startup

wxc_exec_path is the binary that builds every sandbox, but nothing
validated it before spawning: a relative value (including the shipped
default, a bare "wxc-exec.exe") let PATH-lookup or working-directory-
relative resolution execute a decoy binary with the gateway's identity
instead of the approved wxc-exec, turning the containment mechanism
itself into an arbitrary-code-execution primitive.

Add MxcComputeConfig::validate_configuration, wired into the existing
(previously no-op) compute-driver config preflight, rejecting an
empty or non-absolute wxc_exec_path with a clear diagnostic. Change
the default from the relative "wxc-exec.exe" to an empty string so
the field must be explicitly configured -- no usable-but-insecure
fallback survives. Update the architecture doc's stale "else PATH"
discovery claim to match.

Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
(cherry picked from commit d4192a0072f8f0ca0a4035f10cae07889c2b578f)

* test(mxc): cover wxc_exec_path enforcement at the gateway boundary

MxcComputeConfig::validate_configuration already had unit coverage,
but that only proves the validation function itself is correct -- it
says nothing about whether MxcFactory::validate_config (src/lib.rs)
still calls it. Before this fix, that factory method discarded the
parsed config entirely, so a regression back to that no-op shape would
leave every unit test passing while a relative wxc_exec_path again
reached gateway startup.

Add an integration test that spawns the actual compiled
openshell-gateway binary through its config preflight subcommand,
exercising the real chain: CLI parsing, TOML loading, driver
selection, MxcFactory::validate_config, and
MxcComputeConfig::validate_configuration. Assertions check only
pass/fail, not message content: run_effective_config_preflight
replaces any validation failure with a generic message whenever a
config file is used, to keep file-sourced values out of preflight
diagnostics -- pre-existing, deliberate, and covered by its own tests.

Also document the new required-and-absolute wxc_exec_path constraint
in the gateway config reference and the driver README, which
previously only showed example values without stating the
requirement.

Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>

---------

Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
2026-09-22 13:39:43 -07:00
Prekshi VyasandShailendra Singh 651ed7e03d NVBug 6783374: make MXC HTTPS L7 qualification authoritative (#3479)
* test(mxc): qualify HTTPS L7 enforcement (NVBug 6783374)

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* docs(mxc): clarify real qualification failures

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* docs(mxc): align real test skip semantics

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>

---------

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>
Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Co-authored-by: Shailendra Singh <shailendras@nvidia.com>
2026-09-22 12:38:07 -07:00
Prekshi Vyas d78430c66c test(mxc): verify ProcessContainer token isolation (#3430)
Adds real-MXC regression coverage for ProcessContainer token isolation, including an unsandboxed SCM positive control, and documents the AppContainer authorization model.
2026-09-21 15:25:31 -07:00
Drew Newberry 49b4f0eb7f feat(mxc): add UI policy, credentials, relay lifecycle, and proxy auth
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-17 12:06:06 -07:00
Shiju 8751e35e28 fix(supervisor-network): reject malformed OPA policy containers (#3337)
Validate raw OPA object and array containers before normalization and
access-preset expansion can skip malformed values. Return one fixed
structural error without embedding authored policy data.

Preserve versionless and runtime-only OPA data and existing semantic
validation. Cover initial string/file loading, middleware callback order,
valid deny-rule enforcement, and rejected reload state and generation.
Document the loader contract and its engine-local rejection behavior.

Signed-off-by: Shiju <shiju@nvidia.com>
2026-09-15 20:00:51 +00:00
Seth Jennings dbe36eaf85 fix(security): harden Vault credential transport (#3329)
Reject non-loopback plaintext Vault endpoints, disable redirects, and support private CA bundles without weakening hostname verification. Update Helm configuration, documentation, operator skills, and regression coverage for OSSR-002.

Signed-off-by: Seth Jennings <sjenning@redhat.com>
2026-09-15 19:56:09 +00:00
Mrunal Patel 39cf4823f7 feat(api): add structured gateway errors and SDK decoding (#3313)
* feat(api): expose structured gateway errors across SDKs

Refs #3051. Add standard validation, conflict, and retry details; preserve raw transport status in Rust, Go, TypeScript, and Python; document status and recovery guidance.

This is the structured-error foundation only. Mutation result shapes, allow_missing, durable request deduplication, and exec retry semantics remain follow-up work.

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(python): preserve wrapped RPC cleanup handling

Inspect the original gRPC call when handling missing sandboxes during deletion waits and managed cleanup. Add intercepted cleanup regressions and clarify the error-wrapper migration contract.

Addresses the cleanup review on #3313; part of #3051.

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

---------

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
2026-09-15 17:51:11 +00:00
Mrunal Patel b799fccb8b fix(auth): harden OIDC trust root retrieval (#3332)
* fix(auth): harden OIDC trust root retrieval

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(e2e): pass OIDC HTTP acknowledgement value

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

---------

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
2026-09-15 17:00:45 +00:00
26f2f96393 feat(mxc): add Windows host proxy for MXC sandbox network egress (#3163)
* Implement Windows host proxy integration and update dependencies for OpenShell

* Update README and gateway config to clarify egress proxy address handling and allocation

* Refactor ProxyIdentityMode to return Result for static_binary and add tests for binary path and SHA256 hash

* Enhance platform_hosts_path for Windows to use SystemRoot and improve error handling for hosts file reading

* Refactor FileFingerprint to use Option for mtime and ctime, simplifying metadata handling

* Add conditional compilation for Windows host module

* add unit tests for OPA policy evaluation and identity handling

* remove openshell-supervisor-network from unsupported driver package test exclusion list

* feat(mxc): enable host proxy TLS state generation

Generate per-sandbox TLS state for the MXC host proxy so HTTPS L7 enforcement can use the same MITM path as Linux. Grant generated CA material to the MXC process and inject standard trust env vars, while matching Linux behavior by disabling TLS termination on CA setup failure and relying on proxy fail-closed handling.

* fix(docs): remove outdated notes on governed egress from docs

* fix(tests): update TLS environment variable paths to use temporary directory

* fix(examples): make run-mxc-e2e harness correct and orphan-free

The MXC e2e harness never actually exercised the fs scenarios: it started
the gateway once and patched agent_command per scenario AFTERWARDS, so the
running gateway kept launching the default demo agent (not shipped in the
kit) and every fs scenario failed with CreateProcessW error:2. It also
scored on the `sandbox create` exit code (non-zero due to the harmless
interactive attach), wrote sandbox records to the persistent gateway DB
(leaving orphans that collided on later runs), and its deny scenarios never
proved denial.

Changes:
- Start a FRESH gateway per scenario so each scenario's agent_command is
  actually loaded (root cause of CreateProcessW error:2).
- Score by on-disk artifact / expected outcome, not `sandbox create` exit.
- Real deny assertions: a control write to a granted path must succeed
  (proves the agent ran) while the denied write must be absent. fs-empty
  probes an ungranted out-of-share path (share_dir is mapped rw by design).
- Run the gateway on an ephemeral in-memory DB (sqlite::memory:) so the
  harness never writes to the persistent store and cannot leave orphan
  sandbox records; also use unique per-run sandbox names + pre-delete.
- Fix the process_container probe: use a real cwd + absolute cmd.exe
  (canonical wxc-exec does not expand %TEMP% -> 0x8007010B).
- Fix summary counts (@() so a single FAIL is counted and exit is non-zero).

Verified PASS=4 FAIL=0 on 7F203-MXC-003 (no BaseContainer velocity keys)
using a canonical wxc-exec build (AppContainer fallback).

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(e2e): probe timeout is milliseconds (10ms->30000ms)

MXC process.timeout is wall-clock ms (wire.rs). The 10 value meant 10ms,
which the base-container tier (7F203-MXC-001/.181) enforced strictly and
timed the probe out. AppContainer path (.18/-003) happened to slip under
it. Bump to 30000ms so the process_container preflight probe is reliable
across both tiers.

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc): use native paths in real runtime probes

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc): make processcontainer work with mxc-latest-released wxc-exec

Three fixes to support the release wxc-exec binary (BaseContainer dispatcher)
in addition to mxc-fixes-env-vars:

1. Seed process env from host (driver.rs)
   ProcessContainer starts with a completely blank environment -- no PATH,
   SystemRoot, or anything.  Seed the process env from the gateway host
   environment so the agent binary can locate DLLs and run.  Skip internal
   Windows drive-letter variables (keys starting with '=') which cause
   CreateProcessW to return ERROR_ENVVAR_NOT_FOUND.  User agent_env entries
   and TLS CA vars are applied as overrides on top of the host env.

2. Remove TLS readonly_paths grant (driver.rs)
   The release wxc-exec (BaseContainer dispatcher) requires write-DAC
   permission on every path in readonly_paths to set up AppContainer ACLs.
   Adding the proxy's temp TLS directory caused a DACL error and exit -1.
   The CA cert paths remain available to the agent via TLS env vars.

3. Remove allowedHosts from network JSON (mxc.rs)
   The release wxc-exec rejects network.allowedHosts / network.blockedHosts
   on Windows with "not yet supported".  Removed the loopback exemption
   attempt (127.0.0.1, ::1, localhost) from the network section.
   Intra-container loopback works natively in the release binary without
   it -- the spawner can connect to the server at 127.0.0.1:22000 directly.

Additional changes:
- mxc-ws-agent.rs: add relay-debug.txt error capture and relay-ready.txt
  marker for reliable timing of host client connections.
- mxc-ws-gateway.toml: debug = true for JSON config dump during diagnosis.
- run-ws-agent-test.ps1: default port changed to 17670 (gateway default);
  relay-ready.txt polling before ws-echo to avoid connecting before the
  spawner has established the proxy bridge.

Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(e2e): address CodeRabbit review on run-mxc-e2e.ps1 (MR !46)

Four robustness/correctness fixes from CodeRabbit:

1. Start-Gw: kill the spawned gateway before the "did not start within 30s"
   throw. If the process is alive but never binds the port, $gw is not yet
   assigned in the caller, so the finally block cannot reap it -> orphan
   gateway holding the port for the next run.

2. create-fail scoring: a non-zero `sandbox create` exit alone is not proof
   of a policy rejection (gateway-registration/transport/fixture errors also
   exit non-zero and would false-pass). PASS now requires a genuine
   rejection signal (network / invalid_argument / network_policies) AND that
   it is not an infrastructure failure; other non-zero exits go to FAIL with
   output captured.

3. deny scenarios (ControlTarget path): snapshot the deny target AFTER
   Wait-File lands the control artifact, so a late denied write (enforcement
   regression racing the control write) can no longer be recorded as PASS.

4. -KeepRunning: break out of the scenario loop after the first scenario so
   a later scenario does not start a second gateway on the same port
   (previously a reliable port collision instead of a usable debug mode).

Re-verified PASS=4 FAIL=0 on both boxes (7F203-MXC-001 base-container and
7F203-MXC-003 AppContainer fallback); network-policy-rejected correctly
scores as "policy rejection".

Signed-off-by: Akber Raza <akberr@nvidia.com>

* feat(mxc-e2e): collect run-mxc-e2e output into a results bundle

Mirror the sibling run-*.ps1 scripts by collecting every run's logs into a
timestamped results-e2e-<stamp>\ folder and zipping it. The bundle contains the
console transcript, per-scenario gateway stdout/stderr, the exact TOML rendered
for each scenario, the policy fixture used, and a summary.txt with the verdict
table.

Per-scenario gateway logs now land in gateway.<scenario>.log/.err.log inside the
bundle instead of a single fixed gateway.e2e.log in the script directory.

Wrap pre-flight, mode setup, scenario definitions, and the scenario loop in a
single try/catch/finally so the finally always writes the summary, stops the
transcript, and zips the bundle -- even on a pre-flight failure. The existing
per-scenario gateway-cleanup try/finally stays nested inside. All scenario
logic, scoring rules, and comments are preserved.

Signed-off-by: Prashant S Khodade <pkhodade@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc-e2e): address CodeRabbit review on run-mxc-e2e.ps1

- Require -Scenario when -KeepRunning: the loop breaks after the first
  scenario, so a full-suite run would execute only one scenario yet still
  report the suite as PASS. Fail fast so a partial run can't be mislabeled
  complete.
- Start-Transcript now runs inside the guarded try block with a
  $transcriptStarted flag; Stop-Transcript is only called when it actually
  started, so a Start-Transcript failure still yields the results bundle.
- Wrap the -Scenario filter in @() so a single exact match stays an array
  (reliable .Count and a proper array for the scenario loop on PS 5.1).

Signed-off-by: Prashant S Khodade <pkhodade@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(examples): pass gateway config via OPENSHELL_GATEWAY_CONFIG for spaced paths

Start-Process -ArgumentList does not quote array elements, so launching the
gateway with a bare --config <path> token split on any space in the install
path (e.g. C:\Users\First Last\...), and clap rejected the fragment with
'unrecognized subcommand'. Every MXC example launcher that started the gateway
hit this when the kit was unzipped under a path containing a space.

Pass the config path through the OPENSHELL_GATEWAY_CONFIG env var (which the
gateway already reads via clap) and drop the --config token. Env vars carry
spaces safely.

Affected: run-ocsf-audit, run-mxc-e2e, run-demo, run-inference-test,
run-ollama-test. run-mtls-test was not affected (its launch passes no config
path). Root-caused and fix-verified on 7F203-MXC-003 from a spaced path.

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(run-mxc-e2e): improve scoring logic and enhance command execution handling

* fix(mxc): reconcile proxy support after rebase

Restore the proxy-enabled OCSF audit example removed by 13185f6e now that the host CONNECT proxy is present. Adapt the proxy lifecycle test to the target branch's DriverSandboxSpec policy delivery contract.

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc): grant sandbox access to proxy CA

- Share the per-sandbox public CA bundle with the AppContainer
- Add real wxc-exec HTTPS proxy coverage and document trust isolation

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc): reject unsupported network middleware

- Reject middleware-bearing MXC policies before sandbox lifecycle begins
- Guard host proxy startup and document the unsupported registry path
- Add mapper, lifecycle, and host proxy regression coverage

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc): reconcile host proxy with main

- remove obsolete inference routing from the host proxy adapter
- use the workspace AWS-LC provider in host-proxy tests
- adapt the forward-proxy test to ProxyIdentityMode

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(build): switch to bundled Z3 for Windows MSVC builds

* fix(mxc): reconcile host proxy after rebase

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Akber Raza <akberr@nvidia.com>
Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
Signed-off-by: Prashant S Khodade <pkhodade@nvidia.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Jamie King <jamiek@nvidia.com>
Co-authored-by: Shailendra Singh <shailendras@nvidia.com>
Co-authored-by: Prashant Khodade <pkhodade@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
2026-09-14 22:58:49 +00:00
Brandon Squizzato cc780d4e17 feat(helm): add BackendTLSPolicy support (#2728)
* feat(helm): add optional BackendTLSPolicy for e2e TLS

Add grpcRoute.backendTLSPolicy values to optionally create a
BackendTLSPolicy resource that enables end-to-end TLS between the
Gateway proxy and the OpenShell gateway pod. The Gateway proxy
terminates client-facing TLS and re-encrypts when connecting to the
backend, validating the pod's certificate against a user-supplied CA
ConfigMap.

This removes the requirement to set server.disableTls=true when using
HTTPS at the Gateway listener. Supported on OpenShift 4.22+ and other
platforms with BackendTLSPolicy support in the Gateway API
implementation.

Update OpenShift and ingress documentation with e2e TLS instructions.

Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>

* feat(helm,server): auto-create backend CA ConfigMap in certgen hook

Extend the generate-certs command with --backend-ca-configmap-name and
--backend-ca-source-secret flags. When BackendTLSPolicy is enabled, the
certgen pre-install hook creates the CA ConfigMap automatically:

- pkiInitJob mode (default): uses the CA from the generated PKI bundle.
  Fully automatic on first install.
- cert-manager mode: reads ca.crt from the server TLS Secret. On first
  install the Secret does not exist yet (cert-manager reconciles after
  templates are applied), so the ConfigMap is created on the first helm
  upgrade. Logs a warning on the initial skip.

The caCertificateConfigMapName value now defaults to <fullname>-backend-ca
when empty, so users only need to set backendTLSPolicy.enabled=true.

Update certgen RBAC to include configmaps get/create when the feature is
enabled. Add CLI arg parsing tests for the new flags.

Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>

* refactor(helm): add server.tls.enableMtls flag for mTLS control

Replace automatic mTLS disabling based on BackendTLSPolicy with an
explicit server.tls.enableMtls flag that defaults to true. The user is
now responsible for setting this to false when using BackendTLSPolicy,
as ingress proxies cannot present client certificates to backends.

Updated:
- values.yaml: Added server.tls.enableMtls (default true)
- gateway-config.yaml: Check enableMtls instead of backendTLSPolicy
- _gateway-workload.tpl: Check enableMtls for client CA mount
- Tests: Updated to use enableMtls flag
- Docs: Added enableMtls=false to BackendTLSPolicy examples
- README: Document new flag and BackendTLSPolicy requirement

Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>

* docs(helm): clarify cert-manager backend CA ConfigMap workflow

Update documentation to explain the two-step install process required
when using cert-manager with BackendTLSPolicy:

1. helm install - cert-manager issues the server certificate, but the
   certgen hook can't create the backend CA ConfigMap yet (cert-manager
   reconciles after templates are applied)
2. helm upgrade - certgen hook reads the CA from the cert-manager-issued
   certificate and creates the ConfigMap

Previously, the docs said "created on first upgrade" without explaining
why or that the feature won't work until then. The updated docs now:
- Explain the timing issue (cert-manager reconciles after chart install)
- Provide clear steps for the cert-manager workflow
- Note that pkiInitJob (default) creates it immediately on install
- Clarify that users must wait for the Certificate to be Ready before
  running the second upgrade

Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>

* fix(docs): remove incorrect external hostname requirement for BackendTLSPolicy

BackendTLSPolicy validates the backend certificate against the service FQDN
(e.g., openshell.openshell.svc.cluster.local), not the external hostname.
The external hostname only needs to be on the Gateway listener certificate
for client-facing TLS.

The default certManager.serverDnsNames already includes all required service
FQDN variants, so no configuration is needed for BackendTLSPolicy to work.

Fixed incorrect documentation that claimed:
- "The server certificate SAN list must include the external hostname"
- Users need to "configure certManager.serverDnsNames with the external hostname"

Removed the unnecessary pkiInitJob.serverDnsNames override from the example
and clarified that:
- Gateway listener certificate needs the external hostname (for clients)
- Backend certificate needs the service FQDN (for Gateway proxy)
- The service FQDN is already in the defaults

Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>

* docs: clarify ACME with LetsEncrypt reference

Change all references from "ACME issuer" to "LetsEncrypt/ACME issuer"
to help users understand that LetsEncrypt is the most common ACME
provider and what ACME means in practice.

Updated:
- docs/kubernetes/managing-certificates.mdx
- docs/kubernetes/openshift.mdx
- deploy/helm/openshell/values.yaml
- deploy/helm/openshell/README.md
- deploy/helm/openshell/ci/values-openshift-route-cert-manager.yaml

Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>

* docs(openshift): restructure end-to-end TLS options and clarify Gateway hostname

Reorganize the OpenShift production deployment documentation:

1. Changed main section from "Production Deployments" to "Options for
   end-to-end TLS" for better clarity

2. Renamed subsections for consistency and clarity:
   - "End-to-end TLS using Gateway API and BackendTLSPolicy (OpenShift 4.22+)"
   - "End-to-end TLS using pass-through Route (all OpenShift versions)"

3. Clarified that the Gateway hostname is typically a wildcard:
   "typically a wildcard like *.openshell-ingress-gw.example.com"

4. Removed the recommendation to copy the cluster's wildcard certificate
   from openshift-ingress namespace, as this is not a recommended
   security best practice

These changes make it clearer that users have two end-to-end TLS options
and help them understand the typical naming pattern for Gateway hostnames.

Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>

* feat(helm): eliminate two-stage install for BackendTLSPolicy with cert-manager

When using BackendTLSPolicy with cert-manager, the certgen hook now polls
for up to 90 seconds waiting for cert-manager to issue the TLS certificate
before creating the backend CA ConfigMap. This eliminates the need for a
second `helm upgrade` in most cases.

The hook polls every 2 seconds with progress logging every 10 seconds.
If cert-manager takes longer than 90 seconds, the hook times out gracefully
and logs a warning, preserving the fallback to manual ConfigMap creation
or a second upgrade.

The Job's activeDeadlineSeconds is 120s, so the 90s timeout leaves 30s
margin for ConfigMap creation and hook completion.

Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com>

* feat(helm): add configurable timeout for certgen hook

Add `pkiInitJob.timeoutSeconds` Helm value (default 120) to control how long
the certgen hook Job can run. When using cert-manager with BackendTLSPolicy,
the hook polls for (timeoutSeconds - 30) seconds to leave margin for ConfigMap
creation and cleanup.

This allows users to increase the timeout for environments where cert-manager
takes longer than 90 seconds to issue certificates, without requiring code
changes.

Example usage:
```yaml
pkiInitJob:
  timeoutSeconds: 180  # Hook polls for 150 seconds
```

Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com>

* docs(helm): document configurable certgen timeout

Update documentation to mention the pkiInitJob.timeoutSeconds value and
how it affects the cert-manager polling behavior when using BackendTLSPolicy.

Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com>

* feat(helm): add configurable failure behavior for certgen timeout

Add `pkiInitJob.failOnTimeout` Helm value (default false) to control whether
the certgen hook fails or succeeds when cert-manager does not issue a
certificate within the polling timeout.

When false (default), the hook succeeds with a warning and users can run
`helm upgrade` after cert-manager issues the certificate to create the
backend CA ConfigMap. This provides backwards-compatible behavior.

When true, the hook fails immediately if the timeout is reached, providing
clear feedback that BackendTLSPolicy is non-functional. This is useful for
strict validation requirements where incomplete installs should fail fast.

Example usage:
```yaml
pkiInitJob:
  timeoutSeconds: 180
  failOnTimeout: true  # Fail install if cert-manager takes >150s
```

Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com>

* feat(helm): change failOnTimeout default to true and add troubleshooting docs

Change `pkiInitJob.failOnTimeout` default from false to true to provide
immediate feedback when cert-manager does not issue certificates within
the polling timeout. This prevents silent failures where BackendTLSPolicy
is non-functional but the install appears to succeed.

Add comprehensive troubleshooting section to docs/kubernetes/ingress.mdx
documenting the specific error "TLS error: Secret is not supplied by SDS"
that occurs when the backend CA ConfigMap is missing, with step-by-step
resolution instructions.

Updated comments in values.yaml to clearly document the default behavior
and explain when administrators might see connectivity errors if they
override the default to failOnTimeout=false.

BREAKING CHANGE: pkiInitJob.failOnTimeout now defaults to true. Helm
installs will fail if cert-manager takes longer than (timeoutSeconds - 30)
seconds to issue certificates. To restore the old behavior of allowing
installs to succeed with a warning, set `pkiInitJob.failOnTimeout=false`.

Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com>

* fix(helm): make cert-manager resources pre-install hooks to fix ordering

Make Certificate and Issuer resources run as pre-install/pre-upgrade hooks
with weight -30, before the certgen hook (weight -20). This fixes the
chicken-and-egg problem where the certgen hook was waiting for Secrets
created by Certificates that hadn't been created yet.

**Hook ordering:**
1. Certificate and Issuer resources created (weight -30)
2. cert-manager issues certificates and creates Secrets
3. certgen hook runs (weight -20), finds Secrets, creates ConfigMap
4. Main resources (StatefulSet, Service, etc.) created

Previously, the certgen pre-install hook would run before any resources
were created, poll for a non-existent Secret, timeout, and fail. The
Certificate resources would never get created because Helm waits for
all pre-install hooks to succeed before creating main resources.

This fix allows single-stage installs to work reliably as long as
cert-manager can issue certificates within the polling timeout.

Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com>

* feat(helm): add validation to prevent enableMtls with BackendTLSPolicy

Add Helm chart validation that fails the install if both
server.tls.enableMtls=true and grpcRoute.backendTLSPolicy.enabled=true
are set, since this is an invalid configuration.

BackendTLSPolicy requires mTLS to be disabled because the Gateway proxy
cannot present client certificates to the backend. This validation provides
immediate, clear feedback at install time rather than allowing the
misconfiguration to be discovered through runtime errors.

Example error message:
```
Error: grpcRoute.backendTLSPolicy requires mTLS to be disabled because
the Gateway proxy cannot present client certificates to the backend;
set server.tls.enableMtls=false
```

Also updated documentation to mention this validation check.

Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com>

* docs(helm): clarify pkiInitJob.timeoutSeconds polling behavior

Improve documentation to clearly explain that pkiInitJob.timeoutSeconds
controls the Job deadline, but the actual polling timeout is
(timeoutSeconds - 30) to reserve 30 seconds for ConfigMap creation
and cleanup.

Added concrete example: "timeoutSeconds=180 allows 150 seconds of polling"
to make the relationship explicit and avoid confusion where users might
expect the hook to poll for the full timeout value.

Updated both values.yaml inline comments and ingress.mdx documentation
for consistency.

Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com>

* docs(openshift): remove outdated two-stage install instructions

Update OpenShift documentation to reflect that single-stage installs now
work with cert-manager and BackendTLSPolicy. The Certificate resources
run as pre-install hooks (weight -30) before certgen (weight -20),
allowing the hook to poll for and find the issued certificates.

Removed the outdated two-step process:
1. helm install (cert-manager issues cert, hook logs warning)
2. helm upgrade (hook creates ConfigMap)

Replaced with current single-stage behavior:
- Certificate resources created as pre-install hooks
- certgen hook polls for up to 90 seconds (configurable)
- Single helm install succeeds in most cases
- Fails fast by default if timeout reached

This brings openshift.mdx in line with the already-updated ingress.mdx
documentation.

Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com>

* fix(helm): make pkiInitJob.timeoutSeconds the actual polling duration

The timeout value now represents the actual polling time that users
experience when waiting for cert-manager to issue certificates.
The Job activeDeadlineSeconds is set to (timeoutSeconds + 30) to
allow buffer time for ConfigMap creation and cleanup.

Previously, the hook polled for (timeoutSeconds - 30) seconds, which
was confusing when users set timeoutSeconds=180 and only got 150
seconds of actual polling.

Updated documentation in values.yaml, ingress.mdx, and openshift.mdx
to reflect the clearer behavior.

Signed-off-by: Brandon Squizzato <bsquizza@nvidia.com>

* docs(helm): update values.yaml and README with correct polling duration

Updated the caCertificateConfigMapName description to reflect that the
hook polls for exactly pkiInitJob.timeoutSeconds seconds, not
(timeoutSeconds - 30) seconds.

Regenerated README.md with helm-docs.

Signed-off-by: Brandon Squizzato <bsquizza@nvidia.com>

* docs(kubernetes): add OIDC configuration to helm install and CLI examples

Updated all helm install and openshell gateway add examples in ingress.mdx
and openshift.mdx to include OIDC issuer and audience configuration.

Examples now use concrete placeholder values:
- OIDC issuer: https://keycloak.example.com/realms/openshell
- OIDC audience: openshell-cli
- Hostname: gateway.example.com
- ClusterIssuer: letsencrypt-prod

This makes it clearer how to configure OIDC authentication, which is
required when using BackendTLSPolicy or HTTPS termination since the
Gateway proxy cannot present client certificates to the backend.

Signed-off-by: Brandon Squizzato <bsquizza@nvidia.com>

* docs(kubernetes): explicitly list OIDC client ID in gateway add examples

Added --oidc-client-id openshell-cli to all openshell gateway add
commands in ingress.mdx and openshift.mdx, making the default client
ID explicit in the examples even though it's the CLI default.

This improves clarity and helps users understand the complete OIDC
configuration needed for gateway registration.

Signed-off-by: Brandon Squizzato <bsquizza@nvidia.com>

* fix(helm): address PR review feedback for BackendTLSPolicy

- Read backend CA from the authoritative server Secret instead of the
  in-memory PKI bundle so enabling BackendTLSPolicy on an existing
  release uses the CA that actually signed the server certificate.
- Reconcile the backend CA ConfigMap on every hook run (compare and
  update) instead of skipping when it already exists, so CA rotations
  propagate automatically.
- Remove hook annotations from cert-manager Issuer/Certificate resources
  so they remain regular release objects managed by Helm lifecycle. Split
  the cert-manager backend CA ConfigMap creation into a separate
  post-install/post-upgrade hook Job that polls after cert-manager
  Certificate resources are applied.
- Update architecture/gateway.md, docs/reference/gateway-config.mdx,
  debug-openshell-cluster skill, and helm-dev-environment skill with
  BackendTLSPolicy, backend CA ConfigMap, enableMtls, and timeout
  documentation.

Signed-off-by: Brandon Squizzato <bsquizza@nvidia.com>
Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>

* fix(docs): resolve markdown lint errors in helm README and kubernetes docs

Escape inline HTML angle brackets in README.md template placeholders,
remove trailing spaces, and add blank lines around fenced code blocks
in numbered lists.

Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>

* Update docs

* fix(helm): escape inline HTML in values.yaml descriptions and sync mise lockfile

Wrap `<fullname>` and `<namespace>` template placeholders in backticks
so markdownlint does not flag them as inline HTML (MD033). Regenerate
mise.lock to match current mise.toml after rebase onto main.

Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>

* Run 'mise lock'

* fix: align mise.lock with CI mise version output

The lockfile was regenerated locally with mise 2026.8.10 which resolves
uv Linux artifacts to gnu variants and adds provenance_verified fields,
but CI uses v2026.4.25 which produces musl variants without those fields.

Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>

* fix(docs): correct cert-manager hook ordering and clientCaSecretName comment

Update ingress.mdx and openshift.mdx to describe Certificate resources
as regular release objects with a post-install/post-upgrade Job, matching
the current implementation and architecture/gateway.md.

Fix values.yaml clientCaSecretName comment to state that "" disables
client certificate verification, matching the helper and access-control
docs.

Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>

---------

Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>
Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com>
Signed-off-by: Brandon Squizzato <bsquizza@nvidia.com>
2026-09-14 17:39:15 +00:00
Jesse JaggarsandDrew Newberry 02b664bb0d refactor(config): normalize and enforce gateway schema v2 (#2814)
* refactor(config): normalize compute driver field names

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* refactor(config): introduce canonical gateway fields

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* refactor(config): enforce gateway schema version 2

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): preserve compute driver runtime guarantees

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): address schema v2 review regressions

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): complete schema v2 migration safeguards

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(config): expand schema v2 regression coverage

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(config): add schema v2 parity manifest

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): correct parity manifest inventory

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* docs(config): record schema v2 intentional changes

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* docs(config): disposition schema v2 parity gaps

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): add dual schema parity harness

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): establish compute lifecycle parity baseline

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): preserve gateway option compatibility

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): record gateway option parity

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* docs(config): close gateway-wide parity gaps

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(podman): apply configured pids limit

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): validate Podman option parity

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): add Kubernetes option parity harness

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): record Kubernetes option parity

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): disposition VM parity lanes

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): add external driver parity lane

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(e2e): preserve external driver pull policy

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): attest parity artifacts and launches

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): require clean parity build sources

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): bind parity runtime artifacts

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(e2e): use isolated supervisor tags

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(e2e): qualify parity image tags

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(e2e): serve parity supervisor locally

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): isolate parity podman services

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): harden parity evidence provenance

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): pin parity sandbox artifacts

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): attest parity runtime inputs

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): bind parity runtime evidence

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): record compute boundary parity

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): disposition cross-cutting parity lanes

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(packaging): preflight gateway config upgrades

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): preserve rebase integration guarantees

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(ci): isolate temporary git signing config

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): update remaining schema v2 consumers

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(ci): provide e2fs tools to VM tests

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): align preflight with gateway startup

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(vm): preserve rootfs tar configuration

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* chore(config): adopt duration unit constructors

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(packaging): preflight RPM gateway config

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): address driver review findings

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(e2e): require fresh semantic parity evidence

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(docker): update tests for renamed sandbox label

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(gateway): preserve selective driver coverage after rebase

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
2026-09-11 05:00:24 +00:00
ae57979b03 feat(mxc): add Windows ETW-to-OCSF audit trail (#3015)
* feat(mxc): ETW->OCSF audit consumer + Windows OCSF JSONL parity (cp6 P1)

Add a Windows MXC ETW->OCSF audit trail in openshell-driver-mxc: a real-time
Sandboxing-provider ETW consumer that decodes events (TDH), attributes each to
an OpenShell sandbox_id, and maps them to OCSF (lifecycle 6002, config 5019,
process 1007, finding 2004).

cp6 Phase 1 - durable OCSF JSONL audit-file parity with Linux:
- openshell-ocsf: add emit_ocsf_event_routed (populates the event-bridge
  thread-local AND stamps sandbox_id+message in one dispatch) plus public
  set/clear_current_event; OS-aware device (Device::windows/for_current_os) so
  device.os.name reflects the host instead of a hardcoded Linux stub.
- etw_consumer: emit via the routed emit (previously fired a bare info! that
  never populated the bridge, so the structured event was dropped).
- openshell-server: install OcsfJsonlLayer over a synchronous daily-rotated
  appender (durable under force-kill), gated by OPENSHELL_OCSF_JSON, path via
  %PROGRAMDATA%\OpenShell\logs (override OPENSHELL_OCSF_LOG_DIR).
- device.hostname now resolves to the real gateway machine name.

Box-proven on 7F203-MXC-001: JSONL lines == shorthand OCSF rows, all valid
OCSF JSON, per-sandbox attribution intact, disabled state writes nothing.

Signed-off-by: Akber Raza <akberr@nvidia.com>

* feat(mxc): map remaining Sandboxing ETW events to OCSF

Close the last three ETW->OCSF gaps so the audit trail covers the full
set of events the Sandboxing provider emits (12/12):

- ProcessLaunched -> Process Activity [1007] "Launch" (confirmed start;
  carries the real processId/threadId, the twin of CreateProcessInSandbox
  which only has the request + command line).
- SandboxProxyConfigured -> Device Config State Change [5019] (the one
  network-plane setup event; surfaces proxyPort, "no proxy" when 0).
- SandboxConsoleReferencePlumbed -> Device Config State Change [5019]
  (console-handle plumbing).

map_config_state now handles the full config/hardening/setup family and
carries proxyPort/hasConsoleReference/creationFlags as unmapped fields.
Verified on 7F203-MXC-001: 11/12 event types emit OCSF without a proxy
(SandboxProxyConfigured requires proxy config to fire).

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc): seed ETW attribution under registry lock + Device tests

Address CodeRabbit review on !31:

- Prevent stale ETW attribution on a delete/launch race: register the
  wxc-exec pid while holding the registry lock, and bail if the sandbox
  entry is already gone. Previously the attribution key could be seeded
  after `delete` had removed the sandbox, leaving a stale key that could
  misroute later Sandboxing ETW events to a dead sandbox_id. Lock order
  (registry -> attribution) matches the delete path, so no deadlock.
- Add unit tests for the new Device::windows and Device::for_current_os
  constructors to harden Windows/Linux OCSF device parity.

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc-etw): buffer+replay racing events and harden attribution keys

Addresses two ETW->OCSF attribution review items (Shailendra #1, #2).

#2 early-event loss: ETW delivers the sandbox create/config burst the instant wxc-exec starts, which can beat the driver's register_launch (now under the registry lock post-Ready). process_event previously dropped anything unresolved, losing the racing burst. Add a bounded, time-bounded pending buffer (PENDING_MAX=4096, PENDING_TTL=5s): unresolved events are held and replayed once attribution lands, aged-out ones dropped. Consumer switched to a timed recv_timeout(200ms) so the buffer is re-driven after each event and on a tick. Emit path factored into shared emit_resolved().

#1 attribution collisions: a Windows PID is recycled after exit and a command line is commonly identical across sandboxes. register_launch now rebinds by_pid on reuse and clears the stale last_pid_sid hint (warns if the PID still pointed at a different, leaked sandbox); command line is held in by_cmd only while unique and demoted to a new ambiguous_cmds set on a second owner, so a duplicate command refuses to resolve rather than misroute.

Unit tests: buffer replay (direct + cross-link), buffer bound, PID-reuse rebind, duplicate-cmd non-resolution. Box-verified on 7F203-MXC-001 (5 sandboxes, identical cmd -> 5 isolated sandbox_ids, 50/50 OCSF/JSONL, BuffersLost=0).

Signed-off-by: Akber Raza <akberr@nvidia.com>

* docs(mxc-etw): note cmd_line is captured raw with no privacy filtering

Review item #3 (Shailendra): add a PRIVACY NOTE on map_process_launch stating cmd_line is copied verbatim into OCSF process.cmd_line with no redaction, so secrets/PII on a command line land unredacted in the durable audit trail (deliberate audit-fidelity trade-off; treat the log as sensitive). Redaction is owned by an upstream privacy layer, not this path; no general audit-output PII scrubber exists today (openshell_core::secrets [CREDENTIAL] redaction is scoped to the proxy HTTP-target logging, a separate egress path).

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc-etw): open ETW trace on caller thread so start_session reports real status

Review item #4 (Shailendra): start_session previously returned Ok(EtwSession) as soon as the pump thread was spawned, but OpenTraceW ran later inside that thread; if it failed we still handed back a live-looking session and logged 'consumer started' (silent failure = false audit coverage).

Split the two Win32 calls instead of adding a channel handshake (avoids any lost-wakeup/hang risk): the quick, synchronous OpenTraceW now runs on the caller thread (open_trace), and only the blocking ProcessTrace runs on the pump thread (run_trace). start_session returns Err if OpenTraceW fails (reclaiming the boxed Sender so the consumer disconnects, stopping the session, joining the consumer) and returns Ok/logs 'started' only once capture is genuinely open. Opened handle + LoggerName buffer + boxed Sender are carried to the pump via a Send OpenedTrace so they outlive ProcessTrace.

Box-verified on 7F203-MXC-001: consumer started=True, failed-to-start=False, 50 OCSF rows / 50 JSONL, BuffersLost=0 (no regression to capture/emit).

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc-etw): guard pending-event replay against PID recycling

CodeRabbit flagged that drain_resolved() re-resolved buffered events
against the live by_pid map, so if Windows recycled a wxc-exec PID within
PENDING_TTL a stale event from the dead sandbox could be emitted under the
new owner.

Stamp each by_pid registration with its Instant and add resolve_replay(),
used only on the buffered/replay path. It (a) never falls back to the
recycle-/ambiguity-prone by_cmd or last_pid_sid keys, and (b) trusts a PID
match only when the registration is not newer than the buffered event by
more than REPLAY_PID_GRACE (2s) - a recycled PID's registration lands well
outside that window, so the stale event ages out instead of misattributing.
The legitimate #2 seed race (registration lands ~immediately) still replays.

Adds unit tests for the recycle-refusal, in-grace acceptance, and
weak-fallback exclusion.

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc-etw): surface unexpected ProcessTrace termination (review #4)

start_session already returns Err on OpenTraceW failure (runs on the
caller thread since e41a7701), closing the first half of Shailendra's #4.
This closes the second half: ProcessTrace's result was discarded, so if
capture died mid-run the backend had no way to know.

Add a shared CaptureHealth (stopped/stopping/exit_code) between the pump
thread and EtwSession. run_trace now records ProcessTrace's WIN32_ERROR
and, when the pump returns without a deliberate stop, logs at ERROR that
MXC OCSF capture is no longer running. EtwSession::stop() sets `stopping`
before teardown so a normal shutdown isn't misreported, and
EtwSession::is_capture_alive() exposes the state for status/diagnostics.

Box-verified on 7F203-MXC-001: 5 sandboxes, 50 attributed OCSF rows,
JSONL parity 50/50, BuffersLost=0, clean start/stop (no false failure).

Signed-off-by: Akber Raza <akberr@nvidia.com>

* feat(mxc-ocsf): add ETW->OCSF audit-trail example kit; fix proxy-configured message

Add a runnable OCSF audit-trail example under examples/ (run-ocsf-audit.ps1,
mxc-ocsf-audit.toml, ocsf-audit.yaml, README) that spins up sandboxes with the
in-process ETW consumer and egress proxy on, emitting a full OCSF JSONL audit
trail across all four classes (6002/5019/1007/2004).

Fix SandboxProxyConfigured mapping to log "MXC sandbox proxy configured" instead
of a misleading "(no proxy)" when the provider reports proxyPort=0; the event's
presence already indicates proxy configuration. Verified on-box: 26 events, all
mapped ETW event types present.

Signed-off-by: Akber Raza <akberr@nvidia.com>

* feat(mxc-ocsf): clearer audit report + client-safe run-ocsf-audit.ps1

Improve the ETW to OCSF audit-trail example output and make it safe to ship.

Report:
- Add an event-type coverage count ("N of M expected event types fired");
  the denominator auto-adjusts (8 with proxy on, 7 with -NoProxy).
- Split the checklist into expected event types vs anomaly findings
  (ActivityError/FallbackError), which are reported separately and not
  counted toward coverage (a clean run may emit none).
- Verdict is now coverage-based (all expected types must fire) instead of
  the looser "at least 3 OCSF classes".
- Call out the absolute path to the durable OCSF JSONL log prominently.

Client-safety:
- Default -ShareOut to empty (no auto-copy); pass -ShareOut a UNC path to
  opt in. Removes a hardcoded internal share path from a published example.
- Drop internal-team wording ("Hand that zip back for evaluation", "BUNDLE:")
  in favor of neutral "Results bundle:".
- Update README-ocsf-audit.txt to match the opt-in -ShareOut behavior.

Verified on both MXC boxes: 7F203-MXC-001 (base-container) -> PASS, 8 of 8
event types, 26 OCSF events across 4 classes; 7F203-MXC-003 (AppContainer
fallback) -> reduced set as expected, clean output.

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc): configure OCSF audit workloads per sandbox

- remove unsupported gateway-scoped workload fields from the shipped MXC audit example.
- build the command, working directory, and filesystem grant from each run's ShareDir
- pass the workload through --driver-config-json.
- preserve the host CONNECT proxy configuration and conditional audit coverage for the future host_connect_proxy merge
- require the workload output when determining the audit verdict.

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc): omit command arguments from OCSF audit logs

- record only the executable basename for MXC CreateProcessInSandbox audit events
- leave process.cmd_line unset so workload arguments cannot reach shorthand or JSONL logs
- cover tokens, passwords, signed URLs, and PII with a secret-leak regression test
- update the audit example, architecture guidance, and published logging documentation
- preserve ETW attribution and future host_connect_proxy enforcement behavior

Signed-off-by: Akber Raza <akberr@nvidia.com>

* feat(etw): enhance ETW session management with distinct naming for concurrent gateways

* fix(etw): bound the audit queue during overload

- replace the unbounded ETW callback channel with count- and byte-bounded buffering

- keep the ETW callback non-blocking and count records rejected during overload

- emit immediate, rate-limited warnings that identify resulting audit coverage gaps

- make the audit example fail when queue overload causes dropped ETW records

- cover stalled consumers, oversized events, and warning throttling with unit tests

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(etw): harden sandbox audit attribution

- remove command-line and persistent per-PID fallback keys from live and replay resolution

- retire the driver-owned wxc-exec PID before publishing child completion

- retain established identity, activity, and correlation-vector links only for the five-second late-event window

- prevent buffered records from crossing rapid PID retirement and reuse boundaries

- add resolver and lifecycle coverage and document the attribution trust boundary

Signed-off-by: Akber Raza <akberr@nvidia.com>

* chore(mxc): address rebase follow-ups

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(mxc): align OCSF audit example with driver config

- remove unsupported egress proxy settings

- stop requiring the unavailable proxy audit event

- update example documentation for supported event coverage

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(etw): redact command-line secrets in DecodedEtwEvent summary

* fix(etw): enhance PID resolution and event attribution logic for ETW records

* fix(ocsf): restrict gateway-local JSONL sink to Windows/MXC path with opt-in configuration

* address rebase issues

* fix(mxc): address ETW audit review feedback

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(mxc): fail closed across ambiguous PID reuse

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(mxc): bind ETW attribution to process generation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Akber Raza <akberr@nvidia.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Jamie King <jamiek@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
2026-09-11 02:04:26 +00:00
Drew Newberry 38f2aef930 feat(gateway): support selective compute driver builds (#3118)
* feat(gateway): support selective compute driver builds

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(gateway): support selective Windows MXC builds

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-11 00:36:09 +00:00
John T. MyersandJohn Myers 226a83b323 fix(supervisor): classify credential placeholders in request bodies (#3246)
* fix(supervisor): classify credential placeholders in request bodies

Closes #2904

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* fix(supervisor): preserve same-provider placeholders in request bodies

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <johntmyers@users.noreply.github.com>
2026-09-10 23:10:01 +00:00
Shiju ea8eda6d5b feat(supervisor): enforce MCP request protocol versions (#3241)
* feat(supervisor): enforce MCP request protocol versions

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(supervisor): enforce MCP versions across HTTP forwarding

Apply shared request-version guards before authorization and after forward-request rewriting. Require version metadata to survive HTTP header cleanup, and cover valid initialization and selected-revision forwarding through middleware.

Signed-off-by: Shiju <shiju@nvidia.com>

---------

Signed-off-by: Shiju <shiju@nvidia.com>
2026-09-09 20:28:55 +00:00
John T. Myers f4dc6be4b2 refactor(inference): remove managed inference routes (#3195)
* refactor(inference): remove managed inference routes

Closes #3172

Remove the inference route control plane, inference.local data path, built-in router crate, and SDK surface. Move inference workloads to explicitly imported provider profiles and native endpoints, with migration cleanup and updated tests and documentation.

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(policy): preserve alternate upstream isolation

Restore the provider policy activation guard so legacy OpenAI and Anthropic providers configured for alternate base URLs do not grant egress to the built-in public vendor endpoints.

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-09-09 18:47:22 +00:00
Artem Lytvyn 7f4bd49a47 fix(policy): harden landlock.compatibility validation (#2541)
* fix(policy): reject invalid landlock.compatibility values at parse time

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* fix(policy): abort sandbox startup when hard_requirement has no filesystem paths

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* fix(policy): validate landlock.compatibility at gateway and fix zero-path logging

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* fix(policy): reject invalid landlock.compatibility on serialization

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

* fix: fixed linting error

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>

---------

Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>
2026-09-09 18:22:55 +00:00
Gaizka Menendez 8af79a7f4b fix(podman): resolve macOS Podman socket dynamically (#3135)
* docs(podman): document macOS socket path mismatch and dynamic lookup

On macOS, Homebrew-installed Podman does not create the default socket
path that the Podman driver probes. Document the OPENSHELL_PODMAN_SOCKET
override and the podman machine inspect lookup in both the compute
drivers reference and the debug-openshell-cluster skill.

Fixes #1690

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* fix(podman): resolve macOS Podman socket dynamically

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* chore: restore debug skill file

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* chore: drop legacy debug skill path

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* fix(podman): trim unrelated e2e changes

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* fix(e2e): harden shell array expansion

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* chore: remove unrelated skill note

* ci: retrigger checks

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

---------

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>
2026-09-09 14:10:59 +00:00
Evie Howard 6e6b3c8905 refactor(cli): remove local Dockerfile image builds (#3214)
Signed-off-by: Evie Howard <evhoward@redhat.com>
2026-09-09 13:28:17 +00:00
Philippe Martin 118b250f01 feat(sandbox): support rootfs tar as --from source for VM driver (#2863)
* feat(sandbox): support rootfs tar as --from source for VM driver

Accept flat rootfs tar archives (.tar, .tar.gz, .tgz) via the --from
flag for VM-backed gateways. The CLI detects the archive extension,
validates that the gateway uses the VM compute driver, and passes the
tar path through driver_config. The VM driver copies the tar into its
staging area and feeds it into the existing rootfs extraction and ext4
disk creation pipeline, skipping the container image pull/export steps.

Closes #2175

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(sandbox): validate rootfs tar path at the VM driver boundary

The rootfs_tar_path field in driver_config was passed from the API
caller directly to tokio::fs::copy without validation. An authenticated
user bypassing the CLI could supply arbitrary host paths (e.g.
/dev/zero for disk exhaustion, or readable host files for data
exfiltration).

Introduce a trusted staging directory that the VM driver creates on
startup and advertises via GetCapabilities. The CLI now copies the tar
into the staging directory before creating the sandbox, and the driver
validates that the received path is a regular file inside the staging
root and within a configurable size limit (default 10 GiB) before any
I/O.

New VmDriverConfig options:
- rootfs_tar_staging_dir: override the staging directory
  (default: <state_dir>/rootfs-tar-staging)
- rootfs_tar_max_bytes: override the size limit (default: 10 GiB)

Addresses GATOR-28b5152e-01.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(sandbox): request-scoped staging, size pre-check, and cleanup for rootfs tar

Tighten the rootfs tar staging flow to address the remaining GATOR-01
obligations:

- Request-scoped staging: the CLI creates a unique per-request
  subdirectory (req-<pid>) under the staging root instead of placing
  files directly in the shared directory. The driver enforces that the
  tar path is at depth 2 (staging_root/<subdir>/<file>), preventing
  cross-request path selection.

- Size pre-check: the driver advertises rootfs_tar_max_bytes via
  GetCapabilities. The CLI reads this limit and rejects oversized files
  before copying, avoiding disk exhaustion in the staging directory.

- Cleanup: the driver removes the request staging subdirectory after
  consuming the tar (on cache hit, copy success, or copy failure),
  ensuring staged data does not persist beyond the request.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(vm): restore rootfs-tar sandboxes from persisted image identity

On restore or restart, the one-shot staged tar archive has already been
cleaned up. Reading the persisted image identity from the sandbox state
directory and resolving the cached disk path directly avoids re-accessing
the deleted staging path.

Addresses GATOR-168b9210-01.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(cli): use random staging dirs and enforce byte limit during rootfs tar copy

Replace PID-based request staging directories with tempfile-generated
random names to prevent collisions and make paths unpredictable.
Replace bare tokio::fs::copy with a streaming copy loop that enforces
the advertised max_bytes limit during transfer, closing the TOCTOU gap
between the pre-copy size check and the actual copy.

Signed-off-by: Philippe Martin <phmartin@nvidia.com>
Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(sandbox): issue rootfs tar staging slots from the gateway

A caller could name any host path in `driver_config.vm.rootfs_tar_path`,
which the privileged VM driver then read. The CLI-side locality check did
not apply to direct API requests.

The gateway now owns staging. `BeginRootfsTarStaging` allocates a
request-scoped directory under the driver-advertised staging root and
returns an opaque single-use token; `CreateSandbox` carries the token, and
the gateway substitutes the path it allocated before dispatching to the
driver. `template.driver_config.<driver>.rootfs_tar_path` is rejected
outright in request validation, so a caller-supplied path never reaches
privileged I/O.

Tokens are bound to the issuing workspace and subject, consumed once, and
expire after 30 minutes. Outstanding slots are capped per caller and
overall, so one caller can neither exhaust the staging filesystem nor
starve others. An RAII guard reclaims the directory on every failure path
after consumption, and an age-gated sweep runs at startup and on each
reconcile pass for directories whose driver died before its own cleanup.

The token is stripped from the public sandbox before persistence: the
stored copy is returned verbatim by GetSandbox, ListSandboxes and
WatchSandbox to every member of the workspace.

Also fixes two defects this exposed:

- The CLI wrote `rootfs_tar_path` at the top level of `driver_config`, but
  the gateway forwards only `driver_config.<driver_name>`, silently
  dropping unmatched keys. The archive never reached the VM driver, so the
  documented `--from ./rootfs.tar` flow did not work at all. Config is now
  nested under `vm` and deep-merged, so a caller's existing VM settings
  survive instead of being clobbered by a shallow extend.

- Staging previously required `GetGatewayInfo`, which is restricted to
  `platform_admin`, making the feature unusable for ordinary users on any
  RBAC-enabled gateway. The new RPC matches CreateSandbox at
  `sandbox:write` / `workspace_role: user`.

`compute_driver.proto` is unchanged; the gateway reads the staging root
from the capabilities it already stores.

Refs #2175

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(vm): derive rootfs tar cache identity from archive contents

The prepared-disk cache key combined the archive's full path with an mtime
truncated to seconds, then mapped punctuation to `-`. Distinct paths such as
`/tmp/a/b.tar` and `/tmp/a-b.tar` collapsed onto the same key and reused each
other's disk, a rewrite within the same second kept stale contents, and a long
path could exceed filesystem component limits.

Identity is now a SHA-256 of the archive contents. This is also what makes the
cache work at all now that the gateway allocates a fresh staging directory per
request: a path-derived key would miss on every create.

The archive is hashed, the cache checked, and only on a miss copied — so a hit
skips writing a multi-gigabyte file. The copy is hashed as it is written and
rejected if the digest differs from the first pass, which closes the window
where the source changes during staging rather than approximating it with a
re-stat.

Refs #2175

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(vm): decompress gzip rootfs tar archives during staging

`--from` accepts `.tar.gz` and `.tgz`, but the driver staged whatever bytes
it was given and the guest image-prep VM extracts the staged file with a
plain `tar -xpf`. Compressed sources therefore depended on the guest tar
auto-detecting gzip, and the prepared disk was sized from the compressed
length, which is far too small for the expanded rootfs.

Staging now detects gzip from the archive's magic bytes -- the driver only
ever sees a gateway-issued staging path, never the caller's file name -- and
writes an uncompressed tar. The digest still covers the source bytes, so the
"archive changed while staging" check is unaffected, and expansion is bounded
by `rootfs_tar_max_bytes` so a compression bomb cannot fill the host disk.

`extract_rootfs_archive_to` sniffs gzip as well, so the host-side extraction
path matches.

Adds unit coverage for gzip staging, bounded expansion, and gzip extraction,
plus an e2e sandbox created from a gzip-compressed export.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

---------

Signed-off-by: Philippe Martin <phmartin@redhat.com>
Signed-off-by: Philippe Martin <phmartin@nvidia.com>
2026-09-08 22:36:40 +00:00
Jesse Jaggars 2ad86c1b2e fix(cli): fail closed when OIDC refresh fails (#2817)
* fix(cli): fail closed when OIDC refresh fails

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(cli): fall back to unexpired OIDC token on refresh failure

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

---------

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>
2026-09-08 15:37:26 +00:00
Evan Lezar 1510e2c5a7 feat(compute): advertise resource capabilities (#3010)
* refactor(driver-docker): group runtime GPU capabilities

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* feat(compute): advertise resource capabilities

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* docs(compute): describe resource capabilities

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(server): satisfy resource capability clippy lint

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-09-07 13:32:42 +00:00
Shiju 592df3e014 feat(policy): preserve exact MCP revision allowlists (#3027)
* feat(mcp): add version-aware wire profile metadata

Signed-off-by: Shiju <shiju@nvidia.com>

* feat(policy): canonicalize MCP version allowlists

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(policy): align MCP policy tests with current main

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(policy): canonicalize supervisor protobuf ingress

Materialize defaultable MCP revisions before ambiguity checks, OPA construction, and sidecar delivery. Reject invalid sidecar policies with bounded errors.

Signed-off-by: Shiju <shiju@nvidia.com>

---------

Signed-off-by: Shiju <shiju@nvidia.com>
2026-09-05 04:24:49 +00:00
John T. MyersandJohn Myers fc0929749c fix(policy): harden advisor transport proposals (#3136)
Signed-off-by: John Myers <johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <johntmyers@users.noreply.github.com>
2026-09-04 19:38:00 +00:00
Yuedong Wu c93b2fa7da docs(gateway-config): fix stale community sandbox image path (#2800)
* docs(gateway-config): fix stale community sandbox image path

Signed-off-by: Yuedong Wu <dwcn22@outlook.com>

* docs(sandbox-image): purge remaining stale image references

Rebasing onto main surfaced four more instances of the same dead
ghcr.io/nvidia/openshell/sandbox path, introduced by commits merged
after this branch was opened: three test fixtures (driver-docker,
openshell-ocsf, compute::mod) and one user-facing default in the
SPIFFE token-exchange Podman demo README. Correct all four to
ghcr.io/nvidia/openshell-community/sandboxes/base, consistent with
the rest of this fix.

Signed-off-by: Yuedong Wu <dwcn22@outlook.com>

---------

Signed-off-by: Yuedong Wu <dwcn22@outlook.com>
2026-09-04 00:16:42 +00:00
Yuedong Wu 7cc9551677 feat(server): support EC and EdDSA keys in OIDC JWKS validation (#2593)
Signed-off-by: Yuedong Wu <dwcn22@outlook.com>
2026-09-04 00:07:29 +00:00
krishicks 17171cd933 refactor(otel): unify compute driver tracing (#2995)
Centralize compute-driver RPC descriptors, stream instrumentation, provider
routing, and standalone installation in openshell-otel. Use typed RPC
constants so gateway and in-process driver paths cannot panic on unknown
operation strings or repeat runtime method parsing.

Emit semantic-convention rpc.service and rpc.method attributes, preserve
trace context and resource identity across deployment modes, and route both
RPC boundary and backend crate spans to each selected driver provider. Leave
consumer-dropped watch spans unset while recording observed terminal status,
and avoid reboxing untraced external-driver streams.

Derive each driver tracing identity from Cargo package and crate metadata and
attach its descriptor to the compute-driver registration, keeping provider
selection and target routing tied to the registered implementation. Share
tracing setup and round-trip test support across Docker, Podman, Kubernetes,
and VM, and update the gateway tracing documentation.

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-09-03 15:26:14 +00:00
Philippe Martin 857af42a16 feat(vm): support corporate HTTP forward proxy egress for microVM sandboxes (#3090)
* feat(vm): support corporate HTTP forward proxy egress for microVM sandboxes

The corporate forward proxy machinery from #1792 is driver-agnostic and
already merged: openshell-supervisor-network implements CONNECT chaining,
NO_PROXY matching, credentials, https:// proxies and corporate CA trust, and
openshell-sandbox exposes it as six argv-only flags. Podman gained the driver
half in #2245/#2512 and Kubernetes in #2633; the VM driver had none of it, so
VM sandboxes on proxy-only networks could not reach any destination requiring
the proxy even when policy allowed it.

The blocking piece was not proxy logic but delivery: the VM guest init script
runs as PID 1 and execs a fixed supervisor command line, and libkrun's
krun_set_exec receives an empty argv, so there was no channel for driver-owned
supervisor arguments. The supervisor's proxy flags deliberately have no
environment fallback, and build_guest_environment merges user-supplied
environment, so the guest env is not a safe transport either.

Add a driver-authored argument file, mirroring the existing init.d manifest:
the driver writes /opt/openshell/supervisor-args into the overlay upperdir on
every launch and the guest reads it verbatim, one argument per line, appending
it to every supervisor exec. It is written even when empty, which is what makes
the channel unforgeable -- the upperdir always shadows the read-only image
layer, so an image can neither supply its own arguments nor disable the
operator's by omitting the file. Because both launch backends exec the same
init script, this covers libkrun and QEMU without touching either.

A microVM has no bind mounts or container secrets, so the credential and CA
bundle are staged into the per-sandbox overlay the way the gateway JWT already
is: credential root-only at 0600, CA at 0644, both rewritten every launch so a
removed setting clears prior material, and both deleted with the sandbox state
directory. This places the credential at rest in the overlay image on the
gateway host, which differs from the Podman secret model and is documented as
an explicit security consideration.

Validation is fail-closed and shared: a new
openshell_core::driver_utils::validate_upstream_proxy_settings holds the
pairing rules the Podman driver established, and both the gateway and the
driver call it so an invalid table names the offending key instead of
surfacing as an opaque driver-readiness timeout.

Guest egress leaves through gvproxy, so a proxy on the gateway host's loopback
is reachable only through host.openshell.internal; the guest to gateway
callback is unaffected.

Closes #3088

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(vm): bound the proxy CA read and scope the host-loopback recipe

Two review findings on the corporate forward proxy support for microVM
sandboxes.

The driver read the operator's proxy_ca_bundle with an unbounded fs::read and
accepted it on a substring match for the PEM BEGIN CERTIFICATE marker. A
special file such as /dev/zero therefore grew driver memory without bound on
every authorized sandbox create, and a PEM block holding invalid DER passed
the host check but contributes no trust anchor in the guest, so every
supervisor would fail after boot with an error attributed to the sandbox
rather than to the setting.

Move the read into openshell-core as read_upstream_proxy_ca_bundle_file: it
reuses the credential reader's bounded-read path (non-regular files rejected
on fstat, size capped, read bounded even if the file grows), then requires at
least one anchor that RootCertStore::add_parsable_certificates accepts. The
supervisor's own reader now delegates to it, so host acceptance and guest
acceptance are the same function and cannot drift.

The published host-loopback recipe was written for libkrun only. gvproxy NATs
host.openshell.internal to the gateway host's 127.0.0.1, but GPU sandboxes run
on the QEMU/TAP backend where that name resolves to the TAP host address and
the driver's own nftables input chain accepts only the gateway port from the
guest — no proxy on the gateway host is reachable there at any bind address,
so an operator following the generic recipe lost all proxy-required egress
while configuration validation succeeded.

Scope the recipe to libkrun in every reference and reject a gateway-host proxy
URL when a launch plan resolves to QEMU, naming the reason, instead of booting
a sandbox whose policy-approved CONNECTs all time out.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(vm): match the QEMU proxy preflight to the selected TAP host

The gateway-host proxy guard added for the QEMU/TAP backend classified the
wrong set of addresses in both directions.

It ran at the top of configure_qemu_launch_plan, before the subnet allocation
that settles plan.host_ip, so it could not compare against the address the
guest actually reaches the host on. An operator pointing https_proxy at the
sandbox's own TAP host address, such as 10.0.128.1, passed the check, and the
driver's nftables input chain — which accepts only the gateway port from the
guest — then dropped every policy-approved CONNECT, which is exactly the
silent timeout the guard exists to prevent.

In the other direction it rejected 192.168.127.254 unconditionally. That
address is special only to libkrun/gvproxy; on QEMU/TAP it is an ordinary
address that may be routable through the guest's masqueraded egress, so the
guard refused a working configuration.

Run the check after the launch plan's network allocation, on both the
freshly-allocated and already-complete paths, and compare IP literals with
that sandbox's selected TAP host. Loopback literals, localhost, and the
documented host aliases that write_host_gateway_aliases seeds to the TAP host
still classify as the gateway host, and the failure names the address. The
gvproxy host-loopback constant returns to being a documentation anchor.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

---------

Signed-off-by: Philippe Martin <phmartin@redhat.com>
2026-09-02 14:56:08 +00:00
Drew Newberry 9ca19e6c80 refactor(compute): decouple gateway driver composition (#2823)
* refactor(compute): decouple gateway driver composition

Move first-party composition and VM process ownership into openshell-gateway, leaving openshell-server backend-independent. Update packaging and build references with the new crate, simplify the compiled-driver boundary, and keep the driver-free gateway path buildable with bundled Z3 tooling.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(telemetry): bound compute driver categories

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(core): keep runtime transport generic

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(compute): complete server driver decoupling

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(compute): preserve driver integrations after rebase

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>

* fix(compute): preserve docker tracing after decoupling

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>

* fix(compute): preserve driver behavior after extraction

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(compute): remove MXC policy side channel

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(compute): separate policy delivery from readiness

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>
2026-09-01 21:13:45 +00:00
John T. MyersandJohn Myers 07df822090 feat(providers): make profiles authoritative (#2962)
* feat(providers): make profiles authoritative

Closes #1988

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* docs(providers): move profiles into provider navigation

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* docs(providers): clarify provider attachment lifecycle

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* fix(tui): scroll provider profile picker

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* fix(providers): honor profile credential semantics

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* fix(providers): prefer exact profile IDs

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* fix(providers): harden authoritative profile adoption

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* test(oidc): align provider fixtures with profiles

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* fix(providers): preserve authoritative profile lifecycle

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <johntmyers@users.noreply.github.com>
2026-09-01 19:22:07 +00:00
Yuedong Wu e508c169ea fix(helm): honor empty clientCaSecretName for HTTPS-only mode (#2235)
* fix(helm): honor empty clientCaSecretName for HTTPS-only mode

Signed-off-by: Yuedong Wu <dwcn22@outlook.com>

* docs(skills): document HTTPS-only clientCaSecretName in debug-openshell-cluster

Signed-off-by: Yuedong Wu <dwcn22@outlook.com>

---------

Signed-off-by: Yuedong Wu <dwcn22@outlook.com>
2026-09-01 18:10:40 +00:00
John T. Myers 5c541e1e0e fix(kubernetes): prevent stop-start relay race (#3064)
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-08-31 23:40:15 +00:00
Drew Newberry 9b6d904e88 feat(compute): delegate sandbox authentication to drivers (#2968)
* feat(compute): delegate sandbox authentication to drivers

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>

* fix(kubernetes): align sandbox identity annotation

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>

* test(auth): restore sandbox bootstrap coverage

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>

* fix(kubernetes): satisfy ownership test lint

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>

---------

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>
2026-08-31 17:41:56 +00:00
Drew Newberry 69a05ebb3b fix(sandbox): complete successful main processes (#2884) 2026-08-28 18:21:20 -07:00
krishicks f68867b869 feat(gateway): identify gateways in exported traces (#2647)
* feat(gateway): add installation name configuration

Add a first-class operator-assigned gateway name with TOML, CLI, environment,
and Helm configuration surfaces. Local gateways default to openshell, while
Helm defaults to the chart fullname; operators sharing a collector across
namespaces or clusters can set a globally distinct name.

Signed-off-by: Kris Hicks <khicks@nvidia.com>

* feat(gateway): identify gateways in exported traces

Attach the configured gateway installation name and compute driver to the
gateway OpenTelemetry resource so operators can filter traces from multiple
installations that share a collector. Forward the gateway name and OTLP
endpoint to managed external drivers so their distinct service resources carry
the same installation identity.

Keep service.name stable per process type, omit blank resource values, and
leave per-span operation names and request attributes unchanged.

Refs #2507

Signed-off-by: Kris Hicks <khicks@nvidia.com>

---------

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-08-27 17:27:34 +00:00
krishicks d0dfb22baf feat(kubernetes): export driver traces over OTLP (#2958)
Mirror the VM, Podman, and Docker driver tracing setup for Kubernetes.
Export standalone driver spans through OTLP/gRPC as the distinct
openshell-driver-kubernetes service, preserve gateway trace context, record
lifecycle operations and gRPC failures, and flush spans on shutdown.

Kubernetes currently runs in-process when selected as a built-in gateway
driver. Use the temporary server-boundary shim shared with Podman and Docker
so traces retain the shape they will have when Kubernetes moves to a
separate process. Move the common ComputeDriver RPC tracing layer into
openshell-otel to keep all drivers aligned.

Propagate the active W3C context through the controller-reserved Sandbox
annotation and enable Agent Sandbox OTLP export in the local k3s workflow.
This connects asynchronous controller reconciliation spans to the originating
OpenShell create trace.

Expose gateway OTLP configuration through Helm and add an Aspire collector
to the local k3s workflow. Extend helm:k3s:forward with OTLP ingest and trace
UI forwarding for Kubernetes and local container gateway development.

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-08-26 21:22:51 +00:00
Evan Lezar 4e992093f8 fix(docker): trace standalone driver over OTLP (#2923)
Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-08-26 06:27:21 +00:00
Seth Jennings 5206bc51b2 feat(sdk): add OAuth Client Credentials support to SDKs (#2907)
* feat(sdk): add renewable client credentials auth

Implement lazy OAuth client-credentials acquisition and renewal for the Python, TypeScript, and Go SDK clients, with shared security conformance coverage and service-account documentation.\n\nCloses #2803

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(sdk): address client credentials review

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(sdk): address client credentials review

Signed-off-by: Seth Jennings <sjenning@redhat.com>

---------

Signed-off-by: Seth Jennings <sjenning@redhat.com>
2026-08-25 17:53:47 +00:00
Russell Bryant aa848f164d fix(core): fall back to podman CLI when no API socket is found (#1858)
* fix(core): fall back to podman CLI when no API socket responds

Auto-detection only checked well-known Podman socket paths, so a Podman
machine exposing its API socket at a non-standard location went
undetected. The symlink at a well-known path is not always present; it
varies by Podman version, machine provider, and platform.

Extend detect_podman_socket() to fall back to podman CLI discovery when
no well-known candidate responds: podman info --format json determines
whether the service is local or remote, and podman machine inspect
resolves the host-side forwarded socket for VM-backed machines. All
existing callers (driver auto-detection, the Podman driver, and the VM
driver's container-engine fallback) pick this up without change.

Select the machine backing the active Podman connection instead of the
first entry from podman machine inspect, honoring podman's connection
precedence: CONTAINER_CONNECTION, then CONTAINER_HOST (mapped to a
connection by URI), then the containers.conf default. An explicit
endpoint that maps to no known machine is left unresolved rather than
guessing an unrelated machine. When CONTAINER_HOST is an explicit unix://
socket, use that path directly since podman info connects through it.

Update the gateway config reference and the Podman driver README, which
described probe-only detection.

Signed-off-by: Russell Bryant <rbryant@redhat.com>

* fix(core): inspect podman machine by name and bound discovery probes

Address two startup defects in Podman socket auto-detection:

- Remote discovery ran `podman machine inspect` with no arguments, which
  inspects only `podman-machine-default`. On a host whose default
  connection is a different machine, selection fell back to the first
  (wrong) entry, so the gateway could operate against the wrong backend.
  Resolve the active machine first and inspect it by name, trying the
  `-root`-stripped machine for rootful connections and returning None
  when no machine can be mapped instead of substituting another.

- `podman info`, `podman machine inspect`, and `podman system connection
  list` used unbounded `Command::output()`, so a stalled machine, SSH
  connection, or helper could hang gateway startup indefinitely. Route
  all three through a bounded runner that kills and reaps the child on a
  documented deadline and returns None so detection can continue.

Signed-off-by: Russell Bryant <rbryant@redhat.com>

* fix(core): bound podman probe stdout drainage and kill descendants

The previous timeout only bounded waiting for the direct child to exit.
Draining stdout still called read_to_end, which waits for pipe EOF — a
Podman probe can leave a daemonized descendant (SSH multiplexer, gvproxy)
that inherited the stdout pipe, so drainage could wait indefinitely even
after the direct child exited, defeating the deadline.

Run each probe as its own process-group leader, cap stdout drainage by
the same deadline via a channel, and on expiry kill the whole process
group so descendants holding the pipe are terminated and EOF is reached.
Add a regression test where the direct child exits but a backgrounded
descendant keeps stdout open.

Signed-off-by: Russell Bryant <rbryant@redhat.com>

* fix(core): make podman probe deadline absolute against escaped descendants

The prior drainage timeout still ended in an untimed rx.recv() after
killing the process group. A descendant that escaped the probe's process
group (e.g. via setsid) while holding the stdout pipe would not be killed,
so read_to_end never saw EOF, the reader never sent, and that recv() could
block gateway startup indefinitely.

On drainage timeout, best-effort kill the process group and return None
immediately, abandoning the reader thread instead of waiting on it again.
The deadline now bounds the whole call regardless of what descendants do.
Add a regression test whose stdout-holding descendant escapes the process
group (`set -m`) and assert the call still returns promptly.

Signed-off-by: Russell Bryant <rbryant@redhat.com>

---------

Signed-off-by: Russell Bryant <rbryant@redhat.com>
2026-08-25 11:49:22 +00:00
Philippe Martin 0a1f246587 feat(sandbox,podman): trust corporate CA for https:// proxies and intercepted TLS (#2512)
* feat(sandbox,podman): trust corporate CA for https:// proxies and intercepted TLS

The corporate proxy chaining only accepted plain http:// proxy URLs, so
operators whose forward proxy terminates TLS with a private corporate CA
had no way to reach it, and TLS-intercepting proxies (mitmproxy, squid
ssl-bump) that re-sign tunneled server certificates broke every upstream
handshake after CONNECT.

The supervisor now accepts https:// proxy URLs: it wraps the connection to
the proxy in TLS before the CONNECT handshake, verifying the proxy
certificate against the built-in Mozilla roots, the system CA bundle, and an
optional operator corporate CA bundle. The upstream dial returns a
Plain/Tls stream enum consumed generically by the relay paths.

The corporate CA is delivered as a driver-supplied command-line argument
(--upstream-proxy-ca-bundle), never an environment variable, matching the
hardened proxy-config model where a sandbox image cannot influence the
operator's egress boundary. It is folded into the sandbox combined trust
bundle (write_ca_files) and the L7 upstream verification store
(build_upstream_client_config) at startup, so intercepted upstream
handshakes succeed and sandbox workloads trust the re-signed certificates.
Configuration is fail-closed: a CA bundle set without a proxy, or an
unreadable or certificate-free file, is fatal rather than silently
weakening the trust boundary.

The shared parse_upstream_proxy_url validator accepts https:// (recording
the scheme so the driver and supervisor agree), keeping the explicit-port
requirement. The Podman driver gains a proxy_ca_bundle operator setting
(TOML, --sandbox-proxy-ca-bundle, OPENSHELL_SANDBOX_PROXY_CA_BUNDLE) that
bind-mounts the host PEM read-only into the sandbox (a CA certificate is not
secret) and points --upstream-proxy-ca-bundle at it, with a create-time
readability check.

The standalone dev gateway task passes OPENSHELL_SANDBOX_PROXY_CA_BUNDLE
through to the generated podman config, so a local gateway can be pointed at
a TLS-intercepting proxy without hand-editing the regenerated TOML.

Refs #1792

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(sandbox): reject CA bundles with valid PEM framing but invalid X.509 DER

The proxy CA bundle validation counted PEM blocks that base64-decoded
successfully, but did not verify the decoded bytes were accepted as
trust anchors by RootCertStore. A bundle with syntactically valid PEM
framing but invalid DER would pass the startup check while contributing
zero usable anchors, causing opaque TLS failures at runtime instead of
a fail-closed startup error.

Validate decoded certificates through RootCertStore::add_parsable_certificates
and reject the bundle unless at least one is accepted.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(sandbox,podman): allow proxy auth without insecure acknowledgement for https:// proxies

For an https:// proxy the Proxy-Authorization credential travels inside
the verified TLS session, so the proxy_auth_allow_insecure
acknowledgement is unnecessary. Previously both http:// and https://
proxies required it, producing a misleading cleartext-risk diagnostic
for a path that is already encrypted.

Skip the requirement when the proxy URL uses https://; the
acknowledgement is still tolerated if set. Updated in both the
supervisor and Podman driver validation paths, with docs and tests.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix: format

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(kubernetes): use truly unsupported scheme in proxy validation test

https:// is now a supported proxy scheme after a13c4dce, so the
unsupported-scheme test must use a genuinely unsupported scheme.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(e2e): sign the https proxy fixture listener cert with a CA

The corporate-proxy E2E fixture served a single `openssl req -x509`
certificate as its TLS listener identity. OpenSSL marks that certificate
`basicConstraints: critical, CA:TRUE`, and rustls refuses a CA
certificate presented as an end-entity certificate (CaUsedAsEndEntity).
The supervisor's TLS handshake with the proxy therefore failed, the
upstream dial errored, and the workload's CONNECT was dropped without a
response, so podman_corporate_proxy_trusts_ca_bundle_for_https_proxy
failed on the approved destination while policy denial still worked.

Generate a corporate CA and a separate listener leaf signed by it, serve
the leaf chain, and publish only the CA as the bundle the supervisor
trusts. This is what an intercepting proxy actually presents, and it
exercises the corporate-CA trust path rather than pinning the listener
certificate itself.

Refs #1792

Signed-off-by: Philippe Martin <phmartin@redhat.com>

---------

Signed-off-by: Philippe Martin <phmartin@redhat.com>
2026-08-24 21:10:25 +00:00
krishicks e457974a52 feat(docker): export driver traces over OTLP (#2851)
Mirror the VM and Podman driver tracing setup for Docker. Export Docker
driver spans through OTLP/gRPC as the distinct openshell-driver-docker
service, preserve gateway trace context, record lifecycle and asynchronous
provisioning operations, and report gRPC failures.

Docker currently runs in-process when selected as a built-in gateway driver.
Add the same temporary server-boundary shim used by Podman so traces retain
the shape they will have when Docker moves to a separate process. Generalize
the gateway provider selection for both in-process drivers and share the OTLP
collector fixture across Docker, Podman, and VM tracing tests.

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-08-24 18:39:21 +00:00
grs 40d1b48666 feat(provider): support for SPIFFE backed token exchange (#1970)
* feat(provider): add ability to request token exchange instead of client credentials as OAuth grant_type

Signed-off-by: Gordon Sim <gsim@redhat.com>

* test(proxy): add further tests for token exchange

Signed-off-by: Gordon Sim <gsim@redhat.com>

* test(provider): add runnable example for token exchange

Signed-off-by: Gordon Sim <gsim@redhat.com>

* test(e2e): cover Podman token exchange grants

Signed-off-by: Gordon Sim <gsim@redhat.com>

* refactor(oauth): extract duplicated functionality from server and supervisor

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(provider): evict nearest-to-expiry entry from intermediate token cache

Signed-off-by: Gordon Sim <gsim@redhat.com>

* doc(supervisor): add podman example for token exchange

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(provider): withhold token-exchange subject credentials

Signed-off-by: Gordon Sim <gsim@redhat.com>

---------

Signed-off-by: Gordon Sim <gsim@redhat.com>
2026-08-24 05:42:30 +00:00
Drew Newberry ef296806f5 feat(sandbox): add canonical main process (#2726)
* feat(sandbox): add canonical main process

Closes #2710

Persist and supervise one canonical workload per sandbox, attach sandbox connect to its retained session, and make every unexpected main-process exit terminal.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(sandbox): simplify canonical main process contract

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): preserve legacy VM main compatibility

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): preserve main status across driver updates

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): satisfy macOS process lint

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): gate Linux exit acknowledgement publisher

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(sandbox): simplify main process plumbing

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(supervisor): make controlling tty ioctl portable

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(supervisor): initialize canonical process environment

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* perf(supervisor): optimize retained main session

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(sandbox): detach main session on ctrl-c

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): use explicit main detach keys

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(dev): atomically stage Docker supervisor

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(test): align Docker main environment assertion

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(sdk): expose canonical main process fields

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-08-20 22:18:23 +00:00
John T. Myers b2ea81822b feat(network): enable Docker and Podman policy DNS and transparent TCP (#2723)
* feat(network): enable Docker transparent TCP egress

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(e2e): cover Docker transparent TCP egress

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* feat(network): correlate transparent TCP audit events

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(examples): add transparent TCP Redis demo

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(examples): demonstrate blocked TCP connections

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(examples): focus Redis demo audit output

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): close transparent TCP policy bypasses

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(sandbox): reject unsupported TCP policy reloads

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(ci): satisfy Linux transparent TCP lints

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* feat(podman): enable transparent TCP egress

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(podman): permit policy DNS port binding

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(e2e): use qualified transparent TCP hostname

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(podman): preserve exact policy DNS names

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(podman): route policy DNS over TCP

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(dns): serve multiple TCP queries per connection

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(network): explain native DNS and TCP egress

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(sandbox): reconcile runtime reload with upstream

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): harden transparent DNS capture

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(podman): preserve resolver behavior for native tcp

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(network): clarify native tcp runtime constraints

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): remove unused transparent tcp pin

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): admit redirected transparent tcp

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): restore podman transparent networking

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(podman): permit alpine busybox binaries

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(podman): use portable alpine keepalive

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(podman): build musl networking fixture

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(podman): isolate musl DNS probe

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(podman): keep privileged port capability dropped

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): preserve transparent TCP port 53

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): report synthetic pool pressure by family

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(podman): bind tcp fixtures before readiness

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(podman): grant fixture low-port bind

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-08-20 19:26:53 +00:00
John T. Myers 4d7f402ce2 feat(policy): establish direct TCP egress foundation (#2711)
* feat(policy): accept explicit tcp endpoint protocol

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* refactor(network): snapshot authoritative egress decisions

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(policy): document explicit tcp protocol

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(policy): defer transparent TCP release guidance

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* chore(go): regenerate sandbox protobuf bindings

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): complete tcp egress foundation

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(policy): document explicit tcp contract

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): fail closed on authorization errors

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(podman): fence delayed exit events before restart

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(policy): validate network endpoint destinations

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(providers): opt in tcp credential fixture

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(policy): require dns host for transparent tcp

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-08-20 16:43:09 +00:00