Commit Graph
1415 Commits
Author SHA1 Message Date
Shailendra Singh 01ca1d508a test(mxc): authorize curl in the HTTPS egress fixture
Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
2026-10-01 11:24:06 -07:00
pkhodade-NVandShailendra Singh d1ae20a0b6 docs(mxc): add missing demo examples for runbook and mTLS scenario (#3659)
* docs(mxc): add missing demo examples for runbook and mTLS scenario

Add three files present in the internal mirror but absent from the
Windows branch:

- mxc-demo-runbook.md: operator runbook for the MXC demo kit
- README-mtls.txt: usage notes for the mTLS scenario
- run-mtls-test.ps1: PowerShell test runner for the mTLS scenario

These are required by the package-demo.ps1 packager script and are
referenced by the mxc-kit documentation.

No issue required: mechanical sync of missing demo assets.

Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>

* fix(mxc): align demo workflows with current contracts

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>

---------

Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Co-authored-by: Shailendra Singh <shailendras@nvidia.com>
2026-10-01 10:06:38 -07:00
pkhodade-NV 9ceec26d42 fix(mxc): redact injected secrets from gateway diagnostics (#3853)
Redact injected environment values and the per-sandbox proxy password in
captured MXC output and decoded relay launch-failure diagnostics before
logging or publishing sandbox failure status.

Match the original text and redact the union of overlapping occurrences.
Preserve control-channel payloads and avoid allocating for unmatched text.
Add regression coverage and document exact-match and length limits.

Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
2026-10-01 10:02:55 -07:00
araza008 8f22fe84f6 ci(windows): run MXC host probe, WebSocket agent, and OpenClaw forwar… (#3826)
* ci(windows): run MXC host probe, WebSocket agent, and OpenClaw forward examples checks

Add separate hosted CI tasks for the MXC host probe, WebSocket agent,
and OpenClaw forward examples. Use mock workloads to verify gateway,
CLI, driver, and sandbox lifecycle wiring without requiring wxc-exec.

Document that mock passes do not validate forwarding or MXC enforcement.

* fix(tests): enhance environment isolation for WebSocket and OpenClaw mock tests

* fix(mxc): enhance OpenClaw mock validation to require 'Ready' sandbox state

* fix(mxc): restore native process helper in WebSocket example

Restore Invoke-NativeCaptured for the PowerShell argument regression test and delegate CLI execution through it while preserving explicit gateway endpoint selection.

Signed-off-by: Akber Raza <akberr@nvidia.com>

---------

Signed-off-by: Akber Raza <akberr@nvidia.com>
2026-09-30 14:13:26 -07:00
Prekshi Vyas d46a814141 ci(windows): exercise MXC credential, audit, and aggregate E2E flows (#3787)
* ci(windows): exercise MXC provider credential example

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* ci(windows): exercise MXC OCSF audit example

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* ci(windows): enable aggregate MXC example E2E

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* ci(windows): select native MXC mock target

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* fix(windows): isolate MXC example CI harnesses

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

---------

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>
2026-09-29 19:50:53 -07:00
Prekshi Vyas bd49d45cb4 docs(mxc): document Windows host preparation (NVBug 6842834) (#3892)
Document standalone MXC executable placement and condition system-drive ACL preparation on the AppContainer + DACL tier and probe recommendation. Explain the persistent metadata-only grant and link upstream verification and rollback guidance.

NVBug: 6842834

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>
Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
2026-09-29 15:30:25 -07:00
Prekshi Vyas 4cb6054cfe fix(mxc): preserve JSON arguments in PowerShell 5.1 (#3893)
Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>
2026-09-29 15:11:01 -07:00
Prekshi Vyas 20b0ebdea4 docs(mxc): clarify mapper schema versions (#3894)
Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>
2026-09-29 15:10:52 -07:00
Prekshi Vyas d6fea62e76 fix(mxc): add schema version to test configs (#3891)
Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>
2026-09-29 15:08:48 -07:00
Prekshi Vyas 7ba7a39d09 ci(windows): exercise MXC inference demos with mock API (#3780)
* ci(windows): exercise MXC Ollama demo with mock API

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* ci(windows): cover both MXC inference demos

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* fix(mxc): preserve executable extension resolution

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* test(mxc): use absolute PowerShell in lifecycle checks

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* test(mxc): make lifecycle write probes deterministic

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* test(mxc): assert stable lifecycle completion

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

---------

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>
2026-09-29 13:39:23 -07:00
Prekshi Vyas 3451e72700 docs(mxc): correct isolation filesystem support (#3760)
Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>
2026-09-28 08:45:16 -07:00
Shailendra Singh a749bc88c8 docs(windows): add runtime architecture overview (#3575)
* docs(windows): add runtime architecture overview

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>

* docs(windows): consolidate build documentation

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>

* docs(windows): address architecture review feedback

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>

---------

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
2026-09-23 13:52:34 -07:00
pkhodade-NVandShailendra Singh 9e397e8a43 fix(mxc): reject cpu/memory limits instead of silently discarding them (#3548)
* fix(mxc): reject cpu/memory limits instead of silently discarding them

CreateSandbox accepted --cpu/--memory and reached Ready with no Job
Object enforcement and no diagnostic, leaving the SDD's T11
host-exhaustion mitigation silently unmet. MXC's schema does not
expose CPU rate control or memory limiting outside the WSLC backend,
so reject requests carrying cpu/memory limits synchronously at
CreateSandbox, matching the existing fail-closed GPU rejection.

NVBug 6782894

Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>

* fix(server): validate sandbox resource quantities

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>

* docs(skill): clarify MXC resource limit behavior

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>

* chore(go): regenerate protobuf bindings

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>

* fix(go): align generated protobuf comments

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>

---------

Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Co-authored-by: Shailendra Singh <shailendras@nvidia.com>
2026-09-23 11:16:54 -07:00
pkhodade-NV ab64e84bfc fix(core): enforce owner-only Windows ACLs on sensitive files and dirs (#3495)
* fix(core): enforce owner-only Windows ACLs on sensitive files and dirs

set_dir_owner_only/set_file_owner_only were unconditional no-ops on
Windows, so the CLI's mTLS client private key, OIDC/edge tokens, cached
SSH keys, and the gateway's key-encryption key relied entirely on
inherited NTFS ACLs with no OpenShell-applied restriction. Apply an
owner-only DACL via SetEntriesInAclW/SetNamedSecurityInfoW with
PROTECTED_DACL_SECURITY_INFORMATION to strip inherited ACEs, matching
the 0700/0600 guarantee already provided on Unix. is_file_permissions_too_open
now also works on Windows instead of being Unix-only, closing the
detection gap alongside the prevention gap.

Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
(cherry picked from commit 71560e947f85819efbcddf70ddda94befab62b0b)

* fix(core): treat a NULL DACL as too open in is_file_permissions_too_open

has_foreign_trustee conflated a NULL DACL with an unreadable/invalid
ACL and returned Some(false) (not too open) for both. Per the Win32
contract, a NULL DACL means the object grants full access to everyone
-- the most permissive state possible -- so it must be flagged as too
open. Split the null and invalid-ACL branches: null now returns
Some(true), invalid ACL keeps the existing unreadable-ACL fallback
(None, which the caller maps to false via unwrap_or). Adds a
regression test that constructs a real NULL DACL via a
SetNamedSecurityInfoW helper confined to the windows_acl module,
consistent with the existing unsafe-FFI confinement in that module.

Found by CodeRabbit review on MR !113.

Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
(cherry picked from commit 46e635a4ef1d6937cdb088f46aa85baee3d6ad28)

* fix(core): close three false-negative gaps in the Windows ACL audit

restrict_to_current_user() updated only the DACL, leaving a foreign
owner's implicit WRITE_DAC right intact -- they could later replace
the DACL we just set. Query OWNER_SECURITY_INFORMATION and take
ownership in the same SetNamedSecurityInfoW call; if the caller can't
(a genuinely foreign-owned object), the call now fails instead of
silently leaving the object insecure.

is_file_permissions_too_open() mapped every Win32 inspection failure
(missing READ_CONTROL, an invalid ACL, a token-query failure) to
"not too open" via unwrap_or(false). Fail closed instead: an
inspection failure is a security false-negative risk, not a green
light.

has_foreign_trustee()'s ACE loop only recognized plain
ACCESS_ALLOWED_ACE_TYPE and treated every other type as non-granting.
Windows also defines access-allowed object, callback, and
callback-object ACE variants that can grant rights to a foreign
trustee; this audit doesn't parse their wider layouts, so their mere
presence is now conservatively flagged as too open instead of
silently skipped.

Also updates architecture/gateway.md, which still described the
SQLite file-tightening behavior only in terms of Unix mode 0o600, to
distinguish it from the owner-only DACL behavior on Windows.

Addresses review comments on PR #3495.

Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>

* fix(core): conditional owner claim and audit owner in Windows ACL helpers

restrict_to_current_user: query the current owner before calling
SetNamedSecurityInfoW. Include OWNER_SECURITY_INFORMATION only when the
path has a foreign owner -- requesting it unconditionally fails with
ACCESS_DENIED (0x80070005) on standard credentials even when the current
user is already the owner, because WRITE_OWNER is not implied by object
ownership. A foreign-owned path still triggers an ownership claim and
fails hard if the claim is denied, preserving the security contract.

has_foreign_trustee: request OWNER_SECURITY_INFORMATION alongside
DACL_SECURITY_INFORMATION and reject paths with a foreign owner
immediately, before inspecting the DACL. A foreign owner has implicit
WRITE_DAC rights and can replace any DACL we set, so a clean DACL is not
sufficient evidence of safety on a foreign-owned object.

architecture/gateway.md: clarify that the Windows path-hardening behavior
sets mode 0o600 on Unix and applies a protected owner-only DACL on
Windows, with conditional ownership claim and fail-hard semantics for
foreign-owned objects.

Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>

---------

Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
2026-09-23 10:13:32 -07:00
Prekshi Vyas 2d5c9bfaab fix(network): scope Windows egress by socket owner (#3491)
* fix(ocsf): attribute MXC proxy events to sandboxes

NVBug 6783086

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* fix(network): scope Windows egress by socket owner

Resolve each accepted Windows proxy connection to its unique owning PID and executable before evaluating binary-scoped network policy. Fail closed when ownership or process identity cannot be established, and cover allowed and undeclared child processes with a real MXC regression.

* fix(network): preserve sandbox context with socket owners

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>
(cherry picked from commit 32ea316deb6c83c157fdf2de4a2cdc3a7d2da1e3)

* fix(network): harden Windows process identity

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

---------

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>
2026-09-22 20:22:40 -07:00
Prekshi Vyas b031dc037a fix(mxc): reject unsupported live policy updates (#3480)
* fix(mxc): reject unsupported live policy updates (NVBug 6782891)

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* fix(mxc): gate all live policy mutations

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* fix(mxc): gate composed policy mutations

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* fix(ci): satisfy provider update lint

* fix(ci): order provider validation branches

* test(mxc): make policy synchronization deterministic

* fix(server): scope MXC policy synchronization

* fix(server): serialize provider-backed sandbox creation

* test(server): use valid provider create fixture name

---------

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>
2026-09-22 15:44:44 -07:00
pkhodade-NVandShailendra Singh 7305209ac7 fix(mxc): honor the generic sandbox create -- <COMMAND> syntax (#3553)
* fix(mxc): honor the generic sandbox create -- <COMMAND> syntax

sandbox_config only ever read the command from the MXC-specific
driver_config (--driver-config-json), never DriverSandboxSpec.command
-- the portable field the CLI's documented, driver-agnostic
`sandbox create -- <COMMAND>` syntax actually populates, and the same
field every other compute driver honors. A caller following that
syntax got "driver_config.command must contain a non-empty
executable", a message that reads as if no command was supplied at
all.

Add spec.command as a fallback source when no driver_config is
present, and reword the empty-command error to name both ways to
supply one.

NVBug 6782884

Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>

* docs(mxc): document generic sandbox commands

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>

---------

Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Co-authored-by: Shailendra Singh <shailendras@nvidia.com>
2026-09-22 13:49:45 -07:00
Prekshi Vyas 0b8f3821e2 fix(network): normalize Windows binary paths (NVBug 6782969) (#3482)
* fix(network): normalize Windows policy binary paths

Match Windows executable identities using a stable case-insensitive, separator-normalized representation across policy data, L4 input, and L7 relay evaluation. Preserve exact matching on other platforms and keep the original path for hashing and filesystem access.

NVBug 6782969

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* fix(network): harden Windows binary matching

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* fix(ci): scope Windows relay test imports

* fix(network): harden Windows binary path matching

---------

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>
2026-09-22 13:47:25 -07:00
pkhodade-NV e084e1b09e fix(server): keep Error-phase sandbox records through the prune sweep (#3498)
* fix(server): keep Error-phase sandbox records through the prune sweep

The periodic store-vs-backend reconciliation sweep deleted any
persisted sandbox not present in the driver's live backend snapshot,
except for Completed and failed-main-process phases. A driver whose
registry is in-process-only and never rehydrates after a restart (no
persistence of its own) reports every previously-known sandbox as
missing on the very first sweep after startup -- including ones
already correctly, terminally marked Error by earlier crash detection
-- so the sweep silently deleted them shortly after gateway restart,
racing any client (GetSandbox/ListSandboxes/DeleteSandbox) working
with the same sandbox in that window.

Treat Error the same as the existing Completed exemption: it is
already a settled, informational terminal state with no live compute
resource to reclaim, so keep the durable record instead of deleting
it.

Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
(cherry picked from commit 721a1659a822a72e76e3c0dffc6847f17129a3fc)

* fix(server): narrow the prune sweep's Error-phase exemption

The blanket phase == SandboxPhase::Error exemption changed established
missing-backend cleanup for every compute driver, not only the MXC
restart race the PR intended to fix. It also matched
BackendResourceMissing (set by gateway-startup recovery when a
previously-known sandbox's backend resource is already gone),
StartFailed (startup recovery's driver-error case), and
ComputeResourceMissing (this same sweep's own first-pass Error
transition for a Stopping/Stopped/Starting sandbox). All three mark
exactly the orphaned resources this sweep exists to reclaim across
Docker, Podman, VM, Kubernetes, and extension drivers -- exempting
them left orphaned names and gateway-owned records in place
indefinitely and skipped the idempotent driver cleanup for
volumes/secrets until a user explicitly deleted the sandbox.

Add is_missing_compute_resource_reason to inspect the sandbox's Ready
condition and narrow the exemption to a settled Error record only: one
whose reason isn't one of those three. A crashed main process or any
other non-resource failure keeps the exemption (no live resource ever
expected again); a resource-missing reason keeps flowing through the
normal delete-and-cleanup path exactly as before this PR.

Adds regression coverage for BackendResourceMissing and
ComputeResourceMissing confirming they are still pruned with driver
cleanup invoked, and documents the settled-vs-missing-resource
retention distinction in architecture/compute-runtimes.md.

Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>

---------

Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
2026-09-22 13:43:49 -07:00
pkhodade-NV 37c2bae2e6 feat(mxc): warn when the ETW consumer receives zero provider events (#3499)
* feat(mxc): warn when the ETW consumer receives zero provider events

EnableTraceEx2 succeeding only proves the request to enable the
Sandboxing provider succeeded, not that the provider exists on this
host/build or will ever fire. A provider-identity mismatch or a
non-firing provider left the ETW->OCSF audit trail silently empty
across real, successful sandbox lifecycles, with no error, warning,
or diagnostic anywhere.

Track raw provider-matched events received per session and add a
zero-events watchdog on the consumer thread: once real sandbox
activity has happened (register_launch called at least once) and a
grace period elapses with zero events matched, log a warning and
emit a Detection Finding [2004] naming the gap. Gated on actual
activity (not just session uptime) so an idle gateway with
etw_audit=true and no sandboxes created never warns.

Also exposes EtwSession::events_received() alongside the existing
is_capture_alive(), so a status/diagnostics surface can query capture
health directly, not just infer it from tracing output.

The watchdog's decision logic is extracted into a pure function
(should_warn_zero_events) so it's unit-testable without a real ETW
session or elevation.

Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
(cherry picked from commit 4fcfa716a808eaa91c6476e9c69366f91702b1fa)

* fix(mxc): repair build/lint breaks left by the zero-events watchdog commit

warn_zero_events_received() referenced a nonexistent SESSION_NAME
constant (only SESSION_NAME_PREFIX exists), so the crate failed to
compile. Separately, the consumer thread's decode-or-log match on
Option<DecodedEtwEvent> tripped clippy::single_match_else under -D
warnings. Neither issue is specific to a platform or toolchain
version -- both reproduce on a clean checkout of this branch's tip.

Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>

* fix(mxc): attribute ETW Sandboxing-provider events via identity/CV, not PID

The Sandboxing provider's events are logged under two PIDs that are
never the driver's own wxc-exec PID: a short-lived launcher PID
(wxc-exec.exe exits within seconds even while the sandboxed workload
keeps running) and a shared, constant PID hosted by a system-wide OS
broker service across unrelated sandboxes. Anchoring attribution on
`register_launch`'s wxc_pid therefore left every event unattributable,
so the OCSF audit trail stayed empty despite the provider firing
correctly (confirmed against a raw logman/tracerpt capture running
alongside this consumer).

wxc-exec never reports its OS-generated `identity`/`__TlgCV__` back to
the driver, so there is nothing to pre-seed `by_identity` with at
registration time the way `by_pid` is pre-seeded. Track launches still
awaiting their first identity/CV in a `pending_launches` queue instead,
and bind opportunistically in `resolve()`: when an event carries a
never-seen identity/CV and has no `by_pid` registration at all (the
real-world shape of these events) and exactly one launch is pending,
it can only be that launch's burst. Zero or multiple pending launches
stay ambiguous and fall through to the existing unresolved-event
buffer/TTL path rather than guess -- misattributing an audit event to
the wrong sandbox_id is worse than dropping it. A PID registration
that does exist (even generation-mismatched) is treated as positive
evidence of an existing PID-reuse race and takes precedence over the
opportunistic path, preserving the existing generation-key guarantees.

Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>

* fix(mxc): surface dropped-unattributed ETW events above debug level

An event that ages out of the unresolved-event buffer unattributed is
a permanent audit-trail gap: the OS action it represents will never
appear in the OCSF log, and nothing retries it afterward. That was
only visible at --log-level debug, so an operator running with the
default level would never see it. Promote it to warn, matching the
severity already used for the zero-events watchdog's own gap warning.

Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>

* fix(mxc): keep ETW attribution fail-closed, rate-limit drop warnings

The prior opportunistic "single pending launch" fallback attributed
any fresh identity/CV from an unregistered or system-broker PID to
the sole pending OpenShell launch. Queue cardinality is not
attribution evidence: unrelated, non-OpenShell AppContainer or UAC
activity shares this same OS Sandboxing provider, so an event
arriving in the five-second window could cross-link an unrelated
identity and emit subsequent events under the wrong sandbox_id,
corrupting the audit trail. Revert it -- attribution requires the
driver-owned wxc-exec PID and its kernel process start key to both
match, exactly as before. Records without that generation-backed
evidence remain unattributed rather than guessed.

The dropped-unattributed warning also moved from one line per record
to a rate-limited, aggregated warning: unattributed drops are
expected, ordinary system-wide activity, not a rare condition, so
warning per record could flood operator logs during a burst of
unrelated AppContainer/UAC activity. UnattributedDropReporter mirrors
the existing OverloadReporter pattern -- warn immediately on the
first drop, then coalesce to one warning with a running count every
30 seconds while drops continue.

Updates the MXC observability documentation to match: the
unattributed-record sentence now explains why fail-closed is
deliberate, documents the coalesced warning cadence, and documents
the mxc-etw-zero-events OCSF Detection Finding [2004], which was not
mentioned anywhere before.

Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>

---------

Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
2026-09-22 13:41:12 -07:00
pkhodade-NV 27a36b4004 fix(mxc): reject a non-absolute wxc_exec_path at gateway startup (#3497)
* fix(mxc): reject a non-absolute wxc_exec_path at gateway startup

wxc_exec_path is the binary that builds every sandbox, but nothing
validated it before spawning: a relative value (including the shipped
default, a bare "wxc-exec.exe") let PATH-lookup or working-directory-
relative resolution execute a decoy binary with the gateway's identity
instead of the approved wxc-exec, turning the containment mechanism
itself into an arbitrary-code-execution primitive.

Add MxcComputeConfig::validate_configuration, wired into the existing
(previously no-op) compute-driver config preflight, rejecting an
empty or non-absolute wxc_exec_path with a clear diagnostic. Change
the default from the relative "wxc-exec.exe" to an empty string so
the field must be explicitly configured -- no usable-but-insecure
fallback survives. Update the architecture doc's stale "else PATH"
discovery claim to match.

Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
(cherry picked from commit d4192a0072f8f0ca0a4035f10cae07889c2b578f)

* test(mxc): cover wxc_exec_path enforcement at the gateway boundary

MxcComputeConfig::validate_configuration already had unit coverage,
but that only proves the validation function itself is correct -- it
says nothing about whether MxcFactory::validate_config (src/lib.rs)
still calls it. Before this fix, that factory method discarded the
parsed config entirely, so a regression back to that no-op shape would
leave every unit test passing while a relative wxc_exec_path again
reached gateway startup.

Add an integration test that spawns the actual compiled
openshell-gateway binary through its config preflight subcommand,
exercising the real chain: CLI parsing, TOML loading, driver
selection, MxcFactory::validate_config, and
MxcComputeConfig::validate_configuration. Assertions check only
pass/fail, not message content: run_effective_config_preflight
replaces any validation failure with a generic message whenever a
config file is used, to keep file-sourced values out of preflight
diagnostics -- pre-existing, deliberate, and covered by its own tests.

Also document the new required-and-absolute wxc_exec_path constraint
in the gateway config reference and the driver README, which
previously only showed example values without stating the
requirement.

Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>

---------

Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
2026-09-22 13:39:43 -07:00
Prekshi VyasandShailendra Singh 651ed7e03d NVBug 6783374: make MXC HTTPS L7 qualification authoritative (#3479)
* test(mxc): qualify HTTPS L7 enforcement (NVBug 6783374)

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* docs(mxc): clarify real qualification failures

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* docs(mxc): align real test skip semantics

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>

---------

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>
Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Co-authored-by: Shailendra Singh <shailendras@nvidia.com>
2026-09-22 12:38:07 -07:00
Prekshi Vyas 69d73a7eaf fix(ocsf): attribute MXC proxy events to sandboxes (#3434)
* fix(ocsf): attribute MXC proxy events to sandboxes

NVBug 6783086

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* fix(ocsf): preserve sandbox context across proxy denials

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

---------

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>
2026-09-22 10:55:34 -07:00
Prekshi Vyas a69d0319f2 test(windows): define GB300 MXC qualification contract (NVBug 6643699) (#3471)
* test(windows): define GB300 MXC qualification contract

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* test(mxc): harden GB300 qualification provenance

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* test(windows): package portable GB300 qualification

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* test(windows): honor external qualification checkout

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* revert: remove portable GB300 packaging

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* test(mxc): bind GB300 evidence to exact inputs

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

---------

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>
2026-09-22 10:52:21 -07:00
pkhodade-NVandPrekshi Vyas 207eb56242 fix(mxc): stage egress-proxy CA files regardless of env tier (#3496)
* fix(mxc): stage egress-proxy CA files regardless of env tier

stage_tls_ca_files was only invoked when pc_minimal_env was set, but a
curated ProcessContainer can't read the host proxy's private temp
folder under any env tier. With the default pc_minimal_env=false, the
CA env vars pointed at a path the sandboxed process couldn't read at
all, breaking TLS validation for allow-listed HTTPS requests through
the egress proxy. Extract the staging decision into
resolve_agent_proxy_ca_paths, which takes no env-tier argument, so the
gap can't silently regress; existing tests never exercised this path
since mocked invokers skip host-proxy startup entirely.

Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
(cherry picked from commit d0346af96aef9a93aca26a7e1825d8678a25567d)

* fix(mxc): complete proxy CA staging isolation (#3535)

* test(mxc): provide workload dir for CA staging

Signed-off-by: Prekshi Vyas <prekshivyas@nvidia.com>

* fix(mxc): isolate staged proxy CA files

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

---------

Signed-off-by: Prekshi Vyas <prekshivyas@nvidia.com>
Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

---------

Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
Signed-off-by: Prekshi Vyas <prekshivyas@nvidia.com>
Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>
Co-authored-by: Prekshi Vyas <34834085+prekshivyas@users.noreply.github.com>
2026-09-21 20:37:11 -07:00
Prekshi Vyas 3a4d4f98f7 test(mxc): parse OpenClaw config schema portably (#3536)
Signed-off-by: Prekshi Vyas <prekshivyas@nvidia.com>
(cherry picked from commit 3a0efc2d630bc98da297f0fef54ef99f8125662a)

# Conflicts:
#	crates/openshell-driver-mxc/Cargo.toml
2026-09-21 20:35:15 -07:00
Prekshi Vyas 80b35dbfe0 fix(mxc): repair Windows inference demos (NVBug 6782874) (#3473)
* fix(mxc): repair Windows inference demos

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* fix(mxc): address inference demo review feedback

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

---------

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>
2026-09-21 17:28:52 -07:00
Prekshi Vyas b2d6280794 fix(windows): repair OpenClaw ProcessContainer startup (NVBug 6782898) (#3475) 2026-09-21 16:17:48 -07:00
Prekshi Vyas d78430c66c test(mxc): verify ProcessContainer token isolation (#3430)
Adds real-MXC regression coverage for ProcessContainer token isolation, including an unsandboxed SCM positive control, and documents the AppContainer authorization model.
2026-09-21 15:25:31 -07:00
Prekshi VyasandShailendra Singh 4f06e23cc8 fix(mxc): harden governed proxy lifecycle (NVBug 6783325) (#3472)
* fix(mxc): gate host proxy on explicit network policy

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* fix(mxc): reject unrestricted egress fallback

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* chore(mise): refresh lockfile

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>

* fix(mxc): unblock Windows validation

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>

* test(mxc): satisfy Windows clippy

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>

---------

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>
Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Co-authored-by: Shailendra Singh <shailendras@nvidia.com>
2026-09-19 18:12:12 -07:00
Prekshi Vyas fb2980e077 fix(windows): restore MXC qualification and cold-start readiness (#3468)
* fix(windows): restore MXC qualification gates

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* fix(mxc): poll target for full readiness budget

---------

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>
2026-09-18 23:01:55 +05:30
Drew Newberry 49b4f0eb7f feat(mxc): add UI policy, credentials, relay lifecycle, and proxy auth
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-17 12:06:06 -07:00
Shiju 8751e35e28 fix(supervisor-network): reject malformed OPA policy containers (#3337)
Validate raw OPA object and array containers before normalization and
access-preset expansion can skip malformed values. Return one fixed
structural error without embedding authored policy data.

Preserve versionless and runtime-only OPA data and existing semantic
validation. Cover initial string/file loading, middleware callback order,
valid deny-rule enforcement, and rejected reload state and generation.
Document the loader contract and its engine-local rejection behavior.

Signed-off-by: Shiju <shiju@nvidia.com>
2026-09-15 20:00:51 +00:00
Seth Jennings dbe36eaf85 fix(security): harden Vault credential transport (#3329)
Reject non-loopback plaintext Vault endpoints, disable redirects, and support private CA bundles without weakening hostname verification. Update Helm configuration, documentation, operator skills, and regression coverage for OSSR-002.

Signed-off-by: Seth Jennings <sjenning@redhat.com>
2026-09-15 19:56:09 +00:00
John T. Myers 607db99915 fix(deps): refresh gateway Debian runtime image (#3350)
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-09-15 19:38:13 +00:00
John T. MyersandJohn Myers dfd5238d0d fix(gator): make supervised lifecycle sandbox-native (#3343)
* fix(gator): run supervisor as sandbox main process

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* feat(gator): persist supervised state history

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* refactor(gator): remove obsolete background launch mode

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <johntmyers@users.noreply.github.com>
2026-09-15 19:04:13 +00:00
krishicks 481ce566e1 fix(ocsf): correct HTTP activity context (#3316)
Previously, metadata events omitted both HTTP request and response objects,
early proxy rejections used HTTP Activity without request context, and
unsupported-scheme events did not expose enough safe HTTP context to satisfy
the OCSF 1.8 schema.

Now, metadata events include a method-only request and their actual HTTP
response codes without recording the metadata URL. Unsupported-scheme events
also include a method-only request plus the generated 400 response. Authority
mismatches and credential-resolution denials use HTTP Activity with their
generated 403 or 500 responses, and HTTP activity IDs are derived from the
request method.

Additionally, HttpActivityBuilder now enforces the OCSF request-or-response
constraint at compile time.

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-09-15 18:47:55 +00:00
Mrunal Patel 39cf4823f7 feat(api): add structured gateway errors and SDK decoding (#3313)
* feat(api): expose structured gateway errors across SDKs

Refs #3051. Add standard validation, conflict, and retry details; preserve raw transport status in Rust, Go, TypeScript, and Python; document status and recovery guidance.

This is the structured-error foundation only. Mutation result shapes, allow_missing, durable request deduplication, and exec retry semantics remain follow-up work.

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(python): preserve wrapped RPC cleanup handling

Inspect the original gRPC call when handling missing sandboxes during deletion waits and managed cleanup. Add intercepted cleanup regressions and clarify the error-wrapper migration contract.

Addresses the cleanup review on #3313; part of #3051.

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

---------

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
2026-09-15 17:51:11 +00:00
Mrunal Patel b799fccb8b fix(auth): harden OIDC trust root retrieval (#3332)
* fix(auth): harden OIDC trust root retrieval

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(e2e): pass OIDC HTTP acknowledgement value

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

---------

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
2026-09-15 17:00:45 +00:00
Evan Lezar c195e23267 test(conformance): remove plan-driven continuity tests (#3342)
Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-09-15 14:57:07 +00:00
Shiju fd3fd9cf74 feat(sandbox): explain failed calls to external tool servers (#3207)
Show configured tool server addresses and their last observed connection
results together in sandbox status. Keep sandbox lifecycle readiness
separate so an external connection failure does not mark the sandbox unready.

Expose direct endpoint records through the CLI and SDKs, with plain-language
failure explanations and gateway acceptance times. Keep observation tracking,
runtime reporting, and gateway validation in dedicated endpoint status modules.

Preserve bounded reporting, request attribution, retry ordering, and
configuration and supervisor authority checks. Clear obsolete observations
while retaining the configured addresses, and document the distinction
between an observed HTTP response, current availability, and tool success.

Signed-off-by: Shiju <shiju@nvidia.com>
2026-09-15 13:46:38 +00:00
26f2f96393 feat(mxc): add Windows host proxy for MXC sandbox network egress (#3163)
* Implement Windows host proxy integration and update dependencies for OpenShell

* Update README and gateway config to clarify egress proxy address handling and allocation

* Refactor ProxyIdentityMode to return Result for static_binary and add tests for binary path and SHA256 hash

* Enhance platform_hosts_path for Windows to use SystemRoot and improve error handling for hosts file reading

* Refactor FileFingerprint to use Option for mtime and ctime, simplifying metadata handling

* Add conditional compilation for Windows host module

* add unit tests for OPA policy evaluation and identity handling

* remove openshell-supervisor-network from unsupported driver package test exclusion list

* feat(mxc): enable host proxy TLS state generation

Generate per-sandbox TLS state for the MXC host proxy so HTTPS L7 enforcement can use the same MITM path as Linux. Grant generated CA material to the MXC process and inject standard trust env vars, while matching Linux behavior by disabling TLS termination on CA setup failure and relying on proxy fail-closed handling.

* fix(docs): remove outdated notes on governed egress from docs

* fix(tests): update TLS environment variable paths to use temporary directory

* fix(examples): make run-mxc-e2e harness correct and orphan-free

The MXC e2e harness never actually exercised the fs scenarios: it started
the gateway once and patched agent_command per scenario AFTERWARDS, so the
running gateway kept launching the default demo agent (not shipped in the
kit) and every fs scenario failed with CreateProcessW error:2. It also
scored on the `sandbox create` exit code (non-zero due to the harmless
interactive attach), wrote sandbox records to the persistent gateway DB
(leaving orphans that collided on later runs), and its deny scenarios never
proved denial.

Changes:
- Start a FRESH gateway per scenario so each scenario's agent_command is
  actually loaded (root cause of CreateProcessW error:2).
- Score by on-disk artifact / expected outcome, not `sandbox create` exit.
- Real deny assertions: a control write to a granted path must succeed
  (proves the agent ran) while the denied write must be absent. fs-empty
  probes an ungranted out-of-share path (share_dir is mapped rw by design).
- Run the gateway on an ephemeral in-memory DB (sqlite::memory:) so the
  harness never writes to the persistent store and cannot leave orphan
  sandbox records; also use unique per-run sandbox names + pre-delete.
- Fix the process_container probe: use a real cwd + absolute cmd.exe
  (canonical wxc-exec does not expand %TEMP% -> 0x8007010B).
- Fix summary counts (@() so a single FAIL is counted and exit is non-zero).

Verified PASS=4 FAIL=0 on 7F203-MXC-003 (no BaseContainer velocity keys)
using a canonical wxc-exec build (AppContainer fallback).

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(e2e): probe timeout is milliseconds (10ms->30000ms)

MXC process.timeout is wall-clock ms (wire.rs). The 10 value meant 10ms,
which the base-container tier (7F203-MXC-001/.181) enforced strictly and
timed the probe out. AppContainer path (.18/-003) happened to slip under
it. Bump to 30000ms so the process_container preflight probe is reliable
across both tiers.

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc): use native paths in real runtime probes

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc): make processcontainer work with mxc-latest-released wxc-exec

Three fixes to support the release wxc-exec binary (BaseContainer dispatcher)
in addition to mxc-fixes-env-vars:

1. Seed process env from host (driver.rs)
   ProcessContainer starts with a completely blank environment -- no PATH,
   SystemRoot, or anything.  Seed the process env from the gateway host
   environment so the agent binary can locate DLLs and run.  Skip internal
   Windows drive-letter variables (keys starting with '=') which cause
   CreateProcessW to return ERROR_ENVVAR_NOT_FOUND.  User agent_env entries
   and TLS CA vars are applied as overrides on top of the host env.

2. Remove TLS readonly_paths grant (driver.rs)
   The release wxc-exec (BaseContainer dispatcher) requires write-DAC
   permission on every path in readonly_paths to set up AppContainer ACLs.
   Adding the proxy's temp TLS directory caused a DACL error and exit -1.
   The CA cert paths remain available to the agent via TLS env vars.

3. Remove allowedHosts from network JSON (mxc.rs)
   The release wxc-exec rejects network.allowedHosts / network.blockedHosts
   on Windows with "not yet supported".  Removed the loopback exemption
   attempt (127.0.0.1, ::1, localhost) from the network section.
   Intra-container loopback works natively in the release binary without
   it -- the spawner can connect to the server at 127.0.0.1:22000 directly.

Additional changes:
- mxc-ws-agent.rs: add relay-debug.txt error capture and relay-ready.txt
  marker for reliable timing of host client connections.
- mxc-ws-gateway.toml: debug = true for JSON config dump during diagnosis.
- run-ws-agent-test.ps1: default port changed to 17670 (gateway default);
  relay-ready.txt polling before ws-echo to avoid connecting before the
  spawner has established the proxy bridge.

Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(e2e): address CodeRabbit review on run-mxc-e2e.ps1 (MR !46)

Four robustness/correctness fixes from CodeRabbit:

1. Start-Gw: kill the spawned gateway before the "did not start within 30s"
   throw. If the process is alive but never binds the port, $gw is not yet
   assigned in the caller, so the finally block cannot reap it -> orphan
   gateway holding the port for the next run.

2. create-fail scoring: a non-zero `sandbox create` exit alone is not proof
   of a policy rejection (gateway-registration/transport/fixture errors also
   exit non-zero and would false-pass). PASS now requires a genuine
   rejection signal (network / invalid_argument / network_policies) AND that
   it is not an infrastructure failure; other non-zero exits go to FAIL with
   output captured.

3. deny scenarios (ControlTarget path): snapshot the deny target AFTER
   Wait-File lands the control artifact, so a late denied write (enforcement
   regression racing the control write) can no longer be recorded as PASS.

4. -KeepRunning: break out of the scenario loop after the first scenario so
   a later scenario does not start a second gateway on the same port
   (previously a reliable port collision instead of a usable debug mode).

Re-verified PASS=4 FAIL=0 on both boxes (7F203-MXC-001 base-container and
7F203-MXC-003 AppContainer fallback); network-policy-rejected correctly
scores as "policy rejection".

Signed-off-by: Akber Raza <akberr@nvidia.com>

* feat(mxc-e2e): collect run-mxc-e2e output into a results bundle

Mirror the sibling run-*.ps1 scripts by collecting every run's logs into a
timestamped results-e2e-<stamp>\ folder and zipping it. The bundle contains the
console transcript, per-scenario gateway stdout/stderr, the exact TOML rendered
for each scenario, the policy fixture used, and a summary.txt with the verdict
table.

Per-scenario gateway logs now land in gateway.<scenario>.log/.err.log inside the
bundle instead of a single fixed gateway.e2e.log in the script directory.

Wrap pre-flight, mode setup, scenario definitions, and the scenario loop in a
single try/catch/finally so the finally always writes the summary, stops the
transcript, and zips the bundle -- even on a pre-flight failure. The existing
per-scenario gateway-cleanup try/finally stays nested inside. All scenario
logic, scoring rules, and comments are preserved.

Signed-off-by: Prashant S Khodade <pkhodade@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc-e2e): address CodeRabbit review on run-mxc-e2e.ps1

- Require -Scenario when -KeepRunning: the loop breaks after the first
  scenario, so a full-suite run would execute only one scenario yet still
  report the suite as PASS. Fail fast so a partial run can't be mislabeled
  complete.
- Start-Transcript now runs inside the guarded try block with a
  $transcriptStarted flag; Stop-Transcript is only called when it actually
  started, so a Start-Transcript failure still yields the results bundle.
- Wrap the -Scenario filter in @() so a single exact match stays an array
  (reliable .Count and a proper array for the scenario loop on PS 5.1).

Signed-off-by: Prashant S Khodade <pkhodade@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(examples): pass gateway config via OPENSHELL_GATEWAY_CONFIG for spaced paths

Start-Process -ArgumentList does not quote array elements, so launching the
gateway with a bare --config <path> token split on any space in the install
path (e.g. C:\Users\First Last\...), and clap rejected the fragment with
'unrecognized subcommand'. Every MXC example launcher that started the gateway
hit this when the kit was unzipped under a path containing a space.

Pass the config path through the OPENSHELL_GATEWAY_CONFIG env var (which the
gateway already reads via clap) and drop the --config token. Env vars carry
spaces safely.

Affected: run-ocsf-audit, run-mxc-e2e, run-demo, run-inference-test,
run-ollama-test. run-mtls-test was not affected (its launch passes no config
path). Root-caused and fix-verified on 7F203-MXC-003 from a spaced path.

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(run-mxc-e2e): improve scoring logic and enhance command execution handling

* fix(mxc): reconcile proxy support after rebase

Restore the proxy-enabled OCSF audit example removed by 13185f6e now that the host CONNECT proxy is present. Adapt the proxy lifecycle test to the target branch's DriverSandboxSpec policy delivery contract.

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc): grant sandbox access to proxy CA

- Share the per-sandbox public CA bundle with the AppContainer
- Add real wxc-exec HTTPS proxy coverage and document trust isolation

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc): reject unsupported network middleware

- Reject middleware-bearing MXC policies before sandbox lifecycle begins
- Guard host proxy startup and document the unsupported registry path
- Add mapper, lifecycle, and host proxy regression coverage

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc): reconcile host proxy with main

- remove obsolete inference routing from the host proxy adapter
- use the workspace AWS-LC provider in host-proxy tests
- adapt the forward-proxy test to ProxyIdentityMode

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(build): switch to bundled Z3 for Windows MSVC builds

* fix(mxc): reconcile host proxy after rebase

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Akber Raza <akberr@nvidia.com>
Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
Signed-off-by: Prashant S Khodade <pkhodade@nvidia.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Jamie King <jamiek@nvidia.com>
Co-authored-by: Shailendra Singh <shailendras@nvidia.com>
Co-authored-by: Prashant Khodade <pkhodade@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
2026-09-14 22:58:49 +00:00
Drew Newberry 2d5db4c5bd chore(security): document Kubernetes runtime RBAC (#3328)
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-14 22:11:42 +00:00
Jorge 42e9bcf2b0 feat(e2e): support the Vault credential-driver lane on OpenShift (#3312)
Implements https://github.com/NVIDIA/OpenShell/issues/3212

Running e2e:kubernetes with the Vault credential driver failed on
OpenShift in two ways: the OpenBao fixture pod was rejected by the
restricted-v2 SCC, and provider-creating tests hit HTTP 403 from
auth/kubernetes/login because OpenBao's Kubernetes auth was provisioned
only inside a single feature-gated test.

- Deploy OpenBao with the chart's OpenShift mode (global.openshift=true)
  when OpenShift is detected, so the pod inherits a namespace-assigned,
  SCC-compliant security context with no manual SCC grant. Hoist
  OpenShift detection ahead of the credential-driver fixtures so the
  flag is set before the fixture is deployed.
- Provision the OpenBao KV store, Kubernetes auth method, storage
  policy, and gateway login role in the harness (deploy_vault_fixture),
  making a Vault-backed gateway usable by the whole suite instead of
  only the credential_drivers test. Remove the now-redundant
  configure_vault_storage helper from the test.
- Harden openbao_exec so it tolerates only the idempotent "path is
  already in use" error on reruns and fails fast with output on any
  other error, instead of a blanket `|| true` that masked genuine
  failures (e.g. an unresponsive pod) until a later cryptic write.
- Document the OpenShift Vault credential-store SCC and Kubernetes-auth
  403 troubleshooting in the debug-openshell-cluster skill.

Signed-off-by: Jorge Garcia Oncins <jgarciao@redhat.com>
2026-09-14 21:55:43 +00:00
Brandon Squizzato cc780d4e17 feat(helm): add BackendTLSPolicy support (#2728)
* feat(helm): add optional BackendTLSPolicy for e2e TLS

Add grpcRoute.backendTLSPolicy values to optionally create a
BackendTLSPolicy resource that enables end-to-end TLS between the
Gateway proxy and the OpenShell gateway pod. The Gateway proxy
terminates client-facing TLS and re-encrypts when connecting to the
backend, validating the pod's certificate against a user-supplied CA
ConfigMap.

This removes the requirement to set server.disableTls=true when using
HTTPS at the Gateway listener. Supported on OpenShift 4.22+ and other
platforms with BackendTLSPolicy support in the Gateway API
implementation.

Update OpenShift and ingress documentation with e2e TLS instructions.

Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>

* feat(helm,server): auto-create backend CA ConfigMap in certgen hook

Extend the generate-certs command with --backend-ca-configmap-name and
--backend-ca-source-secret flags. When BackendTLSPolicy is enabled, the
certgen pre-install hook creates the CA ConfigMap automatically:

- pkiInitJob mode (default): uses the CA from the generated PKI bundle.
  Fully automatic on first install.
- cert-manager mode: reads ca.crt from the server TLS Secret. On first
  install the Secret does not exist yet (cert-manager reconciles after
  templates are applied), so the ConfigMap is created on the first helm
  upgrade. Logs a warning on the initial skip.

The caCertificateConfigMapName value now defaults to <fullname>-backend-ca
when empty, so users only need to set backendTLSPolicy.enabled=true.

Update certgen RBAC to include configmaps get/create when the feature is
enabled. Add CLI arg parsing tests for the new flags.

Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>

* refactor(helm): add server.tls.enableMtls flag for mTLS control

Replace automatic mTLS disabling based on BackendTLSPolicy with an
explicit server.tls.enableMtls flag that defaults to true. The user is
now responsible for setting this to false when using BackendTLSPolicy,
as ingress proxies cannot present client certificates to backends.

Updated:
- values.yaml: Added server.tls.enableMtls (default true)
- gateway-config.yaml: Check enableMtls instead of backendTLSPolicy
- _gateway-workload.tpl: Check enableMtls for client CA mount
- Tests: Updated to use enableMtls flag
- Docs: Added enableMtls=false to BackendTLSPolicy examples
- README: Document new flag and BackendTLSPolicy requirement

Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>

* docs(helm): clarify cert-manager backend CA ConfigMap workflow

Update documentation to explain the two-step install process required
when using cert-manager with BackendTLSPolicy:

1. helm install - cert-manager issues the server certificate, but the
   certgen hook can't create the backend CA ConfigMap yet (cert-manager
   reconciles after templates are applied)
2. helm upgrade - certgen hook reads the CA from the cert-manager-issued
   certificate and creates the ConfigMap

Previously, the docs said "created on first upgrade" without explaining
why or that the feature won't work until then. The updated docs now:
- Explain the timing issue (cert-manager reconciles after chart install)
- Provide clear steps for the cert-manager workflow
- Note that pkiInitJob (default) creates it immediately on install
- Clarify that users must wait for the Certificate to be Ready before
  running the second upgrade

Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>

* fix(docs): remove incorrect external hostname requirement for BackendTLSPolicy

BackendTLSPolicy validates the backend certificate against the service FQDN
(e.g., openshell.openshell.svc.cluster.local), not the external hostname.
The external hostname only needs to be on the Gateway listener certificate
for client-facing TLS.

The default certManager.serverDnsNames already includes all required service
FQDN variants, so no configuration is needed for BackendTLSPolicy to work.

Fixed incorrect documentation that claimed:
- "The server certificate SAN list must include the external hostname"
- Users need to "configure certManager.serverDnsNames with the external hostname"

Removed the unnecessary pkiInitJob.serverDnsNames override from the example
and clarified that:
- Gateway listener certificate needs the external hostname (for clients)
- Backend certificate needs the service FQDN (for Gateway proxy)
- The service FQDN is already in the defaults

Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>

* docs: clarify ACME with LetsEncrypt reference

Change all references from "ACME issuer" to "LetsEncrypt/ACME issuer"
to help users understand that LetsEncrypt is the most common ACME
provider and what ACME means in practice.

Updated:
- docs/kubernetes/managing-certificates.mdx
- docs/kubernetes/openshift.mdx
- deploy/helm/openshell/values.yaml
- deploy/helm/openshell/README.md
- deploy/helm/openshell/ci/values-openshift-route-cert-manager.yaml

Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>

* docs(openshift): restructure end-to-end TLS options and clarify Gateway hostname

Reorganize the OpenShift production deployment documentation:

1. Changed main section from "Production Deployments" to "Options for
   end-to-end TLS" for better clarity

2. Renamed subsections for consistency and clarity:
   - "End-to-end TLS using Gateway API and BackendTLSPolicy (OpenShift 4.22+)"
   - "End-to-end TLS using pass-through Route (all OpenShift versions)"

3. Clarified that the Gateway hostname is typically a wildcard:
   "typically a wildcard like *.openshell-ingress-gw.example.com"

4. Removed the recommendation to copy the cluster's wildcard certificate
   from openshift-ingress namespace, as this is not a recommended
   security best practice

These changes make it clearer that users have two end-to-end TLS options
and help them understand the typical naming pattern for Gateway hostnames.

Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>

* feat(helm): eliminate two-stage install for BackendTLSPolicy with cert-manager

When using BackendTLSPolicy with cert-manager, the certgen hook now polls
for up to 90 seconds waiting for cert-manager to issue the TLS certificate
before creating the backend CA ConfigMap. This eliminates the need for a
second `helm upgrade` in most cases.

The hook polls every 2 seconds with progress logging every 10 seconds.
If cert-manager takes longer than 90 seconds, the hook times out gracefully
and logs a warning, preserving the fallback to manual ConfigMap creation
or a second upgrade.

The Job's activeDeadlineSeconds is 120s, so the 90s timeout leaves 30s
margin for ConfigMap creation and hook completion.

Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com>

* feat(helm): add configurable timeout for certgen hook

Add `pkiInitJob.timeoutSeconds` Helm value (default 120) to control how long
the certgen hook Job can run. When using cert-manager with BackendTLSPolicy,
the hook polls for (timeoutSeconds - 30) seconds to leave margin for ConfigMap
creation and cleanup.

This allows users to increase the timeout for environments where cert-manager
takes longer than 90 seconds to issue certificates, without requiring code
changes.

Example usage:
```yaml
pkiInitJob:
  timeoutSeconds: 180  # Hook polls for 150 seconds
```

Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com>

* docs(helm): document configurable certgen timeout

Update documentation to mention the pkiInitJob.timeoutSeconds value and
how it affects the cert-manager polling behavior when using BackendTLSPolicy.

Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com>

* feat(helm): add configurable failure behavior for certgen timeout

Add `pkiInitJob.failOnTimeout` Helm value (default false) to control whether
the certgen hook fails or succeeds when cert-manager does not issue a
certificate within the polling timeout.

When false (default), the hook succeeds with a warning and users can run
`helm upgrade` after cert-manager issues the certificate to create the
backend CA ConfigMap. This provides backwards-compatible behavior.

When true, the hook fails immediately if the timeout is reached, providing
clear feedback that BackendTLSPolicy is non-functional. This is useful for
strict validation requirements where incomplete installs should fail fast.

Example usage:
```yaml
pkiInitJob:
  timeoutSeconds: 180
  failOnTimeout: true  # Fail install if cert-manager takes >150s
```

Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com>

* feat(helm): change failOnTimeout default to true and add troubleshooting docs

Change `pkiInitJob.failOnTimeout` default from false to true to provide
immediate feedback when cert-manager does not issue certificates within
the polling timeout. This prevents silent failures where BackendTLSPolicy
is non-functional but the install appears to succeed.

Add comprehensive troubleshooting section to docs/kubernetes/ingress.mdx
documenting the specific error "TLS error: Secret is not supplied by SDS"
that occurs when the backend CA ConfigMap is missing, with step-by-step
resolution instructions.

Updated comments in values.yaml to clearly document the default behavior
and explain when administrators might see connectivity errors if they
override the default to failOnTimeout=false.

BREAKING CHANGE: pkiInitJob.failOnTimeout now defaults to true. Helm
installs will fail if cert-manager takes longer than (timeoutSeconds - 30)
seconds to issue certificates. To restore the old behavior of allowing
installs to succeed with a warning, set `pkiInitJob.failOnTimeout=false`.

Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com>

* fix(helm): make cert-manager resources pre-install hooks to fix ordering

Make Certificate and Issuer resources run as pre-install/pre-upgrade hooks
with weight -30, before the certgen hook (weight -20). This fixes the
chicken-and-egg problem where the certgen hook was waiting for Secrets
created by Certificates that hadn't been created yet.

**Hook ordering:**
1. Certificate and Issuer resources created (weight -30)
2. cert-manager issues certificates and creates Secrets
3. certgen hook runs (weight -20), finds Secrets, creates ConfigMap
4. Main resources (StatefulSet, Service, etc.) created

Previously, the certgen pre-install hook would run before any resources
were created, poll for a non-existent Secret, timeout, and fail. The
Certificate resources would never get created because Helm waits for
all pre-install hooks to succeed before creating main resources.

This fix allows single-stage installs to work reliably as long as
cert-manager can issue certificates within the polling timeout.

Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com>

* feat(helm): add validation to prevent enableMtls with BackendTLSPolicy

Add Helm chart validation that fails the install if both
server.tls.enableMtls=true and grpcRoute.backendTLSPolicy.enabled=true
are set, since this is an invalid configuration.

BackendTLSPolicy requires mTLS to be disabled because the Gateway proxy
cannot present client certificates to the backend. This validation provides
immediate, clear feedback at install time rather than allowing the
misconfiguration to be discovered through runtime errors.

Example error message:
```
Error: grpcRoute.backendTLSPolicy requires mTLS to be disabled because
the Gateway proxy cannot present client certificates to the backend;
set server.tls.enableMtls=false
```

Also updated documentation to mention this validation check.

Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com>

* docs(helm): clarify pkiInitJob.timeoutSeconds polling behavior

Improve documentation to clearly explain that pkiInitJob.timeoutSeconds
controls the Job deadline, but the actual polling timeout is
(timeoutSeconds - 30) to reserve 30 seconds for ConfigMap creation
and cleanup.

Added concrete example: "timeoutSeconds=180 allows 150 seconds of polling"
to make the relationship explicit and avoid confusion where users might
expect the hook to poll for the full timeout value.

Updated both values.yaml inline comments and ingress.mdx documentation
for consistency.

Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com>

* docs(openshift): remove outdated two-stage install instructions

Update OpenShift documentation to reflect that single-stage installs now
work with cert-manager and BackendTLSPolicy. The Certificate resources
run as pre-install hooks (weight -30) before certgen (weight -20),
allowing the hook to poll for and find the issued certificates.

Removed the outdated two-step process:
1. helm install (cert-manager issues cert, hook logs warning)
2. helm upgrade (hook creates ConfigMap)

Replaced with current single-stage behavior:
- Certificate resources created as pre-install hooks
- certgen hook polls for up to 90 seconds (configurable)
- Single helm install succeeds in most cases
- Fails fast by default if timeout reached

This brings openshift.mdx in line with the already-updated ingress.mdx
documentation.

Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com>

* fix(helm): make pkiInitJob.timeoutSeconds the actual polling duration

The timeout value now represents the actual polling time that users
experience when waiting for cert-manager to issue certificates.
The Job activeDeadlineSeconds is set to (timeoutSeconds + 30) to
allow buffer time for ConfigMap creation and cleanup.

Previously, the hook polled for (timeoutSeconds - 30) seconds, which
was confusing when users set timeoutSeconds=180 and only got 150
seconds of actual polling.

Updated documentation in values.yaml, ingress.mdx, and openshift.mdx
to reflect the clearer behavior.

Signed-off-by: Brandon Squizzato <bsquizza@nvidia.com>

* docs(helm): update values.yaml and README with correct polling duration

Updated the caCertificateConfigMapName description to reflect that the
hook polls for exactly pkiInitJob.timeoutSeconds seconds, not
(timeoutSeconds - 30) seconds.

Regenerated README.md with helm-docs.

Signed-off-by: Brandon Squizzato <bsquizza@nvidia.com>

* docs(kubernetes): add OIDC configuration to helm install and CLI examples

Updated all helm install and openshell gateway add examples in ingress.mdx
and openshift.mdx to include OIDC issuer and audience configuration.

Examples now use concrete placeholder values:
- OIDC issuer: https://keycloak.example.com/realms/openshell
- OIDC audience: openshell-cli
- Hostname: gateway.example.com
- ClusterIssuer: letsencrypt-prod

This makes it clearer how to configure OIDC authentication, which is
required when using BackendTLSPolicy or HTTPS termination since the
Gateway proxy cannot present client certificates to the backend.

Signed-off-by: Brandon Squizzato <bsquizza@nvidia.com>

* docs(kubernetes): explicitly list OIDC client ID in gateway add examples

Added --oidc-client-id openshell-cli to all openshell gateway add
commands in ingress.mdx and openshift.mdx, making the default client
ID explicit in the examples even though it's the CLI default.

This improves clarity and helps users understand the complete OIDC
configuration needed for gateway registration.

Signed-off-by: Brandon Squizzato <bsquizza@nvidia.com>

* fix(helm): address PR review feedback for BackendTLSPolicy

- Read backend CA from the authoritative server Secret instead of the
  in-memory PKI bundle so enabling BackendTLSPolicy on an existing
  release uses the CA that actually signed the server certificate.
- Reconcile the backend CA ConfigMap on every hook run (compare and
  update) instead of skipping when it already exists, so CA rotations
  propagate automatically.
- Remove hook annotations from cert-manager Issuer/Certificate resources
  so they remain regular release objects managed by Helm lifecycle. Split
  the cert-manager backend CA ConfigMap creation into a separate
  post-install/post-upgrade hook Job that polls after cert-manager
  Certificate resources are applied.
- Update architecture/gateway.md, docs/reference/gateway-config.mdx,
  debug-openshell-cluster skill, and helm-dev-environment skill with
  BackendTLSPolicy, backend CA ConfigMap, enableMtls, and timeout
  documentation.

Signed-off-by: Brandon Squizzato <bsquizza@nvidia.com>
Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>

* fix(docs): resolve markdown lint errors in helm README and kubernetes docs

Escape inline HTML angle brackets in README.md template placeholders,
remove trailing spaces, and add blank lines around fenced code blocks
in numbered lists.

Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>

* Update docs

* fix(helm): escape inline HTML in values.yaml descriptions and sync mise lockfile

Wrap `<fullname>` and `<namespace>` template placeholders in backticks
so markdownlint does not flag them as inline HTML (MD033). Regenerate
mise.lock to match current mise.toml after rebase onto main.

Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>

* Run 'mise lock'

* fix: align mise.lock with CI mise version output

The lockfile was regenerated locally with mise 2026.8.10 which resolves
uv Linux artifacts to gnu variants and adds provenance_verified fields,
but CI uses v2026.4.25 which produces musl variants without those fields.

Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>

* fix(docs): correct cert-manager hook ordering and clientCaSecretName comment

Update ingress.mdx and openshift.mdx to describe Certificate resources
as regular release objects with a post-install/post-upgrade Job, matching
the current implementation and architecture/gateway.md.

Fix values.yaml clientCaSecretName comment to state that "" disables
client certificate verification, matching the helper and access-control
docs.

Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>

---------

Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>
Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com>
Signed-off-by: Brandon Squizzato <bsquizza@nvidia.com>
2026-09-14 17:39:15 +00:00
Max Dubrinsky 8d19308c08 fix(bootstrap): emit RFC 5280 extensions on generated gateway PKI (#3286)
generate_pki minted a CA with no key usage and server and client leaves
with no Authority Key Identifier. RFC 5280 requires both, and verifiers
that enforce it reject the chain: OpenSSL X509_STRICT fails with
"Missing Authority Key Identifier", and Python 3.13 turned that flag on
by default in ssl.create_default_context(). rustls and BoringSSL do not
enforce it, so gRPC clients kept working while an HTTPS client built on
Python 3.13 (for example a platform proxying to an exposed sandbox
service) could not complete a handshake with a pkiInitJob-provisioned
gateway at all. cert-manager PKI was unaffected.

Set keyCertSign and cRLSign on the CA and use_authority_key_identifier
on both leaves, matching what the sandbox L7 CA already does. Add a
test that parses the bundle and asserts the extensions, including that
each leaf AKI matches the CA SKI.

Verified: openssl verify -x509_strict accepts both leaves, and a strict
Python 3.13 client completes an mTLS handshake against a server using
the new bundle where the previous bundle reproduces the failure.

Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com>
2026-09-14 16:49:44 +00:00
krishicks 5b9daab935 fix(ci): restore mise run ci on macOS (#3294)
- Replace BSD-incompatible in-place sed calls with portable temp-file rewrites.
- Remove test-only shell interception and capture generated gateway config
  directly.
- Allow parity tests to use supplied supervisor binaries without resolving a
  Linux target.
- Normalize temporary-directory paths and use portable RPM config installation.
- Set a valid setuptools-scm version for Python protobuf generation in Jujutsu
  checkouts.

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-09-12 00:17:15 +00:00
Piotr Mlocek 5b57f0d154 fix(ci): restore Windows test portability (#3288)
* fix(ci): restore Windows test portability

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(tasks): skip Unix lockfile check on Windows

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(tasks): use buf shim on Windows

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

---------

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-09-11 21:40:08 +00:00
krishicks bcf4558cfc fix(mise): run mise lock --platform linux-x64 (#3291)
Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-09-11 20:38:01 +00:00
Polite_realism 00f02127ad test(supervisor-network): show response body on ssrf_denied assertion failure (#3290)
forward_handler_preserves_ssrf_response_and_denial_stage's ssrf_denied
check was the only assertion in the test without a failure message,
unlike its sibling assertions. This test has failed intermittently in
CI with no way to see what response was actually returned, since the
message is where that gets surfaced.

Signed-off-by: politerealism <burdcat17@gmail.com>
2026-09-11 19:29:04 +00:00