* docs(mxc): add missing demo examples for runbook and mTLS scenario
Add three files present in the internal mirror but absent from the
Windows branch:
- mxc-demo-runbook.md: operator runbook for the MXC demo kit
- README-mtls.txt: usage notes for the mTLS scenario
- run-mtls-test.ps1: PowerShell test runner for the mTLS scenario
These are required by the package-demo.ps1 packager script and are
referenced by the mxc-kit documentation.
No issue required: mechanical sync of missing demo assets.
Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
* fix(mxc): align demo workflows with current contracts
Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
---------
Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Co-authored-by: Shailendra Singh <shailendras@nvidia.com>
Redact injected environment values and the per-sandbox proxy password in
captured MXC output and decoded relay launch-failure diagnostics before
logging or publishing sandbox failure status.
Match the original text and redact the union of overlapping occurrences.
Preserve control-channel payloads and avoid allocating for unmatched text.
Add regression coverage and document exact-match and length limits.
Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
* ci(windows): run MXC host probe, WebSocket agent, and OpenClaw forward examples checks
Add separate hosted CI tasks for the MXC host probe, WebSocket agent,
and OpenClaw forward examples. Use mock workloads to verify gateway,
CLI, driver, and sandbox lifecycle wiring without requiring wxc-exec.
Document that mock passes do not validate forwarding or MXC enforcement.
* fix(tests): enhance environment isolation for WebSocket and OpenClaw mock tests
* fix(mxc): enhance OpenClaw mock validation to require 'Ready' sandbox state
* fix(mxc): restore native process helper in WebSocket example
Restore Invoke-NativeCaptured for the PowerShell argument regression test and delegate CLI execution through it while preserving explicit gateway endpoint selection.
Signed-off-by: Akber Raza <akberr@nvidia.com>
---------
Signed-off-by: Akber Raza <akberr@nvidia.com>
Document standalone MXC executable placement and condition system-drive ACL preparation on the AppContainer + DACL tier and probe recommendation. Explain the persistent metadata-only grant and link upstream verification and rollback guidance.
NVBug: 6842834
Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>
Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
* fix(core): enforce owner-only Windows ACLs on sensitive files and dirs
set_dir_owner_only/set_file_owner_only were unconditional no-ops on
Windows, so the CLI's mTLS client private key, OIDC/edge tokens, cached
SSH keys, and the gateway's key-encryption key relied entirely on
inherited NTFS ACLs with no OpenShell-applied restriction. Apply an
owner-only DACL via SetEntriesInAclW/SetNamedSecurityInfoW with
PROTECTED_DACL_SECURITY_INFORMATION to strip inherited ACEs, matching
the 0700/0600 guarantee already provided on Unix. is_file_permissions_too_open
now also works on Windows instead of being Unix-only, closing the
detection gap alongside the prevention gap.
Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
(cherry picked from commit 71560e947f85819efbcddf70ddda94befab62b0b)
* fix(core): treat a NULL DACL as too open in is_file_permissions_too_open
has_foreign_trustee conflated a NULL DACL with an unreadable/invalid
ACL and returned Some(false) (not too open) for both. Per the Win32
contract, a NULL DACL means the object grants full access to everyone
-- the most permissive state possible -- so it must be flagged as too
open. Split the null and invalid-ACL branches: null now returns
Some(true), invalid ACL keeps the existing unreadable-ACL fallback
(None, which the caller maps to false via unwrap_or). Adds a
regression test that constructs a real NULL DACL via a
SetNamedSecurityInfoW helper confined to the windows_acl module,
consistent with the existing unsafe-FFI confinement in that module.
Found by CodeRabbit review on MR !113.
Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
(cherry picked from commit 46e635a4ef1d6937cdb088f46aa85baee3d6ad28)
* fix(core): close three false-negative gaps in the Windows ACL audit
restrict_to_current_user() updated only the DACL, leaving a foreign
owner's implicit WRITE_DAC right intact -- they could later replace
the DACL we just set. Query OWNER_SECURITY_INFORMATION and take
ownership in the same SetNamedSecurityInfoW call; if the caller can't
(a genuinely foreign-owned object), the call now fails instead of
silently leaving the object insecure.
is_file_permissions_too_open() mapped every Win32 inspection failure
(missing READ_CONTROL, an invalid ACL, a token-query failure) to
"not too open" via unwrap_or(false). Fail closed instead: an
inspection failure is a security false-negative risk, not a green
light.
has_foreign_trustee()'s ACE loop only recognized plain
ACCESS_ALLOWED_ACE_TYPE and treated every other type as non-granting.
Windows also defines access-allowed object, callback, and
callback-object ACE variants that can grant rights to a foreign
trustee; this audit doesn't parse their wider layouts, so their mere
presence is now conservatively flagged as too open instead of
silently skipped.
Also updates architecture/gateway.md, which still described the
SQLite file-tightening behavior only in terms of Unix mode 0o600, to
distinguish it from the owner-only DACL behavior on Windows.
Addresses review comments on PR #3495.
Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
* fix(core): conditional owner claim and audit owner in Windows ACL helpers
restrict_to_current_user: query the current owner before calling
SetNamedSecurityInfoW. Include OWNER_SECURITY_INFORMATION only when the
path has a foreign owner -- requesting it unconditionally fails with
ACCESS_DENIED (0x80070005) on standard credentials even when the current
user is already the owner, because WRITE_OWNER is not implied by object
ownership. A foreign-owned path still triggers an ownership claim and
fails hard if the claim is denied, preserving the security contract.
has_foreign_trustee: request OWNER_SECURITY_INFORMATION alongside
DACL_SECURITY_INFORMATION and reject paths with a foreign owner
immediately, before inspecting the DACL. A foreign owner has implicit
WRITE_DAC rights and can replace any DACL we set, so a clean DACL is not
sufficient evidence of safety on a foreign-owned object.
architecture/gateway.md: clarify that the Windows path-hardening behavior
sets mode 0o600 on Unix and applies a protected owner-only DACL on
Windows, with conditional ownership claim and fail-hard semantics for
foreign-owned objects.
Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
---------
Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
* fix(ocsf): attribute MXC proxy events to sandboxes
NVBug 6783086
Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>
* fix(network): scope Windows egress by socket owner
Resolve each accepted Windows proxy connection to its unique owning PID and executable before evaluating binary-scoped network policy. Fail closed when ownership or process identity cannot be established, and cover allowed and undeclared child processes with a real MXC regression.
* fix(network): preserve sandbox context with socket owners
Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>
(cherry picked from commit 32ea316deb6c83c157fdf2de4a2cdc3a7d2da1e3)
* fix(network): harden Windows process identity
Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>
---------
Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>
* fix(mxc): honor the generic sandbox create -- <COMMAND> syntax
sandbox_config only ever read the command from the MXC-specific
driver_config (--driver-config-json), never DriverSandboxSpec.command
-- the portable field the CLI's documented, driver-agnostic
`sandbox create -- <COMMAND>` syntax actually populates, and the same
field every other compute driver honors. A caller following that
syntax got "driver_config.command must contain a non-empty
executable", a message that reads as if no command was supplied at
all.
Add spec.command as a fallback source when no driver_config is
present, and reword the empty-command error to name both ways to
supply one.
NVBug 6782884
Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
* docs(mxc): document generic sandbox commands
Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
---------
Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Co-authored-by: Shailendra Singh <shailendras@nvidia.com>
* fix(network): normalize Windows policy binary paths
Match Windows executable identities using a stable case-insensitive, separator-normalized representation across policy data, L4 input, and L7 relay evaluation. Preserve exact matching on other platforms and keep the original path for hashing and filesystem access.
NVBug 6782969
Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>
* fix(network): harden Windows binary matching
Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>
* fix(ci): scope Windows relay test imports
* fix(network): harden Windows binary path matching
---------
Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>
* fix(server): keep Error-phase sandbox records through the prune sweep
The periodic store-vs-backend reconciliation sweep deleted any
persisted sandbox not present in the driver's live backend snapshot,
except for Completed and failed-main-process phases. A driver whose
registry is in-process-only and never rehydrates after a restart (no
persistence of its own) reports every previously-known sandbox as
missing on the very first sweep after startup -- including ones
already correctly, terminally marked Error by earlier crash detection
-- so the sweep silently deleted them shortly after gateway restart,
racing any client (GetSandbox/ListSandboxes/DeleteSandbox) working
with the same sandbox in that window.
Treat Error the same as the existing Completed exemption: it is
already a settled, informational terminal state with no live compute
resource to reclaim, so keep the durable record instead of deleting
it.
Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
(cherry picked from commit 721a1659a822a72e76e3c0dffc6847f17129a3fc)
* fix(server): narrow the prune sweep's Error-phase exemption
The blanket phase == SandboxPhase::Error exemption changed established
missing-backend cleanup for every compute driver, not only the MXC
restart race the PR intended to fix. It also matched
BackendResourceMissing (set by gateway-startup recovery when a
previously-known sandbox's backend resource is already gone),
StartFailed (startup recovery's driver-error case), and
ComputeResourceMissing (this same sweep's own first-pass Error
transition for a Stopping/Stopped/Starting sandbox). All three mark
exactly the orphaned resources this sweep exists to reclaim across
Docker, Podman, VM, Kubernetes, and extension drivers -- exempting
them left orphaned names and gateway-owned records in place
indefinitely and skipped the idempotent driver cleanup for
volumes/secrets until a user explicitly deleted the sandbox.
Add is_missing_compute_resource_reason to inspect the sandbox's Ready
condition and narrow the exemption to a settled Error record only: one
whose reason isn't one of those three. A crashed main process or any
other non-resource failure keeps the exemption (no live resource ever
expected again); a resource-missing reason keeps flowing through the
normal delete-and-cleanup path exactly as before this PR.
Adds regression coverage for BackendResourceMissing and
ComputeResourceMissing confirming they are still pruned with driver
cleanup invoked, and documents the settled-vs-missing-resource
retention distinction in architecture/compute-runtimes.md.
Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
---------
Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
* feat(mxc): warn when the ETW consumer receives zero provider events
EnableTraceEx2 succeeding only proves the request to enable the
Sandboxing provider succeeded, not that the provider exists on this
host/build or will ever fire. A provider-identity mismatch or a
non-firing provider left the ETW->OCSF audit trail silently empty
across real, successful sandbox lifecycles, with no error, warning,
or diagnostic anywhere.
Track raw provider-matched events received per session and add a
zero-events watchdog on the consumer thread: once real sandbox
activity has happened (register_launch called at least once) and a
grace period elapses with zero events matched, log a warning and
emit a Detection Finding [2004] naming the gap. Gated on actual
activity (not just session uptime) so an idle gateway with
etw_audit=true and no sandboxes created never warns.
Also exposes EtwSession::events_received() alongside the existing
is_capture_alive(), so a status/diagnostics surface can query capture
health directly, not just infer it from tracing output.
The watchdog's decision logic is extracted into a pure function
(should_warn_zero_events) so it's unit-testable without a real ETW
session or elevation.
Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
(cherry picked from commit 4fcfa716a808eaa91c6476e9c69366f91702b1fa)
* fix(mxc): repair build/lint breaks left by the zero-events watchdog commit
warn_zero_events_received() referenced a nonexistent SESSION_NAME
constant (only SESSION_NAME_PREFIX exists), so the crate failed to
compile. Separately, the consumer thread's decode-or-log match on
Option<DecodedEtwEvent> tripped clippy::single_match_else under -D
warnings. Neither issue is specific to a platform or toolchain
version -- both reproduce on a clean checkout of this branch's tip.
Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
* fix(mxc): attribute ETW Sandboxing-provider events via identity/CV, not PID
The Sandboxing provider's events are logged under two PIDs that are
never the driver's own wxc-exec PID: a short-lived launcher PID
(wxc-exec.exe exits within seconds even while the sandboxed workload
keeps running) and a shared, constant PID hosted by a system-wide OS
broker service across unrelated sandboxes. Anchoring attribution on
`register_launch`'s wxc_pid therefore left every event unattributable,
so the OCSF audit trail stayed empty despite the provider firing
correctly (confirmed against a raw logman/tracerpt capture running
alongside this consumer).
wxc-exec never reports its OS-generated `identity`/`__TlgCV__` back to
the driver, so there is nothing to pre-seed `by_identity` with at
registration time the way `by_pid` is pre-seeded. Track launches still
awaiting their first identity/CV in a `pending_launches` queue instead,
and bind opportunistically in `resolve()`: when an event carries a
never-seen identity/CV and has no `by_pid` registration at all (the
real-world shape of these events) and exactly one launch is pending,
it can only be that launch's burst. Zero or multiple pending launches
stay ambiguous and fall through to the existing unresolved-event
buffer/TTL path rather than guess -- misattributing an audit event to
the wrong sandbox_id is worse than dropping it. A PID registration
that does exist (even generation-mismatched) is treated as positive
evidence of an existing PID-reuse race and takes precedence over the
opportunistic path, preserving the existing generation-key guarantees.
Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
* fix(mxc): surface dropped-unattributed ETW events above debug level
An event that ages out of the unresolved-event buffer unattributed is
a permanent audit-trail gap: the OS action it represents will never
appear in the OCSF log, and nothing retries it afterward. That was
only visible at --log-level debug, so an operator running with the
default level would never see it. Promote it to warn, matching the
severity already used for the zero-events watchdog's own gap warning.
Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
* fix(mxc): keep ETW attribution fail-closed, rate-limit drop warnings
The prior opportunistic "single pending launch" fallback attributed
any fresh identity/CV from an unregistered or system-broker PID to
the sole pending OpenShell launch. Queue cardinality is not
attribution evidence: unrelated, non-OpenShell AppContainer or UAC
activity shares this same OS Sandboxing provider, so an event
arriving in the five-second window could cross-link an unrelated
identity and emit subsequent events under the wrong sandbox_id,
corrupting the audit trail. Revert it -- attribution requires the
driver-owned wxc-exec PID and its kernel process start key to both
match, exactly as before. Records without that generation-backed
evidence remain unattributed rather than guessed.
The dropped-unattributed warning also moved from one line per record
to a rate-limited, aggregated warning: unattributed drops are
expected, ordinary system-wide activity, not a rare condition, so
warning per record could flood operator logs during a burst of
unrelated AppContainer/UAC activity. UnattributedDropReporter mirrors
the existing OverloadReporter pattern -- warn immediately on the
first drop, then coalesce to one warning with a running count every
30 seconds while drops continue.
Updates the MXC observability documentation to match: the
unattributed-record sentence now explains why fail-closed is
deliberate, documents the coalesced warning cadence, and documents
the mxc-etw-zero-events OCSF Detection Finding [2004], which was not
mentioned anywhere before.
Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
---------
Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
* fix(mxc): reject a non-absolute wxc_exec_path at gateway startup
wxc_exec_path is the binary that builds every sandbox, but nothing
validated it before spawning: a relative value (including the shipped
default, a bare "wxc-exec.exe") let PATH-lookup or working-directory-
relative resolution execute a decoy binary with the gateway's identity
instead of the approved wxc-exec, turning the containment mechanism
itself into an arbitrary-code-execution primitive.
Add MxcComputeConfig::validate_configuration, wired into the existing
(previously no-op) compute-driver config preflight, rejecting an
empty or non-absolute wxc_exec_path with a clear diagnostic. Change
the default from the relative "wxc-exec.exe" to an empty string so
the field must be explicitly configured -- no usable-but-insecure
fallback survives. Update the architecture doc's stale "else PATH"
discovery claim to match.
Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
(cherry picked from commit d4192a0072f8f0ca0a4035f10cae07889c2b578f)
* test(mxc): cover wxc_exec_path enforcement at the gateway boundary
MxcComputeConfig::validate_configuration already had unit coverage,
but that only proves the validation function itself is correct -- it
says nothing about whether MxcFactory::validate_config (src/lib.rs)
still calls it. Before this fix, that factory method discarded the
parsed config entirely, so a regression back to that no-op shape would
leave every unit test passing while a relative wxc_exec_path again
reached gateway startup.
Add an integration test that spawns the actual compiled
openshell-gateway binary through its config preflight subcommand,
exercising the real chain: CLI parsing, TOML loading, driver
selection, MxcFactory::validate_config, and
MxcComputeConfig::validate_configuration. Assertions check only
pass/fail, not message content: run_effective_config_preflight
replaces any validation failure with a generic message whenever a
config file is used, to keep file-sourced values out of preflight
diagnostics -- pre-existing, deliberate, and covered by its own tests.
Also document the new required-and-absolute wxc_exec_path constraint
in the gateway config reference and the driver README, which
previously only showed example values without stating the
requirement.
Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
---------
Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
* fix(mxc): stage egress-proxy CA files regardless of env tier
stage_tls_ca_files was only invoked when pc_minimal_env was set, but a
curated ProcessContainer can't read the host proxy's private temp
folder under any env tier. With the default pc_minimal_env=false, the
CA env vars pointed at a path the sandboxed process couldn't read at
all, breaking TLS validation for allow-listed HTTPS requests through
the egress proxy. Extract the staging decision into
resolve_agent_proxy_ca_paths, which takes no env-tier argument, so the
gap can't silently regress; existing tests never exercised this path
since mocked invokers skip host-proxy startup entirely.
Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
(cherry picked from commit d0346af96aef9a93aca26a7e1825d8678a25567d)
* fix(mxc): complete proxy CA staging isolation (#3535)
* test(mxc): provide workload dir for CA staging
Signed-off-by: Prekshi Vyas <prekshivyas@nvidia.com>
* fix(mxc): isolate staged proxy CA files
Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>
---------
Signed-off-by: Prekshi Vyas <prekshivyas@nvidia.com>
Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>
---------
Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
Signed-off-by: Prekshi Vyas <prekshivyas@nvidia.com>
Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>
Co-authored-by: Prekshi Vyas <34834085+prekshivyas@users.noreply.github.com>
Adds real-MXC regression coverage for ProcessContainer token isolation, including an unsandboxed SCM positive control, and documents the AppContainer authorization model.
Validate raw OPA object and array containers before normalization and
access-preset expansion can skip malformed values. Return one fixed
structural error without embedding authored policy data.
Preserve versionless and runtime-only OPA data and existing semantic
validation. Cover initial string/file loading, middleware callback order,
valid deny-rule enforcement, and rejected reload state and generation.
Document the loader contract and its engine-local rejection behavior.
Signed-off-by: Shiju <shiju@nvidia.com>
Previously, metadata events omitted both HTTP request and response objects,
early proxy rejections used HTTP Activity without request context, and
unsupported-scheme events did not expose enough safe HTTP context to satisfy
the OCSF 1.8 schema.
Now, metadata events include a method-only request and their actual HTTP
response codes without recording the metadata URL. Unsupported-scheme events
also include a method-only request plus the generated 400 response. Authority
mismatches and credential-resolution denials use HTTP Activity with their
generated 403 or 500 responses, and HTTP activity IDs are derived from the
request method.
Additionally, HttpActivityBuilder now enforces the OCSF request-or-response
constraint at compile time.
Signed-off-by: Kris Hicks <khicks@nvidia.com>
* feat(api): expose structured gateway errors across SDKs
Refs #3051. Add standard validation, conflict, and retry details; preserve raw transport status in Rust, Go, TypeScript, and Python; document status and recovery guidance.
This is the structured-error foundation only. Mutation result shapes, allow_missing, durable request deduplication, and exec retry semantics remain follow-up work.
Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
* fix(python): preserve wrapped RPC cleanup handling
Inspect the original gRPC call when handling missing sandboxes during deletion waits and managed cleanup. Add intercepted cleanup regressions and clarify the error-wrapper migration contract.
Addresses the cleanup review on #3313; part of #3051.
Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
---------
Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
Show configured tool server addresses and their last observed connection
results together in sandbox status. Keep sandbox lifecycle readiness
separate so an external connection failure does not mark the sandbox unready.
Expose direct endpoint records through the CLI and SDKs, with plain-language
failure explanations and gateway acceptance times. Keep observation tracking,
runtime reporting, and gateway validation in dedicated endpoint status modules.
Preserve bounded reporting, request attribution, retry ordering, and
configuration and supervisor authority checks. Clear obsolete observations
while retaining the configured addresses, and document the distinction
between an observed HTTP response, current availability, and tool success.
Signed-off-by: Shiju <shiju@nvidia.com>
* Implement Windows host proxy integration and update dependencies for OpenShell
* Update README and gateway config to clarify egress proxy address handling and allocation
* Refactor ProxyIdentityMode to return Result for static_binary and add tests for binary path and SHA256 hash
* Enhance platform_hosts_path for Windows to use SystemRoot and improve error handling for hosts file reading
* Refactor FileFingerprint to use Option for mtime and ctime, simplifying metadata handling
* Add conditional compilation for Windows host module
* add unit tests for OPA policy evaluation and identity handling
* remove openshell-supervisor-network from unsupported driver package test exclusion list
* feat(mxc): enable host proxy TLS state generation
Generate per-sandbox TLS state for the MXC host proxy so HTTPS L7 enforcement can use the same MITM path as Linux. Grant generated CA material to the MXC process and inject standard trust env vars, while matching Linux behavior by disabling TLS termination on CA setup failure and relying on proxy fail-closed handling.
* fix(docs): remove outdated notes on governed egress from docs
* fix(tests): update TLS environment variable paths to use temporary directory
* fix(examples): make run-mxc-e2e harness correct and orphan-free
The MXC e2e harness never actually exercised the fs scenarios: it started
the gateway once and patched agent_command per scenario AFTERWARDS, so the
running gateway kept launching the default demo agent (not shipped in the
kit) and every fs scenario failed with CreateProcessW error:2. It also
scored on the `sandbox create` exit code (non-zero due to the harmless
interactive attach), wrote sandbox records to the persistent gateway DB
(leaving orphans that collided on later runs), and its deny scenarios never
proved denial.
Changes:
- Start a FRESH gateway per scenario so each scenario's agent_command is
actually loaded (root cause of CreateProcessW error:2).
- Score by on-disk artifact / expected outcome, not `sandbox create` exit.
- Real deny assertions: a control write to a granted path must succeed
(proves the agent ran) while the denied write must be absent. fs-empty
probes an ungranted out-of-share path (share_dir is mapped rw by design).
- Run the gateway on an ephemeral in-memory DB (sqlite::memory:) so the
harness never writes to the persistent store and cannot leave orphan
sandbox records; also use unique per-run sandbox names + pre-delete.
- Fix the process_container probe: use a real cwd + absolute cmd.exe
(canonical wxc-exec does not expand %TEMP% -> 0x8007010B).
- Fix summary counts (@() so a single FAIL is counted and exit is non-zero).
Verified PASS=4 FAIL=0 on 7F203-MXC-003 (no BaseContainer velocity keys)
using a canonical wxc-exec build (AppContainer fallback).
Signed-off-by: Akber Raza <akberr@nvidia.com>
* fix(e2e): probe timeout is milliseconds (10ms->30000ms)
MXC process.timeout is wall-clock ms (wire.rs). The 10 value meant 10ms,
which the base-container tier (7F203-MXC-001/.181) enforced strictly and
timed the probe out. AppContainer path (.18/-003) happened to slip under
it. Bump to 30000ms so the process_container preflight probe is reliable
across both tiers.
Signed-off-by: Akber Raza <akberr@nvidia.com>
* fix(mxc): use native paths in real runtime probes
Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>
* fix(mxc): make processcontainer work with mxc-latest-released wxc-exec
Three fixes to support the release wxc-exec binary (BaseContainer dispatcher)
in addition to mxc-fixes-env-vars:
1. Seed process env from host (driver.rs)
ProcessContainer starts with a completely blank environment -- no PATH,
SystemRoot, or anything. Seed the process env from the gateway host
environment so the agent binary can locate DLLs and run. Skip internal
Windows drive-letter variables (keys starting with '=') which cause
CreateProcessW to return ERROR_ENVVAR_NOT_FOUND. User agent_env entries
and TLS CA vars are applied as overrides on top of the host env.
2. Remove TLS readonly_paths grant (driver.rs)
The release wxc-exec (BaseContainer dispatcher) requires write-DAC
permission on every path in readonly_paths to set up AppContainer ACLs.
Adding the proxy's temp TLS directory caused a DACL error and exit -1.
The CA cert paths remain available to the agent via TLS env vars.
3. Remove allowedHosts from network JSON (mxc.rs)
The release wxc-exec rejects network.allowedHosts / network.blockedHosts
on Windows with "not yet supported". Removed the loopback exemption
attempt (127.0.0.1, ::1, localhost) from the network section.
Intra-container loopback works natively in the release binary without
it -- the spawner can connect to the server at 127.0.0.1:22000 directly.
Additional changes:
- mxc-ws-agent.rs: add relay-debug.txt error capture and relay-ready.txt
marker for reliable timing of host client connections.
- mxc-ws-gateway.toml: debug = true for JSON config dump during diagnosis.
- run-ws-agent-test.ps1: default port changed to 17670 (gateway default);
relay-ready.txt polling before ws-echo to avoid connecting before the
spawner has established the proxy bridge.
Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>
* fix(e2e): address CodeRabbit review on run-mxc-e2e.ps1 (MR !46)
Four robustness/correctness fixes from CodeRabbit:
1. Start-Gw: kill the spawned gateway before the "did not start within 30s"
throw. If the process is alive but never binds the port, $gw is not yet
assigned in the caller, so the finally block cannot reap it -> orphan
gateway holding the port for the next run.
2. create-fail scoring: a non-zero `sandbox create` exit alone is not proof
of a policy rejection (gateway-registration/transport/fixture errors also
exit non-zero and would false-pass). PASS now requires a genuine
rejection signal (network / invalid_argument / network_policies) AND that
it is not an infrastructure failure; other non-zero exits go to FAIL with
output captured.
3. deny scenarios (ControlTarget path): snapshot the deny target AFTER
Wait-File lands the control artifact, so a late denied write (enforcement
regression racing the control write) can no longer be recorded as PASS.
4. -KeepRunning: break out of the scenario loop after the first scenario so
a later scenario does not start a second gateway on the same port
(previously a reliable port collision instead of a usable debug mode).
Re-verified PASS=4 FAIL=0 on both boxes (7F203-MXC-001 base-container and
7F203-MXC-003 AppContainer fallback); network-policy-rejected correctly
scores as "policy rejection".
Signed-off-by: Akber Raza <akberr@nvidia.com>
* feat(mxc-e2e): collect run-mxc-e2e output into a results bundle
Mirror the sibling run-*.ps1 scripts by collecting every run's logs into a
timestamped results-e2e-<stamp>\ folder and zipping it. The bundle contains the
console transcript, per-scenario gateway stdout/stderr, the exact TOML rendered
for each scenario, the policy fixture used, and a summary.txt with the verdict
table.
Per-scenario gateway logs now land in gateway.<scenario>.log/.err.log inside the
bundle instead of a single fixed gateway.e2e.log in the script directory.
Wrap pre-flight, mode setup, scenario definitions, and the scenario loop in a
single try/catch/finally so the finally always writes the summary, stops the
transcript, and zips the bundle -- even on a pre-flight failure. The existing
per-scenario gateway-cleanup try/finally stays nested inside. All scenario
logic, scoring rules, and comments are preserved.
Signed-off-by: Prashant S Khodade <pkhodade@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>
* fix(mxc-e2e): address CodeRabbit review on run-mxc-e2e.ps1
- Require -Scenario when -KeepRunning: the loop breaks after the first
scenario, so a full-suite run would execute only one scenario yet still
report the suite as PASS. Fail fast so a partial run can't be mislabeled
complete.
- Start-Transcript now runs inside the guarded try block with a
$transcriptStarted flag; Stop-Transcript is only called when it actually
started, so a Start-Transcript failure still yields the results bundle.
- Wrap the -Scenario filter in @() so a single exact match stays an array
(reliable .Count and a proper array for the scenario loop on PS 5.1).
Signed-off-by: Prashant S Khodade <pkhodade@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>
* fix(examples): pass gateway config via OPENSHELL_GATEWAY_CONFIG for spaced paths
Start-Process -ArgumentList does not quote array elements, so launching the
gateway with a bare --config <path> token split on any space in the install
path (e.g. C:\Users\First Last\...), and clap rejected the fragment with
'unrecognized subcommand'. Every MXC example launcher that started the gateway
hit this when the kit was unzipped under a path containing a space.
Pass the config path through the OPENSHELL_GATEWAY_CONFIG env var (which the
gateway already reads via clap) and drop the --config token. Env vars carry
spaces safely.
Affected: run-ocsf-audit, run-mxc-e2e, run-demo, run-inference-test,
run-ollama-test. run-mtls-test was not affected (its launch passes no config
path). Root-caused and fix-verified on 7F203-MXC-003 from a spaced path.
Signed-off-by: Akber Raza <akberr@nvidia.com>
* fix(run-mxc-e2e): improve scoring logic and enhance command execution handling
* fix(mxc): reconcile proxy support after rebase
Restore the proxy-enabled OCSF audit example removed by 13185f6e now that the host CONNECT proxy is present. Adapt the proxy lifecycle test to the target branch's DriverSandboxSpec policy delivery contract.
Signed-off-by: Akber Raza <akberr@nvidia.com>
* fix(mxc): grant sandbox access to proxy CA
- Share the per-sandbox public CA bundle with the AppContainer
- Add real wxc-exec HTTPS proxy coverage and document trust isolation
Signed-off-by: Akber Raza <akberr@nvidia.com>
* fix(mxc): reject unsupported network middleware
- Reject middleware-bearing MXC policies before sandbox lifecycle begins
- Guard host proxy startup and document the unsupported registry path
- Add mapper, lifecycle, and host proxy regression coverage
Signed-off-by: Akber Raza <akberr@nvidia.com>
* fix(mxc): reconcile host proxy with main
- remove obsolete inference routing from the host proxy adapter
- use the workspace AWS-LC provider in host-proxy tests
- adapt the forward-proxy test to ProxyIdentityMode
Signed-off-by: Akber Raza <akberr@nvidia.com>
* fix(build): switch to bundled Z3 for Windows MSVC builds
* fix(mxc): reconcile host proxy after rebase
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
---------
Signed-off-by: Akber Raza <akberr@nvidia.com>
Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
Signed-off-by: Prashant S Khodade <pkhodade@nvidia.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Jamie King <jamiek@nvidia.com>
Co-authored-by: Shailendra Singh <shailendras@nvidia.com>
Co-authored-by: Prashant Khodade <pkhodade@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
Implements https://github.com/NVIDIA/OpenShell/issues/3212
Running e2e:kubernetes with the Vault credential driver failed on
OpenShift in two ways: the OpenBao fixture pod was rejected by the
restricted-v2 SCC, and provider-creating tests hit HTTP 403 from
auth/kubernetes/login because OpenBao's Kubernetes auth was provisioned
only inside a single feature-gated test.
- Deploy OpenBao with the chart's OpenShift mode (global.openshift=true)
when OpenShift is detected, so the pod inherits a namespace-assigned,
SCC-compliant security context with no manual SCC grant. Hoist
OpenShift detection ahead of the credential-driver fixtures so the
flag is set before the fixture is deployed.
- Provision the OpenBao KV store, Kubernetes auth method, storage
policy, and gateway login role in the harness (deploy_vault_fixture),
making a Vault-backed gateway usable by the whole suite instead of
only the credential_drivers test. Remove the now-redundant
configure_vault_storage helper from the test.
- Harden openbao_exec so it tolerates only the idempotent "path is
already in use" error on reruns and fails fast with output on any
other error, instead of a blanket `|| true` that masked genuine
failures (e.g. an unresponsive pod) until a later cryptic write.
- Document the OpenShift Vault credential-store SCC and Kubernetes-auth
403 troubleshooting in the debug-openshell-cluster skill.
Signed-off-by: Jorge Garcia Oncins <jgarciao@redhat.com>
* feat(helm): add optional BackendTLSPolicy for e2e TLS
Add grpcRoute.backendTLSPolicy values to optionally create a
BackendTLSPolicy resource that enables end-to-end TLS between the
Gateway proxy and the OpenShell gateway pod. The Gateway proxy
terminates client-facing TLS and re-encrypts when connecting to the
backend, validating the pod's certificate against a user-supplied CA
ConfigMap.
This removes the requirement to set server.disableTls=true when using
HTTPS at the Gateway listener. Supported on OpenShift 4.22+ and other
platforms with BackendTLSPolicy support in the Gateway API
implementation.
Update OpenShift and ingress documentation with e2e TLS instructions.
Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>
* feat(helm,server): auto-create backend CA ConfigMap in certgen hook
Extend the generate-certs command with --backend-ca-configmap-name and
--backend-ca-source-secret flags. When BackendTLSPolicy is enabled, the
certgen pre-install hook creates the CA ConfigMap automatically:
- pkiInitJob mode (default): uses the CA from the generated PKI bundle.
Fully automatic on first install.
- cert-manager mode: reads ca.crt from the server TLS Secret. On first
install the Secret does not exist yet (cert-manager reconciles after
templates are applied), so the ConfigMap is created on the first helm
upgrade. Logs a warning on the initial skip.
The caCertificateConfigMapName value now defaults to <fullname>-backend-ca
when empty, so users only need to set backendTLSPolicy.enabled=true.
Update certgen RBAC to include configmaps get/create when the feature is
enabled. Add CLI arg parsing tests for the new flags.
Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>
* refactor(helm): add server.tls.enableMtls flag for mTLS control
Replace automatic mTLS disabling based on BackendTLSPolicy with an
explicit server.tls.enableMtls flag that defaults to true. The user is
now responsible for setting this to false when using BackendTLSPolicy,
as ingress proxies cannot present client certificates to backends.
Updated:
- values.yaml: Added server.tls.enableMtls (default true)
- gateway-config.yaml: Check enableMtls instead of backendTLSPolicy
- _gateway-workload.tpl: Check enableMtls for client CA mount
- Tests: Updated to use enableMtls flag
- Docs: Added enableMtls=false to BackendTLSPolicy examples
- README: Document new flag and BackendTLSPolicy requirement
Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>
* docs(helm): clarify cert-manager backend CA ConfigMap workflow
Update documentation to explain the two-step install process required
when using cert-manager with BackendTLSPolicy:
1. helm install - cert-manager issues the server certificate, but the
certgen hook can't create the backend CA ConfigMap yet (cert-manager
reconciles after templates are applied)
2. helm upgrade - certgen hook reads the CA from the cert-manager-issued
certificate and creates the ConfigMap
Previously, the docs said "created on first upgrade" without explaining
why or that the feature won't work until then. The updated docs now:
- Explain the timing issue (cert-manager reconciles after chart install)
- Provide clear steps for the cert-manager workflow
- Note that pkiInitJob (default) creates it immediately on install
- Clarify that users must wait for the Certificate to be Ready before
running the second upgrade
Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>
* fix(docs): remove incorrect external hostname requirement for BackendTLSPolicy
BackendTLSPolicy validates the backend certificate against the service FQDN
(e.g., openshell.openshell.svc.cluster.local), not the external hostname.
The external hostname only needs to be on the Gateway listener certificate
for client-facing TLS.
The default certManager.serverDnsNames already includes all required service
FQDN variants, so no configuration is needed for BackendTLSPolicy to work.
Fixed incorrect documentation that claimed:
- "The server certificate SAN list must include the external hostname"
- Users need to "configure certManager.serverDnsNames with the external hostname"
Removed the unnecessary pkiInitJob.serverDnsNames override from the example
and clarified that:
- Gateway listener certificate needs the external hostname (for clients)
- Backend certificate needs the service FQDN (for Gateway proxy)
- The service FQDN is already in the defaults
Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>
* docs: clarify ACME with LetsEncrypt reference
Change all references from "ACME issuer" to "LetsEncrypt/ACME issuer"
to help users understand that LetsEncrypt is the most common ACME
provider and what ACME means in practice.
Updated:
- docs/kubernetes/managing-certificates.mdx
- docs/kubernetes/openshift.mdx
- deploy/helm/openshell/values.yaml
- deploy/helm/openshell/README.md
- deploy/helm/openshell/ci/values-openshift-route-cert-manager.yaml
Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>
* docs(openshift): restructure end-to-end TLS options and clarify Gateway hostname
Reorganize the OpenShift production deployment documentation:
1. Changed main section from "Production Deployments" to "Options for
end-to-end TLS" for better clarity
2. Renamed subsections for consistency and clarity:
- "End-to-end TLS using Gateway API and BackendTLSPolicy (OpenShift 4.22+)"
- "End-to-end TLS using pass-through Route (all OpenShift versions)"
3. Clarified that the Gateway hostname is typically a wildcard:
"typically a wildcard like *.openshell-ingress-gw.example.com"
4. Removed the recommendation to copy the cluster's wildcard certificate
from openshift-ingress namespace, as this is not a recommended
security best practice
These changes make it clearer that users have two end-to-end TLS options
and help them understand the typical naming pattern for Gateway hostnames.
Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>
* feat(helm): eliminate two-stage install for BackendTLSPolicy with cert-manager
When using BackendTLSPolicy with cert-manager, the certgen hook now polls
for up to 90 seconds waiting for cert-manager to issue the TLS certificate
before creating the backend CA ConfigMap. This eliminates the need for a
second `helm upgrade` in most cases.
The hook polls every 2 seconds with progress logging every 10 seconds.
If cert-manager takes longer than 90 seconds, the hook times out gracefully
and logs a warning, preserving the fallback to manual ConfigMap creation
or a second upgrade.
The Job's activeDeadlineSeconds is 120s, so the 90s timeout leaves 30s
margin for ConfigMap creation and hook completion.
Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com>
* feat(helm): add configurable timeout for certgen hook
Add `pkiInitJob.timeoutSeconds` Helm value (default 120) to control how long
the certgen hook Job can run. When using cert-manager with BackendTLSPolicy,
the hook polls for (timeoutSeconds - 30) seconds to leave margin for ConfigMap
creation and cleanup.
This allows users to increase the timeout for environments where cert-manager
takes longer than 90 seconds to issue certificates, without requiring code
changes.
Example usage:
```yaml
pkiInitJob:
timeoutSeconds: 180 # Hook polls for 150 seconds
```
Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com>
* docs(helm): document configurable certgen timeout
Update documentation to mention the pkiInitJob.timeoutSeconds value and
how it affects the cert-manager polling behavior when using BackendTLSPolicy.
Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com>
* feat(helm): add configurable failure behavior for certgen timeout
Add `pkiInitJob.failOnTimeout` Helm value (default false) to control whether
the certgen hook fails or succeeds when cert-manager does not issue a
certificate within the polling timeout.
When false (default), the hook succeeds with a warning and users can run
`helm upgrade` after cert-manager issues the certificate to create the
backend CA ConfigMap. This provides backwards-compatible behavior.
When true, the hook fails immediately if the timeout is reached, providing
clear feedback that BackendTLSPolicy is non-functional. This is useful for
strict validation requirements where incomplete installs should fail fast.
Example usage:
```yaml
pkiInitJob:
timeoutSeconds: 180
failOnTimeout: true # Fail install if cert-manager takes >150s
```
Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com>
* feat(helm): change failOnTimeout default to true and add troubleshooting docs
Change `pkiInitJob.failOnTimeout` default from false to true to provide
immediate feedback when cert-manager does not issue certificates within
the polling timeout. This prevents silent failures where BackendTLSPolicy
is non-functional but the install appears to succeed.
Add comprehensive troubleshooting section to docs/kubernetes/ingress.mdx
documenting the specific error "TLS error: Secret is not supplied by SDS"
that occurs when the backend CA ConfigMap is missing, with step-by-step
resolution instructions.
Updated comments in values.yaml to clearly document the default behavior
and explain when administrators might see connectivity errors if they
override the default to failOnTimeout=false.
BREAKING CHANGE: pkiInitJob.failOnTimeout now defaults to true. Helm
installs will fail if cert-manager takes longer than (timeoutSeconds - 30)
seconds to issue certificates. To restore the old behavior of allowing
installs to succeed with a warning, set `pkiInitJob.failOnTimeout=false`.
Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com>
* fix(helm): make cert-manager resources pre-install hooks to fix ordering
Make Certificate and Issuer resources run as pre-install/pre-upgrade hooks
with weight -30, before the certgen hook (weight -20). This fixes the
chicken-and-egg problem where the certgen hook was waiting for Secrets
created by Certificates that hadn't been created yet.
**Hook ordering:**
1. Certificate and Issuer resources created (weight -30)
2. cert-manager issues certificates and creates Secrets
3. certgen hook runs (weight -20), finds Secrets, creates ConfigMap
4. Main resources (StatefulSet, Service, etc.) created
Previously, the certgen pre-install hook would run before any resources
were created, poll for a non-existent Secret, timeout, and fail. The
Certificate resources would never get created because Helm waits for
all pre-install hooks to succeed before creating main resources.
This fix allows single-stage installs to work reliably as long as
cert-manager can issue certificates within the polling timeout.
Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com>
* feat(helm): add validation to prevent enableMtls with BackendTLSPolicy
Add Helm chart validation that fails the install if both
server.tls.enableMtls=true and grpcRoute.backendTLSPolicy.enabled=true
are set, since this is an invalid configuration.
BackendTLSPolicy requires mTLS to be disabled because the Gateway proxy
cannot present client certificates to the backend. This validation provides
immediate, clear feedback at install time rather than allowing the
misconfiguration to be discovered through runtime errors.
Example error message:
```
Error: grpcRoute.backendTLSPolicy requires mTLS to be disabled because
the Gateway proxy cannot present client certificates to the backend;
set server.tls.enableMtls=false
```
Also updated documentation to mention this validation check.
Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com>
* docs(helm): clarify pkiInitJob.timeoutSeconds polling behavior
Improve documentation to clearly explain that pkiInitJob.timeoutSeconds
controls the Job deadline, but the actual polling timeout is
(timeoutSeconds - 30) to reserve 30 seconds for ConfigMap creation
and cleanup.
Added concrete example: "timeoutSeconds=180 allows 150 seconds of polling"
to make the relationship explicit and avoid confusion where users might
expect the hook to poll for the full timeout value.
Updated both values.yaml inline comments and ingress.mdx documentation
for consistency.
Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com>
* docs(openshift): remove outdated two-stage install instructions
Update OpenShift documentation to reflect that single-stage installs now
work with cert-manager and BackendTLSPolicy. The Certificate resources
run as pre-install hooks (weight -30) before certgen (weight -20),
allowing the hook to poll for and find the issued certificates.
Removed the outdated two-step process:
1. helm install (cert-manager issues cert, hook logs warning)
2. helm upgrade (hook creates ConfigMap)
Replaced with current single-stage behavior:
- Certificate resources created as pre-install hooks
- certgen hook polls for up to 90 seconds (configurable)
- Single helm install succeeds in most cases
- Fails fast by default if timeout reached
This brings openshift.mdx in line with the already-updated ingress.mdx
documentation.
Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com>
* fix(helm): make pkiInitJob.timeoutSeconds the actual polling duration
The timeout value now represents the actual polling time that users
experience when waiting for cert-manager to issue certificates.
The Job activeDeadlineSeconds is set to (timeoutSeconds + 30) to
allow buffer time for ConfigMap creation and cleanup.
Previously, the hook polled for (timeoutSeconds - 30) seconds, which
was confusing when users set timeoutSeconds=180 and only got 150
seconds of actual polling.
Updated documentation in values.yaml, ingress.mdx, and openshift.mdx
to reflect the clearer behavior.
Signed-off-by: Brandon Squizzato <bsquizza@nvidia.com>
* docs(helm): update values.yaml and README with correct polling duration
Updated the caCertificateConfigMapName description to reflect that the
hook polls for exactly pkiInitJob.timeoutSeconds seconds, not
(timeoutSeconds - 30) seconds.
Regenerated README.md with helm-docs.
Signed-off-by: Brandon Squizzato <bsquizza@nvidia.com>
* docs(kubernetes): add OIDC configuration to helm install and CLI examples
Updated all helm install and openshell gateway add examples in ingress.mdx
and openshift.mdx to include OIDC issuer and audience configuration.
Examples now use concrete placeholder values:
- OIDC issuer: https://keycloak.example.com/realms/openshell
- OIDC audience: openshell-cli
- Hostname: gateway.example.com
- ClusterIssuer: letsencrypt-prod
This makes it clearer how to configure OIDC authentication, which is
required when using BackendTLSPolicy or HTTPS termination since the
Gateway proxy cannot present client certificates to the backend.
Signed-off-by: Brandon Squizzato <bsquizza@nvidia.com>
* docs(kubernetes): explicitly list OIDC client ID in gateway add examples
Added --oidc-client-id openshell-cli to all openshell gateway add
commands in ingress.mdx and openshift.mdx, making the default client
ID explicit in the examples even though it's the CLI default.
This improves clarity and helps users understand the complete OIDC
configuration needed for gateway registration.
Signed-off-by: Brandon Squizzato <bsquizza@nvidia.com>
* fix(helm): address PR review feedback for BackendTLSPolicy
- Read backend CA from the authoritative server Secret instead of the
in-memory PKI bundle so enabling BackendTLSPolicy on an existing
release uses the CA that actually signed the server certificate.
- Reconcile the backend CA ConfigMap on every hook run (compare and
update) instead of skipping when it already exists, so CA rotations
propagate automatically.
- Remove hook annotations from cert-manager Issuer/Certificate resources
so they remain regular release objects managed by Helm lifecycle. Split
the cert-manager backend CA ConfigMap creation into a separate
post-install/post-upgrade hook Job that polls after cert-manager
Certificate resources are applied.
- Update architecture/gateway.md, docs/reference/gateway-config.mdx,
debug-openshell-cluster skill, and helm-dev-environment skill with
BackendTLSPolicy, backend CA ConfigMap, enableMtls, and timeout
documentation.
Signed-off-by: Brandon Squizzato <bsquizza@nvidia.com>
Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>
* fix(docs): resolve markdown lint errors in helm README and kubernetes docs
Escape inline HTML angle brackets in README.md template placeholders,
remove trailing spaces, and add blank lines around fenced code blocks
in numbered lists.
Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>
* Update docs
* fix(helm): escape inline HTML in values.yaml descriptions and sync mise lockfile
Wrap `<fullname>` and `<namespace>` template placeholders in backticks
so markdownlint does not flag them as inline HTML (MD033). Regenerate
mise.lock to match current mise.toml after rebase onto main.
Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>
* Run 'mise lock'
* fix: align mise.lock with CI mise version output
The lockfile was regenerated locally with mise 2026.8.10 which resolves
uv Linux artifacts to gnu variants and adds provenance_verified fields,
but CI uses v2026.4.25 which produces musl variants without those fields.
Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>
* fix(docs): correct cert-manager hook ordering and clientCaSecretName comment
Update ingress.mdx and openshift.mdx to describe Certificate resources
as regular release objects with a post-install/post-upgrade Job, matching
the current implementation and architecture/gateway.md.
Fix values.yaml clientCaSecretName comment to state that "" disables
client certificate verification, matching the helper and access-control
docs.
Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>
---------
Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>
Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com>
Signed-off-by: Brandon Squizzato <bsquizza@nvidia.com>
generate_pki minted a CA with no key usage and server and client leaves
with no Authority Key Identifier. RFC 5280 requires both, and verifiers
that enforce it reject the chain: OpenSSL X509_STRICT fails with
"Missing Authority Key Identifier", and Python 3.13 turned that flag on
by default in ssl.create_default_context(). rustls and BoringSSL do not
enforce it, so gRPC clients kept working while an HTTPS client built on
Python 3.13 (for example a platform proxying to an exposed sandbox
service) could not complete a handshake with a pkiInitJob-provisioned
gateway at all. cert-manager PKI was unaffected.
Set keyCertSign and cRLSign on the CA and use_authority_key_identifier
on both leaves, matching what the sandbox L7 CA already does. Add a
test that parses the bundle and asserts the extensions, including that
each leaf AKI matches the CA SKI.
Verified: openssl verify -x509_strict accepts both leaves, and a strict
Python 3.13 client completes an mTLS handshake against a server using
the new bundle where the previous bundle reproduces the failure.
Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com>
- Replace BSD-incompatible in-place sed calls with portable temp-file rewrites.
- Remove test-only shell interception and capture generated gateway config
directly.
- Allow parity tests to use supplied supervisor binaries without resolving a
Linux target.
- Normalize temporary-directory paths and use portable RPM config installation.
- Set a valid setuptools-scm version for Python protobuf generation in Jujutsu
checkouts.
Signed-off-by: Kris Hicks <khicks@nvidia.com>
* fix(ci): restore Windows test portability
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
* fix(tasks): skip Unix lockfile check on Windows
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
* fix(tasks): use buf shim on Windows
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
---------
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
forward_handler_preserves_ssrf_response_and_denial_stage's ssrf_denied
check was the only assertion in the test without a failure message,
unlike its sibling assertions. This test has failed intermittently in
CI with no way to see what response was actually returned, since the
message is where that gets surfaced.
Signed-off-by: politerealism <burdcat17@gmail.com>