mirror of
https://github.com/NVIDIA/OpenShell.git
synced 2026-10-02 07:34:45 +08:00
main
185
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
4784e79451 |
refactor(sandbox): remove unreachable root-side identity and workspace code (#3979)
* refactor(sandbox): remove unreachable root-side identity and workspace code RFC 0012 moved the workload into its own capability-free container that starts as the final sandbox identity. The sandbox no longer runs a root supervisor that prepares the filesystem, rewrites account files, resolves OCI USER entries, or drops privileges before launching the workload, so that code had no production callers. Remove the unreachable paths and their tests: - prepare_filesystem / prepare_filesystem_with_identity, the /sandbox and OCI workspace chown preparation, and the root-side workspace validation (validate_oci_workspace and its privilege-dropped subprocess) - the hidden validate-workspace subcommand - drop_privileges / drop_privileges_with_identity, capability bounding set clearing, validate_sandbox_user/group, and /etc/passwd and /etc/group rewriting - the sandbox-side OCI USER resolver (identity.rs) and ResolvedProcessIdentity; the boundary now writes the driver-resolved UID/GID into the policy directly The workspace check that still runs inside the capability-free boundary (validate_oci_workspace_as_effective_identity) is unchanged. Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> * chore(sandbox): remove unused capability dependency and refresh Landlock comments Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> --------- Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> |
||
|
|
b8ffe5244c |
test(podman): move podman_preflight into driver-podman integration tests (#3783)
* test(podman): move podman_preflight into driver-podman integration tests podman_preflight verifies that openshell-driver-podman fails fast when its Podman socket is unreachable. It only needs the standalone driver binary, not a gateway, so it never fit the gateway-backed e2e-podman harness it lived under and never ran anywhere in CI. Move it into crates/openshell-driver-podman/tests/ as a plain Cargo integration test. It now runs via the existing required workspace test job with no special mise task, workflow step, or coverage exception. Signed-off-by: politerealism <burdcat17@gmail.com> * test(podman): make preflight diagnostics portable Signed-off-by: Evan Lezar <elezar@nvidia.com> --------- Signed-off-by: politerealism <burdcat17@gmail.com> Signed-off-by: Evan Lezar <elezar@nvidia.com> Co-authored-by: Evan Lezar <elezar@nvidia.com> |
||
|
|
798500ccdb |
fix(policy): validate raw OPA settings and redact startup errors (#3788)
* test(policy): reproduce raw OPA loading gaps against the typed schema The supervisor loads a sandbox policy in two ways: through the typed schema (parse_sandbox_policy, then from_proto) or directly into OPA (from_strings and from_files). The raw path fills in defaults where the typed schema is strict, so the same policy text can produce a different sandbox configuration, or load when it should be rejected. Add two regression tests that fail on the current code: - An empty filesystem_policy loads with include_workdir true through raw OPA and false through the typed schema. An absent stanza gives true on both paths and must keep doing so. - Raw OPA accepts a string include_workdir, a non-string read_only entry, an unknown Landlock compatibility and an explicit null json_rpc, with or without a version key. The typed schema rejects each. Every case has a valid twin that both paths must accept. A follow-up change makes raw loading apply the typed schema's rules. Refs #3092. Signed-off-by: Shiju <shiju@nvidia.com> * fix(policy): align raw OPA loading with typed settings Validate raw filesystem, Landlock, and process settings with the canonical authored schema before normalization. Preserve the absent filesystem default while applying the present-stanza default, and canonicalize valid Landlock enum representations before runtime evaluation. Reject explicit null JSON-RPC options through the shared parser. Preserve versionless and runtime OPA data, and keep rejected reloads from replacing the active policy or advancing its generation. Add raw-versus-typed, file-loader, and rejected-reload regressions and document the local loading contract. Refs #3092. Signed-off-by: Shiju <shiju@nvidia.com> * fix(policy): validate raw OPA settings and redact startup errors Validate raw network fields through the authored schema before normalization. Preserve custom Rego data and supported runtime forms. Apply the shared filesystem path checks and non-root identity predicate to raw static settings. Discard authored Rego source and nested errors from static configuration evaluation. Cover malformed inputs, valid controls, file loading, and rejected reloads retaining active decisions and generation. Refs #3092. Signed-off-by: Shiju <shiju@nvidia.com> * test(policy): satisfy unit-returning assertion lint Terminate the two error-assertion match arms with semicolons, as required by Clippy. Preserve the existing checks and runtime behavior. Refs #3092. Signed-off-by: Shiju <shiju@nvidia.com> --------- Signed-off-by: Shiju <shiju@nvidia.com> |
||
|
|
2fe5a0e19c |
perf(kubernetes): use a TCP readiness probe for the supervisor (#3700)
- Kubernetes now checks supervisor readiness by connecting to TCP port 5501 - Stop starting a supervisor process in every sandbox each second - The supervisor opens the port only while its gateway session is up - Accept IPv4 and IPv6 probes, even when net.ipv6.bindv6only is set - Keep the health socket for Docker, Podman, and debugging - Add tests and update the docs Signed-off-by: divesh <dgude@nvidia.com> |
||
|
|
1358941b81 |
feat(mcp): inspect requests with Tower-selected protocol profiles (#3335)
* fix(sandbox-backend): sort boundary request objects before hashing Sort boundary request objects recursively before hashing so serde_json's preserve_order feature cannot change digest identity. Cover canonical bytes, envelope round trips, and rejection of modified provider values and operations. Signed-off-by: Shiju <shiju@nvidia.com> * feat(mcp): upgrade tower-mcp-types to 0.22.2 Upgrade tower-mcp-types from 0.12.0 to an exact-pinned 0.22.2 and use its inspection APIs to validate MCP requests against the selected revision. Carry inspection metadata into policy evaluation and validate requests after header rewriting, before forwarding. Add explicit support for the sessionless 2026-07-28 revision while keeping 2025-11-25 as the default. Validate per-request metadata and standard HTTP header mirrors, and support discovery, tools, and subscription requests. Delegate batch availability and parameter schemas to Tower. Share typed request names between policy and HTTP checks, retain the local batch resource cap, and centralize MCP policy version parsing and ordering. Keep supported MCP revisions and shared allowlist parsing in the canonical policy schema; core re-exports those types. Tower owns wire-profile semantics, and every supported policy revision must map to the matching inspector profile. Reject duplicate JSON keys, invalid known-method parameters, unavailable methods, and unsupported batches. Keep exact extension allow rules and deny precedence. Document request inspection boundaries and add unit, forwarding, and sandbox coverage. Refs #2174. Signed-off-by: Shiju <shiju@nvidia.com> * test(mcp): prove authorization at the forwarding boundary Cover March batch denial in both member orders, valid and malformed controls, and audit behavior across both relay entry paths. Exercise real middleware tool rewrites with matching metadata and assert the exact upstream representation or zero forwarded bytes. Verify legacy bodyless SSE GET remains usable while GET tool bodies and unsupported DELETE cleanup are rejected. Clarify request-selected profile and middleware mutation comments without changing production behavior. Signed-off-by: Shiju <shiju@nvidia.com> * test(mcp): exercise permitted profiles through the sandbox proxy Cover March and June singleton policies and select November and July separately under one endpoint allowlist. Capture upstream tool receipts to distinguish proxy policy denial from an upstream rejection. Extend middleware rewrite coverage to June and multi-version policies, and preserve the sessionless discovery and subscription checks through the shared fixture helpers. Signed-off-by: Shiju <shiju@nvidia.com> * test(kubernetes): box the admission check future Keep the admission test future below Clippy's size limit when the workspace dependency features are unified. Signed-off-by: Shiju <shiju@nvidia.com> * test(mcp): reuse the forwarding fixture identity cache Share the binary identity cache across protocol-profile cases, matching the proxy lifecycle and avoiding repeated hashes of the test executable. Keep procfs authorization and all forwarding assertions intact. Signed-off-by: Shiju <shiju@nvidia.com> --------- Signed-off-by: Shiju <shiju@nvidia.com> |
||
|
|
7a50c0899f |
fix(cli): stream piped exec stdin beyond gRPC request limit (#3687)
* fix(cli): stream piped exec stdin across gRPC messages Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(cli): preserve small exec requests and surface stdin errors Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(cli): preserve exec stdin limit across gRPC streaming Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
1e34e8c576 |
fix(drivers): require admission labels for external resources (#3538)
* fix(drivers): require admission labels for external resources Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(drivers): address resource admission review findings Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(core): reserve driver-owned admission labels Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(core): clarify workspace admission label Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(drivers): clarify resource admission failures Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(e2e): configure resource admission fixtures Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): retry forbidden admission lookups Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(e2e): preserve external driver admission defaults Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
470a34635d |
fix(api): make WatchSandbox loss-aware and resumable (#3209)
* fix(api): emit warning on WatchSandbox broadcast lag instead of terminating Broadcast lag on the status, log, and platform receivers was converted to a RESOURCE_EXHAUSTED status that terminated the whole watch stream. Lag is recoverable: the receiver resumes at the oldest surviving message. Emit a SandboxStreamWarning and continue streaming instead; keep terminating on Closed. Add helpers and unit tests covering the warning payload and receiver recovery after lag. Partially addresses #3055 (cursor/resume follow up separately). Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * refactor(server): group per-sandbox log bus state and stamp sequence numbers Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * feat(proto): add resume cursor fields to sandbox watch API Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * feat(server): stamp watch cursors from a shared per-sandbox sequence Allocate cursors from a single SeqAllocator shared by the log and platform event buses, so a sandbox's merged watch stream carries unique, strictly increasing cursors. A single resume_after_cursor can then unambiguously locate a client's position across both sources. Rewrite both publish paths to allocate the sequence, stamp event.cursor, send, and append to the tail under one lock. This removes the previous get_mut().expect() TOCTOU race where a concurrent remove() between the two lock sections could panic. Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * feat(server): serve WatchSandbox resume from cursor with gap detection Add tail_after() to the log and platform event buses, returning every buffered event newer than a client's resume cursor. Each PerSandbox now tracks last_trimmed_seq (the highest seq it has evicted) so a resume is reported as an unrecoverable ResumeGap only when this bus dropped an event the client still needs. Judging gaps by evictions, not by the tail's oldest seq, is required under the shared cursor space: each bus's tail is non-contiguous in the global sequence because the other bus owns the missing seqs, so comparing against tail.front() would flag false gaps. Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * feat(server): resume WatchSandbox from cursor across log and platform buses Wire resume_after_cursor into the watch producer. On a non-zero cursor, replay events strictly after it from both the log and platform buses, merge by shared cursor, and emit in order before entering the live loop. A trimmed range on either bus is an unrecoverable gap and terminates the stream with OUT_OF_RANGE carrying the requested and earliest-available cursors, distinct from recoverable lag which warns and continues. Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * test(server): cover WatchSandbox cursor resume paths Add handler-level tests for the resumable watch stream: replay strictly after the client cursor, merge log and platform events in shared-cursor order, suppress duplicates when resuming at the latest cursor, and terminate with OUT_OF_RANGE when the requested cursor has been trimmed. Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * docs(api): document WatchSandbox loss-awareness and resume Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * fix(server): deliver watch events once and harden cursor teardown Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * feat(sdk): add loss-aware resumable watch_logs to Rust SDK client Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * fix(server): keep watch cursors monotonic across teardown and restart Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * fix(server): merge live watch sources by cursor before emission Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * fix(api): bind watch cursors to a cursor space and merge tail sources Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * fix(server): revalidate the watch cursor space after collecting replay Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * test(server): update the public RPC schema fingerprint for the string cursor Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * fix(server): hold watch events above the publication watermark and emit the watch lag warning before its batch Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * test(server): synchronize the watch live-order test with the end of initialization Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * fix(sdk): use canonical sandbox name in watch_logs Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * test(sdk): guard canonical-name addressing in watch_logs Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * fix(server): fix public rpc schema Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * fix(api): reconcile watch resume rebase Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(server): bound interactive relay cleanup Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> --------- Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> Co-authored-by: John Myers <9696606+johntmyers@users.noreply.github.com> |
||
|
|
3107ff1f82 |
ci(security): stage release finding enforcement (#3552)
Signed-off-by: Evan Lezar <elezar@nvidia.com> |
||
|
|
50230616d5 |
refactor(runtime): retire Community image dependencies (#3386)
* feat(sandbox): default to official Alpine sandbox image default_sandbox_image() now returns docker.io/library/alpine:3.22, a generic version-qualified official image, so a fresh install no longer depends on the community sandbox image catalog. All compute drivers (docker, podman, kubernetes, vm) inherit this fallback. Part of #3116. Signed-off-by: Akram Signed-off-by: Akram <akram.benaissi@gmail.com> * feat(deploy): default deployment configs to the official Alpine sandbox image Update the shared gateway default_image, Helm chart values, the standalone Kubernetes manifest, and the dev gateway task scripts to use docker.io/library/alpine:3.22 instead of the community base image, consistent with default_sandbox_image(). GPU e2e image-build base is left unchanged (CUDA needs a glibc base). Part of #3116. Signed-off-by: Akram Signed-off-by: Akram <akram.benaissi@gmail.com> * feat(driver): default to numeric non-root identity for USER-less images With the default sandbox image now Alpine, images that declare no OCI USER must start instead of being rejected. When the image declares no USER and the policy requests none, the Podman and Docker drivers now supply a numeric non-root identity (DEFAULT_SANDBOX_UID/GID = 1000) instead of rejecting, matching the numeric-identity behavior of the Kubernetes and VM drivers. The supervisor's resolved-identity path runs the sandbox as a synthesized non-root account without the account existing in the image. Images that declare a USER keep the OCI resolution path unchanged. Part of #3116. Signed-off-by: Akram <akram.benaissi@gmail.com> Signed-off-by: Evan Lezar <elezar@nvidia.com> * test(conformance): use Alpine workload image Signed-off-by: Evan Lezar <elezar@nvidia.com> * refactor(policy): drop community image /app path from default policy The restrictive default policy granted read-only access to /app, a directory that only existed in the community base image. A generic Alpine default has no /app, so remove it. Landlock best-effort already ignores absent paths; this just stops advertising a community-specific layout in the default. Part of #3116. Signed-off-by: Akram Signed-off-by: Akram <akram.benaissi@gmail.com> * docs(config): document Alpine default images Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(podman): report early sandbox termination Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(podman): initialize rootless workspace ownership Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(sandbox): qualify NVIDIA Ubuntu default Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): initialize rootful default workspace Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(sftp): add native sandbox adapter Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sftp): gate runtime helper support to Linux Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sftp): support standard OpenSSH file operations Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sftp): harden rename and special file handling Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(runtime): remove community image dependencies Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(e2e): build provider readiness tool fixture Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(e2e): use a dedicated Noble fixture for Docker tests Signed-off-by: Evan Lezar <elezar@nvidia.com> --------- Signed-off-by: Akram Signed-off-by: Akram <akram.benaissi@gmail.com> Signed-off-by: Evan Lezar <elezar@nvidia.com> Signed-off-by: Drew Newberry <anewberry@nvidia.com> Co-authored-by: Evan Lezar <elezar@nvidia.com> Co-authored-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
96c08f111b |
refactor(isolation)!: make confirmation backend-neutral (#3366)
* refactor(isolation): make confirmation backend-neutral Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): validate confirmation evidence at host boundary Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(isolation): keep fence evidence driver-owned Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation)!: validate explicit fence projections Require each compute driver to map its native evidence to individual outer-fence guarantees, and reject incomplete projections before a boundary becomes ready. Exercise the assembled remote confirmation path for invalid audit, property, generation, and digest evidence. BREAKING CHANGE: BoundaryConfig and SandboxRuntimeDescriptor use outer_fence projections rather than the earlier driver_fence representation. State written by earlier builds cannot be decoded; operators must stop and recreate affected sandboxes after upgrading. Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(isolation): clarify outer fence ownership Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): update confirmation test fixtures Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
4cd5e54780 |
fix(deps): update rustls past RUSTSEC-2026-0285 (#3484)
Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
58b5f8f976 |
feat(prover): add standalone policy boundary checker (#3289)
* feat(prover): add standalone policy maximum checker Signed-off-by: Johnny Greco <jogreco@nvidia.com> * refactor(prover): simplify check scope schema Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(prover): align containment and cancellation with runtime Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(prover): stabilize containment checks in CI Signed-off-by: Johnny Greco <jogreco@nvidia.com> * test(prover): avoid solver in fast-path guard test Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(prover): align string containment with runtime Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(prover): reject ambiguous z3 string escapes Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(prover): align containment with runtime boundaries Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(prover-cli): harden cancellation and invalid input Signed-off-by: Johnny Greco <jogreco@nvidia.com> * feat(packaging): install policy prover with OpenShell Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(prover): clarify installation and check results Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(build): describe prover distribution directly Signed-off-by: Johnny Greco <jogreco@nvidia.com> * refactor(prover): rename maximum policy to boundary Signed-off-by: Johnny Greco <jogreco@nvidia.com> * refactor(prover): localize fail-closed validation Signed-off-by: Johnny Greco <jogreco@nvidia.com> * test(prover): cover fail-closed CLI surfaces Signed-off-by: Johnny Greco <jogreco@nvidia.com> * test(prover): allow CI load for REST solver proof Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(prover): use canonical policy schema for containment Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(prover): preserve uncertainty for runtime binary globs Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(prover): bound policy validation work Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(prover): document validation resource limits Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(prover): make containment API extensible Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(prover): define containment API contract Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(ci): integrate prover with consolidated builds Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(ci): declare release packaging dependency Signed-off-by: Johnny Greco <jogreco@nvidia.com> --------- Signed-off-by: Johnny Greco <jogreco@nvidia.com> |
||
|
|
d68b7069c3 |
refactor(proto)!: use well-known time types (#3113)
* refactor(proto)!: use well-known time types Signed-off-by: Derek Carr <decarr@redhat.com> * fix(proto): preserve time migration behavior Signed-off-by: Derek Carr <decarr@redhat.com> * fix(proto): preserve timestamp boundary semantics Signed-off-by: Derek Carr <decarr@redhat.com> * fix(proto): convert sandbox token expiry to timestamp Signed-off-by: Derek Carr <decarr@redhat.com> * fix(proto): preserve time compatibility semantics Signed-off-by: Derek Carr <decarr@redhat.com> * fix(proto): preserve exact endpoint and profile times Signed-off-by: Derek Carr <decarr@redhat.com> * test(e2e): use duration for interactive exec timeout Signed-off-by: Derek Carr <decarr@redhat.com> * fix(sdk-go)!: remove legacy profile duration fields BREAKING CHANGE: Go provider profile callers must use RefreshBefore, MaxLifetime, and CacheTTL with ProfileDuration instead of the whole-second fields. Signed-off-by: Derek Carr <decarr@redhat.com> --------- Signed-off-by: Derek Carr <decarr@redhat.com> |
||
|
|
2ccef97769 |
feat(policy): establish one canonical authored policy representation (#3334)
* chore(policy): restart schema implementation Signed-off-by: Johnny Greco <jogreco@nvidia.com> * feat(policy-schema): add canonical authored policy model Signed-off-by: Johnny Greco <jogreco@nvidia.com> * refactor(policy): use canonical authored schema Signed-off-by: Johnny Greco <jogreco@nvidia.com> * refactor(prover): project canonical policy documents Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): describe shared schema boundary Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(policy): close reviewed parser gaps Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(policy-schema): fail closed on unsupported fields Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(policy): preserve partial process identities Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(policy-schema): harden authored policy inspection Signed-off-by: Johnny Greco <jogreco@nvidia.com> * refactor(policy): rename policy schema crate Signed-off-by: Johnny Greco <jogreco@nvidia.com> * revert(policy): restore policy schema crate Signed-off-by: Johnny Greco <jogreco@nvidia.com> * test(e2e): serialize OIDC PKCE scenarios Signed-off-by: Johnny Greco <jogreco@nvidia.com> --------- Signed-off-by: Johnny Greco <jogreco@nvidia.com> |
||
|
|
c1f2e7189f |
feat(isolation): implement the RFC 0012 sandbox architecture (#2942)
* feat(isolation): add RFC 0012 backend contract Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> * refactor(isolation): name the interface crate explicitly Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): expose trusted host gateway Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(agents): inventory the MXC driver Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): add mediated DNS transport Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): tighten interface error and digest contracts Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(isolation): remove unrelated driver inventory Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): define capability-free launch contract Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): seal confirmed boundary state Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): validate confirmation for external backend implementations Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): clarify mediated DNS identity Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(isolation): generalize loopback connector Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(isolation): unify typed network mediation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): bind launches to sandbox sessions Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(mxc): initialize extended sandbox status Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): add boundary protocol and Linux primitives Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): harden signals and separate process status from transport Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): validate remote confirmation through public contract Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): validate wire state and propagate snapshot failures Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(isolation): import owned agent specification explicitly Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(isolation): describe mediated DNS channel Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): bound mediation attach without nested retries Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(isolation): generalize loopback protocol Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): add transport-neutral session authentication Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(isolation): separate sandbox backend protocol Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): harden runtime boundary controls Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): add terminal boundary operation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): split supervisor and sandbox runtimes Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): harden boundary isolation and lifecycle ownership Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): reject private root redirects and adopt typed errors Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): preserve accept thread ownership on musl Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(sandbox): isolate credential probes from filtered threads Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): return retained exec exit status to independent waiters Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): bound network mediation and preserve socket authorization Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): bound control admission and retire stale mediation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * ci(e2e): select migrated drivers per stack layer Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(sandbox): implement loopback connector Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): authenticate the Sandbox Protocol Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(supervisor): rotate launch-scoped authentication Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(sandbox): consume dedicated backend crate Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(sandbox): align topology session fixture Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): align projected bootstrap bundle Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(auth): validate refreshed credentials before rotation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): fail closed across supervisor disconnects Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): repair rebased sandbox CI Signed-off-by: Drew Newberry <anewberry@nvidia.com> * build(runtime): publish separate sandbox and supervisor images Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(config): configure the sandbox runtime image Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(ci): validate sandbox binary linkage Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(isolation): use backend and runtime terminology Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(sandbox): use a scratch runtime image Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(ci): refresh schema and dependency policy Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): bind reconnects to supervisor process Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs: align runtime split operational guidance Signed-off-by: Drew Newberry <anewberry@nvidia.com> * chore(security): document Kubernetes runtime RBAC Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): enforce runtime lifecycle invariants Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(compute): identify sandbox start generations Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(server): restore sandbox launch sessions Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): support authenticated runtime replacement Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(auth): bind sandbox session successors Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(auth): retry pending sandbox successors Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(vm): run the supervisor outside the guest workload Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(vm): use sandbox backend protocol Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): repair rebase integration Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): use unified build toolchain Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(vm): use sandbox runtime terminology Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(vm): own guest network bootstrap Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): expose guest init version Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): select native supervisor artifacts Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): guard guest init Linux symbols Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): scope Linux test imports Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): avoid guest interface casts Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): reconcile admitted sandbox identity Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): share resolved sandbox identity Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): surface host supervisor failures Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): include guest logs on supervisor exit Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): rotate and clean runtime generations Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): make sandbox starts generation-aware Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): rotate restored sandbox sessions Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(vm): keep shared paths in the base layer Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): bind sandbox session lineage Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(docker): isolate workloads behind the host supervisor Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(docker): rotate launch-scoped authentication Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(docker): use sandbox backend protocol Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(docker): use host networking for supervisor Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): preserve host gateway alias resolution Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(docker): use separate sandbox and supervisor images Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): restore startup validation after rebase Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(docker): name the sandbox runtime directly Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): narrow supervisor CA runtime storage Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): close companion isolation gaps Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(docker): align mediated network expectations Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(docker): exercise mediated network paths Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): attach supervisor to managed network Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): defer supervisor recovery until gateway is ready Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): make sandbox starts generation-aware Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): rotate restored sandbox sessions Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): preserve workloads during session rotation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(docker): remove unrelated configuration RFC changes Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): bind sandbox session lineage Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(kubernetes): add proxy-pod isolation topology Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(kubernetes): use sandbox backend protocol Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): use stable sandbox service authority Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(kubernetes): split sandbox and supervisor images Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): adapt proxy pods to current runtime APIs Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(kubernetes): describe the single runtime placement Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(kubernetes): simplify sandbox orchestration Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): validate deployment prerequisites Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): update Trivy Helm profile inventory Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(kubernetes): update Trivy scan inventory count Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): reuse preloaded runtime images in e2e Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): type and clean runtime resources Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): make sandbox restarts recoverable Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): rotate restored sandbox sessions Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): preserve supervisor egress Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): bind sandbox session lineage Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(podman): adopt isolated sandbox and supervisor containers Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): stage bootstrap archives at named volume destinations Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(podman): rotate launch-scoped authentication Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(podman): use sandbox backend protocol Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(podman): use host networking for supervisor Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(podman): split sandbox and supervisor images Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): repair rebase integration Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(podman): name the sandbox runtime directly Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): provision supervisor CA runtime storage Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): address isolation review findings Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): inspect Debian supervisor provenance Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): use libpod-compatible tmpfs options Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): bind verified sandbox runtime binary Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): provide external driver data directory Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): start sandbox before joining user namespace Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): separate supervisor user namespace Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): make sandbox starts generation-aware Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): rotate restored sandbox sessions Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): bind sandbox session lineage Signed-off-by: Drew Newberry <anewberry@nvidia.com> * perf(isolation): add TCP and DNS benchmark harnesses Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(perf): align benchmark timing and supported protocols Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(perf): report TCP benchmark metrics accurately Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(perf): cancel failed worker startup Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): build matching local supervisor image Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): make local sandbox smoke test runnable Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): wire local sandbox runtime image Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): narrow sandbox service RBAC Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(ci): validate split runtime artifacts Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): harden runtime session handling Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(supervisor): add standalone network proxy role Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(rfc): remove implementation companion notes Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(vm): standardize runtime release name Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): pin renamed runtime artifacts Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(auth): persist sandbox runtime identity Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(runtime): restore branch validation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): reconcile main after rebase Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(network): close unframed HTTP 1.0 responses Signed-off-by: Drew Newberry <anewberry@nvidia.com> * chore(isolation): preserve upstream OCSF updates Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(security): close credential and TLS replay paths Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(auth): make sandbox refresh retries idempotent Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
dbe36eaf85 |
fix(security): harden Vault credential transport (#3329)
Reject non-loopback plaintext Vault endpoints, disable redirects, and support private CA bundles without weakening hostname verification. Update Helm configuration, documentation, operator skills, and regression coverage for OSSR-002. Signed-off-by: Seth Jennings <sjenning@redhat.com> |
||
|
|
481ce566e1 |
fix(ocsf): correct HTTP activity context (#3316)
Previously, metadata events omitted both HTTP request and response objects, early proxy rejections used HTTP Activity without request context, and unsupported-scheme events did not expose enough safe HTTP context to satisfy the OCSF 1.8 schema. Now, metadata events include a method-only request and their actual HTTP response codes without recording the metadata URL. Unsupported-scheme events also include a method-only request plus the generated 400 response. Authority mismatches and credential-resolution denials use HTTP Activity with their generated 403 or 500 responses, and HTTP activity IDs are derived from the request method. Additionally, HttpActivityBuilder now enforces the OCSF request-or-response constraint at compile time. Signed-off-by: Kris Hicks <khicks@nvidia.com> |
||
|
|
39cf4823f7 |
feat(api): add structured gateway errors and SDK decoding (#3313)
* feat(api): expose structured gateway errors across SDKs Refs #3051. Add standard validation, conflict, and retry details; preserve raw transport status in Rust, Go, TypeScript, and Python; document status and recovery guidance. This is the structured-error foundation only. Mutation result shapes, allow_missing, durable request deduplication, and exec retry semantics remain follow-up work. Signed-off-by: Mrunal Patel <mrunalp@gmail.com> * fix(python): preserve wrapped RPC cleanup handling Inspect the original gRPC call when handling missing sandboxes during deletion waits and managed cleanup. Add intercepted cleanup regressions and clarify the error-wrapper migration contract. Addresses the cleanup review on #3313; part of #3051. Signed-off-by: Mrunal Patel <mrunalp@gmail.com> --------- Signed-off-by: Mrunal Patel <mrunalp@gmail.com> |
||
|
|
c195e23267 |
test(conformance): remove plan-driven continuity tests (#3342)
Signed-off-by: Evan Lezar <elezar@nvidia.com> |
||
|
|
fd3fd9cf74 |
feat(sandbox): explain failed calls to external tool servers (#3207)
Show configured tool server addresses and their last observed connection results together in sandbox status. Keep sandbox lifecycle readiness separate so an external connection failure does not mark the sandbox unready. Expose direct endpoint records through the CLI and SDKs, with plain-language failure explanations and gateway acceptance times. Keep observation tracking, runtime reporting, and gateway validation in dedicated endpoint status modules. Preserve bounded reporting, request attribution, retry ordering, and configuration and supervisor authority checks. Clear obsolete observations while retaining the configured addresses, and document the distinction between an observed HTTP response, current availability, and tool success. Signed-off-by: Shiju <shiju@nvidia.com> |
||
|
|
26f2f96393 |
feat(mxc): add Windows host proxy for MXC sandbox network egress (#3163)
* Implement Windows host proxy integration and update dependencies for OpenShell
* Update README and gateway config to clarify egress proxy address handling and allocation
* Refactor ProxyIdentityMode to return Result for static_binary and add tests for binary path and SHA256 hash
* Enhance platform_hosts_path for Windows to use SystemRoot and improve error handling for hosts file reading
* Refactor FileFingerprint to use Option for mtime and ctime, simplifying metadata handling
* Add conditional compilation for Windows host module
* add unit tests for OPA policy evaluation and identity handling
* remove openshell-supervisor-network from unsupported driver package test exclusion list
* feat(mxc): enable host proxy TLS state generation
Generate per-sandbox TLS state for the MXC host proxy so HTTPS L7 enforcement can use the same MITM path as Linux. Grant generated CA material to the MXC process and inject standard trust env vars, while matching Linux behavior by disabling TLS termination on CA setup failure and relying on proxy fail-closed handling.
* fix(docs): remove outdated notes on governed egress from docs
* fix(tests): update TLS environment variable paths to use temporary directory
* fix(examples): make run-mxc-e2e harness correct and orphan-free
The MXC e2e harness never actually exercised the fs scenarios: it started
the gateway once and patched agent_command per scenario AFTERWARDS, so the
running gateway kept launching the default demo agent (not shipped in the
kit) and every fs scenario failed with CreateProcessW error:2. It also
scored on the `sandbox create` exit code (non-zero due to the harmless
interactive attach), wrote sandbox records to the persistent gateway DB
(leaving orphans that collided on later runs), and its deny scenarios never
proved denial.
Changes:
- Start a FRESH gateway per scenario so each scenario's agent_command is
actually loaded (root cause of CreateProcessW error:2).
- Score by on-disk artifact / expected outcome, not `sandbox create` exit.
- Real deny assertions: a control write to a granted path must succeed
(proves the agent ran) while the denied write must be absent. fs-empty
probes an ungranted out-of-share path (share_dir is mapped rw by design).
- Run the gateway on an ephemeral in-memory DB (sqlite::memory:) so the
harness never writes to the persistent store and cannot leave orphan
sandbox records; also use unique per-run sandbox names + pre-delete.
- Fix the process_container probe: use a real cwd + absolute cmd.exe
(canonical wxc-exec does not expand %TEMP% -> 0x8007010B).
- Fix summary counts (@() so a single FAIL is counted and exit is non-zero).
Verified PASS=4 FAIL=0 on 7F203-MXC-003 (no BaseContainer velocity keys)
using a canonical wxc-exec build (AppContainer fallback).
Signed-off-by: Akber Raza <akberr@nvidia.com>
* fix(e2e): probe timeout is milliseconds (10ms->30000ms)
MXC process.timeout is wall-clock ms (wire.rs). The 10 value meant 10ms,
which the base-container tier (7F203-MXC-001/.181) enforced strictly and
timed the probe out. AppContainer path (.18/-003) happened to slip under
it. Bump to 30000ms so the process_container preflight probe is reliable
across both tiers.
Signed-off-by: Akber Raza <akberr@nvidia.com>
* fix(mxc): use native paths in real runtime probes
Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>
* fix(mxc): make processcontainer work with mxc-latest-released wxc-exec
Three fixes to support the release wxc-exec binary (BaseContainer dispatcher)
in addition to mxc-fixes-env-vars:
1. Seed process env from host (driver.rs)
ProcessContainer starts with a completely blank environment -- no PATH,
SystemRoot, or anything. Seed the process env from the gateway host
environment so the agent binary can locate DLLs and run. Skip internal
Windows drive-letter variables (keys starting with '=') which cause
CreateProcessW to return ERROR_ENVVAR_NOT_FOUND. User agent_env entries
and TLS CA vars are applied as overrides on top of the host env.
2. Remove TLS readonly_paths grant (driver.rs)
The release wxc-exec (BaseContainer dispatcher) requires write-DAC
permission on every path in readonly_paths to set up AppContainer ACLs.
Adding the proxy's temp TLS directory caused a DACL error and exit -1.
The CA cert paths remain available to the agent via TLS env vars.
3. Remove allowedHosts from network JSON (mxc.rs)
The release wxc-exec rejects network.allowedHosts / network.blockedHosts
on Windows with "not yet supported". Removed the loopback exemption
attempt (127.0.0.1, ::1, localhost) from the network section.
Intra-container loopback works natively in the release binary without
it -- the spawner can connect to the server at 127.0.0.1:22000 directly.
Additional changes:
- mxc-ws-agent.rs: add relay-debug.txt error capture and relay-ready.txt
marker for reliable timing of host client connections.
- mxc-ws-gateway.toml: debug = true for JSON config dump during diagnosis.
- run-ws-agent-test.ps1: default port changed to 17670 (gateway default);
relay-ready.txt polling before ws-echo to avoid connecting before the
spawner has established the proxy bridge.
Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>
* fix(e2e): address CodeRabbit review on run-mxc-e2e.ps1 (MR !46)
Four robustness/correctness fixes from CodeRabbit:
1. Start-Gw: kill the spawned gateway before the "did not start within 30s"
throw. If the process is alive but never binds the port, $gw is not yet
assigned in the caller, so the finally block cannot reap it -> orphan
gateway holding the port for the next run.
2. create-fail scoring: a non-zero `sandbox create` exit alone is not proof
of a policy rejection (gateway-registration/transport/fixture errors also
exit non-zero and would false-pass). PASS now requires a genuine
rejection signal (network / invalid_argument / network_policies) AND that
it is not an infrastructure failure; other non-zero exits go to FAIL with
output captured.
3. deny scenarios (ControlTarget path): snapshot the deny target AFTER
Wait-File lands the control artifact, so a late denied write (enforcement
regression racing the control write) can no longer be recorded as PASS.
4. -KeepRunning: break out of the scenario loop after the first scenario so
a later scenario does not start a second gateway on the same port
(previously a reliable port collision instead of a usable debug mode).
Re-verified PASS=4 FAIL=0 on both boxes (7F203-MXC-001 base-container and
7F203-MXC-003 AppContainer fallback); network-policy-rejected correctly
scores as "policy rejection".
Signed-off-by: Akber Raza <akberr@nvidia.com>
* feat(mxc-e2e): collect run-mxc-e2e output into a results bundle
Mirror the sibling run-*.ps1 scripts by collecting every run's logs into a
timestamped results-e2e-<stamp>\ folder and zipping it. The bundle contains the
console transcript, per-scenario gateway stdout/stderr, the exact TOML rendered
for each scenario, the policy fixture used, and a summary.txt with the verdict
table.
Per-scenario gateway logs now land in gateway.<scenario>.log/.err.log inside the
bundle instead of a single fixed gateway.e2e.log in the script directory.
Wrap pre-flight, mode setup, scenario definitions, and the scenario loop in a
single try/catch/finally so the finally always writes the summary, stops the
transcript, and zips the bundle -- even on a pre-flight failure. The existing
per-scenario gateway-cleanup try/finally stays nested inside. All scenario
logic, scoring rules, and comments are preserved.
Signed-off-by: Prashant S Khodade <pkhodade@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>
* fix(mxc-e2e): address CodeRabbit review on run-mxc-e2e.ps1
- Require -Scenario when -KeepRunning: the loop breaks after the first
scenario, so a full-suite run would execute only one scenario yet still
report the suite as PASS. Fail fast so a partial run can't be mislabeled
complete.
- Start-Transcript now runs inside the guarded try block with a
$transcriptStarted flag; Stop-Transcript is only called when it actually
started, so a Start-Transcript failure still yields the results bundle.
- Wrap the -Scenario filter in @() so a single exact match stays an array
(reliable .Count and a proper array for the scenario loop on PS 5.1).
Signed-off-by: Prashant S Khodade <pkhodade@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>
* fix(examples): pass gateway config via OPENSHELL_GATEWAY_CONFIG for spaced paths
Start-Process -ArgumentList does not quote array elements, so launching the
gateway with a bare --config <path> token split on any space in the install
path (e.g. C:\Users\First Last\...), and clap rejected the fragment with
'unrecognized subcommand'. Every MXC example launcher that started the gateway
hit this when the kit was unzipped under a path containing a space.
Pass the config path through the OPENSHELL_GATEWAY_CONFIG env var (which the
gateway already reads via clap) and drop the --config token. Env vars carry
spaces safely.
Affected: run-ocsf-audit, run-mxc-e2e, run-demo, run-inference-test,
run-ollama-test. run-mtls-test was not affected (its launch passes no config
path). Root-caused and fix-verified on 7F203-MXC-003 from a spaced path.
Signed-off-by: Akber Raza <akberr@nvidia.com>
* fix(run-mxc-e2e): improve scoring logic and enhance command execution handling
* fix(mxc): reconcile proxy support after rebase
Restore the proxy-enabled OCSF audit example removed by
|
||
|
|
8d19308c08 |
fix(bootstrap): emit RFC 5280 extensions on generated gateway PKI (#3286)
generate_pki minted a CA with no key usage and server and client leaves with no Authority Key Identifier. RFC 5280 requires both, and verifiers that enforce it reject the chain: OpenSSL X509_STRICT fails with "Missing Authority Key Identifier", and Python 3.13 turned that flag on by default in ssl.create_default_context(). rustls and BoringSSL do not enforce it, so gRPC clients kept working while an HTTPS client built on Python 3.13 (for example a platform proxying to an exposed sandbox service) could not complete a handshake with a pkiInitJob-provisioned gateway at all. cert-manager PKI was unaffected. Set keyCertSign and cRLSign on the CA and use_authority_key_identifier on both leaves, matching what the sandbox L7 CA already does. Add a test that parses the bundle and asserts the extensions, including that each leaf AKI matches the CA SKI. Verified: openssl verify -x509_strict accepts both leaves, and a strict Python 3.13 client completes an mTLS handshake against a server using the new bundle where the previous bundle reproduces the failure. Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com> |
||
|
|
02b664bb0d |
refactor(config): normalize and enforce gateway schema v2 (#2814)
* refactor(config): normalize compute driver field names Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * refactor(config): introduce canonical gateway fields Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * refactor(config): enforce gateway schema version 2 Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): preserve compute driver runtime guarantees Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): address schema v2 review regressions Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): complete schema v2 migration safeguards Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(config): expand schema v2 regression coverage Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(config): add schema v2 parity manifest Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): correct parity manifest inventory Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * docs(config): record schema v2 intentional changes Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * docs(config): disposition schema v2 parity gaps Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): add dual schema parity harness Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): establish compute lifecycle parity baseline Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): preserve gateway option compatibility Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): record gateway option parity Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * docs(config): close gateway-wide parity gaps Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(podman): apply configured pids limit Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): validate Podman option parity Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): add Kubernetes option parity harness Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): record Kubernetes option parity Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): disposition VM parity lanes Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): add external driver parity lane Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(e2e): preserve external driver pull policy Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): attest parity artifacts and launches Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): require clean parity build sources Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): bind parity runtime artifacts Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(e2e): use isolated supervisor tags Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(e2e): qualify parity image tags Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(e2e): serve parity supervisor locally Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): isolate parity podman services Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): harden parity evidence provenance Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): pin parity sandbox artifacts Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): attest parity runtime inputs Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): bind parity runtime evidence Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): record compute boundary parity Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): disposition cross-cutting parity lanes Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(packaging): preflight gateway config upgrades Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): preserve rebase integration guarantees Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(ci): isolate temporary git signing config Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): update remaining schema v2 consumers Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(ci): provide e2fs tools to VM tests Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): align preflight with gateway startup Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(vm): preserve rootfs tar configuration Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * chore(config): adopt duration unit constructors Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(packaging): preflight RPM gateway config Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): address driver review findings Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(e2e): require fresh semantic parity evidence Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(docker): update tests for renamed sandbox label Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(gateway): preserve selective driver coverage after rebase Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> Signed-off-by: Drew Newberry <anewberry@nvidia.com> Co-authored-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
ae57979b03 |
feat(mxc): add Windows ETW-to-OCSF audit trail (#3015)
* feat(mxc): ETW->OCSF audit consumer + Windows OCSF JSONL parity (cp6 P1) Add a Windows MXC ETW->OCSF audit trail in openshell-driver-mxc: a real-time Sandboxing-provider ETW consumer that decodes events (TDH), attributes each to an OpenShell sandbox_id, and maps them to OCSF (lifecycle 6002, config 5019, process 1007, finding 2004). cp6 Phase 1 - durable OCSF JSONL audit-file parity with Linux: - openshell-ocsf: add emit_ocsf_event_routed (populates the event-bridge thread-local AND stamps sandbox_id+message in one dispatch) plus public set/clear_current_event; OS-aware device (Device::windows/for_current_os) so device.os.name reflects the host instead of a hardcoded Linux stub. - etw_consumer: emit via the routed emit (previously fired a bare info! that never populated the bridge, so the structured event was dropped). - openshell-server: install OcsfJsonlLayer over a synchronous daily-rotated appender (durable under force-kill), gated by OPENSHELL_OCSF_JSON, path via %PROGRAMDATA%\OpenShell\logs (override OPENSHELL_OCSF_LOG_DIR). - device.hostname now resolves to the real gateway machine name. Box-proven on 7F203-MXC-001: JSONL lines == shorthand OCSF rows, all valid OCSF JSON, per-sandbox attribution intact, disabled state writes nothing. Signed-off-by: Akber Raza <akberr@nvidia.com> * feat(mxc): map remaining Sandboxing ETW events to OCSF Close the last three ETW->OCSF gaps so the audit trail covers the full set of events the Sandboxing provider emits (12/12): - ProcessLaunched -> Process Activity [1007] "Launch" (confirmed start; carries the real processId/threadId, the twin of CreateProcessInSandbox which only has the request + command line). - SandboxProxyConfigured -> Device Config State Change [5019] (the one network-plane setup event; surfaces proxyPort, "no proxy" when 0). - SandboxConsoleReferencePlumbed -> Device Config State Change [5019] (console-handle plumbing). map_config_state now handles the full config/hardening/setup family and carries proxyPort/hasConsoleReference/creationFlags as unmapped fields. Verified on 7F203-MXC-001: 11/12 event types emit OCSF without a proxy (SandboxProxyConfigured requires proxy config to fire). Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(mxc): seed ETW attribution under registry lock + Device tests Address CodeRabbit review on !31: - Prevent stale ETW attribution on a delete/launch race: register the wxc-exec pid while holding the registry lock, and bail if the sandbox entry is already gone. Previously the attribution key could be seeded after `delete` had removed the sandbox, leaving a stale key that could misroute later Sandboxing ETW events to a dead sandbox_id. Lock order (registry -> attribution) matches the delete path, so no deadlock. - Add unit tests for the new Device::windows and Device::for_current_os constructors to harden Windows/Linux OCSF device parity. Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(mxc-etw): buffer+replay racing events and harden attribution keys Addresses two ETW->OCSF attribution review items (Shailendra #1, #2). #2 early-event loss: ETW delivers the sandbox create/config burst the instant wxc-exec starts, which can beat the driver's register_launch (now under the registry lock post-Ready). process_event previously dropped anything unresolved, losing the racing burst. Add a bounded, time-bounded pending buffer (PENDING_MAX=4096, PENDING_TTL=5s): unresolved events are held and replayed once attribution lands, aged-out ones dropped. Consumer switched to a timed recv_timeout(200ms) so the buffer is re-driven after each event and on a tick. Emit path factored into shared emit_resolved(). #1 attribution collisions: a Windows PID is recycled after exit and a command line is commonly identical across sandboxes. register_launch now rebinds by_pid on reuse and clears the stale last_pid_sid hint (warns if the PID still pointed at a different, leaked sandbox); command line is held in by_cmd only while unique and demoted to a new ambiguous_cmds set on a second owner, so a duplicate command refuses to resolve rather than misroute. Unit tests: buffer replay (direct + cross-link), buffer bound, PID-reuse rebind, duplicate-cmd non-resolution. Box-verified on 7F203-MXC-001 (5 sandboxes, identical cmd -> 5 isolated sandbox_ids, 50/50 OCSF/JSONL, BuffersLost=0). Signed-off-by: Akber Raza <akberr@nvidia.com> * docs(mxc-etw): note cmd_line is captured raw with no privacy filtering Review item #3 (Shailendra): add a PRIVACY NOTE on map_process_launch stating cmd_line is copied verbatim into OCSF process.cmd_line with no redaction, so secrets/PII on a command line land unredacted in the durable audit trail (deliberate audit-fidelity trade-off; treat the log as sensitive). Redaction is owned by an upstream privacy layer, not this path; no general audit-output PII scrubber exists today (openshell_core::secrets [CREDENTIAL] redaction is scoped to the proxy HTTP-target logging, a separate egress path). Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(mxc-etw): open ETW trace on caller thread so start_session reports real status Review item #4 (Shailendra): start_session previously returned Ok(EtwSession) as soon as the pump thread was spawned, but OpenTraceW ran later inside that thread; if it failed we still handed back a live-looking session and logged 'consumer started' (silent failure = false audit coverage). Split the two Win32 calls instead of adding a channel handshake (avoids any lost-wakeup/hang risk): the quick, synchronous OpenTraceW now runs on the caller thread (open_trace), and only the blocking ProcessTrace runs on the pump thread (run_trace). start_session returns Err if OpenTraceW fails (reclaiming the boxed Sender so the consumer disconnects, stopping the session, joining the consumer) and returns Ok/logs 'started' only once capture is genuinely open. Opened handle + LoggerName buffer + boxed Sender are carried to the pump via a Send OpenedTrace so they outlive ProcessTrace. Box-verified on 7F203-MXC-001: consumer started=True, failed-to-start=False, 50 OCSF rows / 50 JSONL, BuffersLost=0 (no regression to capture/emit). Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(mxc-etw): guard pending-event replay against PID recycling CodeRabbit flagged that drain_resolved() re-resolved buffered events against the live by_pid map, so if Windows recycled a wxc-exec PID within PENDING_TTL a stale event from the dead sandbox could be emitted under the new owner. Stamp each by_pid registration with its Instant and add resolve_replay(), used only on the buffered/replay path. It (a) never falls back to the recycle-/ambiguity-prone by_cmd or last_pid_sid keys, and (b) trusts a PID match only when the registration is not newer than the buffered event by more than REPLAY_PID_GRACE (2s) - a recycled PID's registration lands well outside that window, so the stale event ages out instead of misattributing. The legitimate #2 seed race (registration lands ~immediately) still replays. Adds unit tests for the recycle-refusal, in-grace acceptance, and weak-fallback exclusion. Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(mxc-etw): surface unexpected ProcessTrace termination (review #4) start_session already returns Err on OpenTraceW failure (runs on the caller thread since e41a7701), closing the first half of Shailendra's #4. This closes the second half: ProcessTrace's result was discarded, so if capture died mid-run the backend had no way to know. Add a shared CaptureHealth (stopped/stopping/exit_code) between the pump thread and EtwSession. run_trace now records ProcessTrace's WIN32_ERROR and, when the pump returns without a deliberate stop, logs at ERROR that MXC OCSF capture is no longer running. EtwSession::stop() sets `stopping` before teardown so a normal shutdown isn't misreported, and EtwSession::is_capture_alive() exposes the state for status/diagnostics. Box-verified on 7F203-MXC-001: 5 sandboxes, 50 attributed OCSF rows, JSONL parity 50/50, BuffersLost=0, clean start/stop (no false failure). Signed-off-by: Akber Raza <akberr@nvidia.com> * feat(mxc-ocsf): add ETW->OCSF audit-trail example kit; fix proxy-configured message Add a runnable OCSF audit-trail example under examples/ (run-ocsf-audit.ps1, mxc-ocsf-audit.toml, ocsf-audit.yaml, README) that spins up sandboxes with the in-process ETW consumer and egress proxy on, emitting a full OCSF JSONL audit trail across all four classes (6002/5019/1007/2004). Fix SandboxProxyConfigured mapping to log "MXC sandbox proxy configured" instead of a misleading "(no proxy)" when the provider reports proxyPort=0; the event's presence already indicates proxy configuration. Verified on-box: 26 events, all mapped ETW event types present. Signed-off-by: Akber Raza <akberr@nvidia.com> * feat(mxc-ocsf): clearer audit report + client-safe run-ocsf-audit.ps1 Improve the ETW to OCSF audit-trail example output and make it safe to ship. Report: - Add an event-type coverage count ("N of M expected event types fired"); the denominator auto-adjusts (8 with proxy on, 7 with -NoProxy). - Split the checklist into expected event types vs anomaly findings (ActivityError/FallbackError), which are reported separately and not counted toward coverage (a clean run may emit none). - Verdict is now coverage-based (all expected types must fire) instead of the looser "at least 3 OCSF classes". - Call out the absolute path to the durable OCSF JSONL log prominently. Client-safety: - Default -ShareOut to empty (no auto-copy); pass -ShareOut a UNC path to opt in. Removes a hardcoded internal share path from a published example. - Drop internal-team wording ("Hand that zip back for evaluation", "BUNDLE:") in favor of neutral "Results bundle:". - Update README-ocsf-audit.txt to match the opt-in -ShareOut behavior. Verified on both MXC boxes: 7F203-MXC-001 (base-container) -> PASS, 8 of 8 event types, 26 OCSF events across 4 classes; 7F203-MXC-003 (AppContainer fallback) -> reduced set as expected, clean output. Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(mxc): configure OCSF audit workloads per sandbox - remove unsupported gateway-scoped workload fields from the shipped MXC audit example. - build the command, working directory, and filesystem grant from each run's ShareDir - pass the workload through --driver-config-json. - preserve the host CONNECT proxy configuration and conditional audit coverage for the future host_connect_proxy merge - require the workload output when determining the audit verdict. Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(mxc): omit command arguments from OCSF audit logs - record only the executable basename for MXC CreateProcessInSandbox audit events - leave process.cmd_line unset so workload arguments cannot reach shorthand or JSONL logs - cover tokens, passwords, signed URLs, and PII with a secret-leak regression test - update the audit example, architecture guidance, and published logging documentation - preserve ETW attribution and future host_connect_proxy enforcement behavior Signed-off-by: Akber Raza <akberr@nvidia.com> * feat(etw): enhance ETW session management with distinct naming for concurrent gateways * fix(etw): bound the audit queue during overload - replace the unbounded ETW callback channel with count- and byte-bounded buffering - keep the ETW callback non-blocking and count records rejected during overload - emit immediate, rate-limited warnings that identify resulting audit coverage gaps - make the audit example fail when queue overload causes dropped ETW records - cover stalled consumers, oversized events, and warning throttling with unit tests Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(etw): harden sandbox audit attribution - remove command-line and persistent per-PID fallback keys from live and replay resolution - retire the driver-owned wxc-exec PID before publishing child completion - retain established identity, activity, and correlation-vector links only for the five-second late-event window - prevent buffered records from crossing rapid PID retirement and reuse boundaries - add resolver and lifecycle coverage and document the attribution trust boundary Signed-off-by: Akber Raza <akberr@nvidia.com> * chore(mxc): address rebase follow-ups Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(mxc): align OCSF audit example with driver config - remove unsupported egress proxy settings - stop requiring the unavailable proxy audit event - update example documentation for supported event coverage Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(etw): redact command-line secrets in DecodedEtwEvent summary * fix(etw): enhance PID resolution and event attribution logic for ETW records * fix(ocsf): restrict gateway-local JSONL sink to Windows/MXC path with opt-in configuration * address rebase issues * fix(mxc): address ETW audit review feedback Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(mxc): fail closed across ambiguous PID reuse Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(mxc): bind ETW attribution to process generation Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Akber Raza <akberr@nvidia.com> Signed-off-by: Drew Newberry <anewberry@nvidia.com> Co-authored-by: Jamie King <jamiek@nvidia.com> Co-authored-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
38f2aef930 |
feat(gateway): support selective compute driver builds (#3118)
* feat(gateway): support selective compute driver builds Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(gateway): support selective Windows MXC builds Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
33bbda3d33 |
refactor(persistence): adopt continuation-token pagination (#3249)
* refactor(persistence): adopt continuation-token pagination Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(pagination): address continuation review findings Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(tui): recover completed list refreshes Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(pagination): address review scalability findings Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(pagination): repair branch validation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(go): use page size in template example Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
ddc8bba967 |
ci(windows): add Windows MSVC CI jobs (#2738)
* fix(ci): preserve Windows Rust build cache Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(ci): invalidate empty Windows caches Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * perf(ci): cache Windows builds with sccache Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * ci(windows): restore target directory caching Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * perf(ci): use prebuilt Z3 on Windows Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * perf(ci): layer sccache on Windows target cache Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * ci(windows): split PR checks from main validation Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * ci(windows): separate checks builds and cache seeding Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(ci): simplify Windows build dependency Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(ci): rely on Windows job dependency status Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(ci): use valid opt-in Windows ARM runner Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * ci(windows): keep ARM64 validation local Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(ci): install Clippy for Windows validation Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * ci(windows): focus platform lint coverage Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(licenses): explain bzip2 allowance Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * ci(windows): simplify workflow name Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(windows): allow async platform stub Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(network): make file fingerprints portable Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * ci(windows): lint supported deliverables Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * ci(windows): allow platform-gated lint Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * chore(ci): align Windows cache action with main Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(ci): align Windows validation with prerequisites Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * ci(windows): pin Rust toolchain action Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(ci): use enterprise-approved Windows actions Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(windows): restore strict MSVC validation Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * ci(windows): run Rust tests with nextest Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(ci): normalize nextest lock provenance Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * ci(windows): add native arm64 validation Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(ci): lock nextest for Windows ARM64 Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(windows): resolve duplicate MXC authentication method Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * test(conformance): use native absolute paths on Windows Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(windows): address MSVC review feedback Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * ci(windows): simplify cache key names Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(ci): isolate Windows Rust toolchains for stable caches Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(ci): configure Rustup home in runner setup Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * ci(windows): surface sccache server write diagnostics Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * ci(windows): remove temporary cache diagnostics Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * ci(windows): address review feedback Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(windows): reconcile merged driver capabilities Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(deps): preserve AWS-LC-only lockfile Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * chore(deps): allow z3 prebuilt TLS wrapper Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> --------- Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> |
||
|
|
25021ee31d |
fix(docker): reclaim sandbox token files on out-of-band removal (#3220)
Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> |
||
|
|
0357daee31 |
refactor(proto): isolate gateway storage messages (#3169)
* refactor(proto): isolate gateway storage messages Move persistence-only protobufs into a server-private versioned package, remove them from generated public SDKs, and gate durable/public schema compatibility with legacy database fixtures. Closes #3053 Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> * docs(gateway): sync protobuf schema inventory Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> --------- Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> |
||
|
|
67374efdf8 |
fix(ocsf): emit schema-valid event identities (#3247)
Previously, every event from a sandbox reused the sandbox ID as its event ID. Consumers deduplicating security records could mistake separate events for the same record, and missing device types or empty image objects could prevent schema validation. Give each event its own ID, retain the sandbox association separately, and classify the environment as Other/Sandbox while keeping the OS separate. Omit unknown container details instead of emitting empty objects. Security tooling can now distinguish events from the same sandbox and read their identity consistently after serialization. Refs #1055 Signed-off-by: Kris Hicks <khicks@nvidia.com> |
||
|
|
f4dc6be4b2 |
refactor(inference): remove managed inference routes (#3195)
* refactor(inference): remove managed inference routes Closes #3172 Remove the inference route control plane, inference.local data path, built-in router crate, and SDK surface. Move inference workloads to explicitly imported provider profiles and native endpoints, with migration cleanup and updated tests and documentation. Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(policy): preserve alternate upstream isolation Restore the provider policy activation guard so legacy OpenAI and Anthropic providers configured for alternate base URLs do not grant egress to the built-in public vendor endpoints. Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> --------- Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> |
||
|
|
48c449d8c8 |
chore(deps): replace ring with AWS-LC (#3243)
* chore(deps): replace ring with AWS-LC Signed-off-by: Simon Scatton <sscatton@nvidia.com> * fix(lint): address warnings after dependency upgrades Signed-off-by: Simon Scatton <sscatton@nvidia.com> * fix(tls): limit provider initialization to reqwest clients Signed-off-by: Simon Scatton <sscatton@nvidia.com> --------- Signed-off-by: Simon Scatton <sscatton@nvidia.com> |
||
|
|
6e6b3c8905 |
refactor(cli): remove local Dockerfile image builds (#3214)
Signed-off-by: Evie Howard <evhoward@redhat.com> |
||
|
|
e4369adcd0 |
chore(deps): replace serde_yml with noyalib (#3031)
* chore(deps): replace serde_yml with noyalib Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(providers): annotate generic YAML test values Signed-off-by: Evan Lezar <elezar@nvidia.com> --------- Signed-off-by: Evan Lezar <elezar@nvidia.com> |
||
|
|
17171cd933 |
refactor(otel): unify compute driver tracing (#2995)
Centralize compute-driver RPC descriptors, stream instrumentation, provider routing, and standalone installation in openshell-otel. Use typed RPC constants so gateway and in-process driver paths cannot panic on unknown operation strings or repeat runtime method parsing. Emit semantic-convention rpc.service and rpc.method attributes, preserve trace context and resource identity across deployment modes, and route both RPC boundary and backend crate spans to each selected driver provider. Leave consumer-dropped watch spans unset while recording observed terminal status, and avoid reboxing untraced external-driver streams. Derive each driver tracing identity from Cargo package and crate metadata and attach its descriptor to the compute-driver registration, keeping provider selection and target routing tied to the registered implementation. Share tracing setup and round-trip test support across Docker, Podman, Kubernetes, and VM, and update the gateway tracing documentation. Signed-off-by: Kris Hicks <khicks@nvidia.com> |
||
|
|
487b26574d |
test(conformance): add plan-driven continuity verification (#3107)
* refactor(test-guest): compose Ansible provisioner roles Signed-off-by: Evan Lezar <elezar@nvidia.com> * test(conformance): add plan-driven sandbox continuity Signed-off-by: Evan Lezar <elezar@nvidia.com> * test(test-guest): add gateway continuity actions Signed-off-by: Evan Lezar <elezar@nvidia.com> * test(test-guest): add RPM gateway reinstall action Signed-off-by: Evan Lezar <elezar@nvidia.com> * test(test-guest): add RPM gateway upgrade action Signed-off-by: Evan Lezar <elezar@nvidia.com> * test(test-guest): install latest-release RPM baseline Signed-off-by: Evan Lezar <elezar@nvidia.com> * test(test-guest): add gateway upgrade-restart plan Signed-off-by: Evan Lezar <elezar@nvidia.com> * ci(conformance): run Fedora gateway upgrade plan Signed-off-by: Evan Lezar <elezar@nvidia.com> --------- Signed-off-by: Evan Lezar <elezar@nvidia.com> |
||
|
|
8e73f1db99 |
fix(deps): remediate h2 advisory (#3085)
* fix(deps): remediate h2 advisory Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(deps): update h2 to 0.4.19 Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> --------- Signed-off-by: Evan Lezar <elezar@nvidia.com> Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> Co-authored-by: Piotr Mlocek <pmlocek@nvidia.com> |
||
|
|
857af42a16 |
feat(vm): support corporate HTTP forward proxy egress for microVM sandboxes (#3090)
* feat(vm): support corporate HTTP forward proxy egress for microVM sandboxes The corporate forward proxy machinery from #1792 is driver-agnostic and already merged: openshell-supervisor-network implements CONNECT chaining, NO_PROXY matching, credentials, https:// proxies and corporate CA trust, and openshell-sandbox exposes it as six argv-only flags. Podman gained the driver half in #2245/#2512 and Kubernetes in #2633; the VM driver had none of it, so VM sandboxes on proxy-only networks could not reach any destination requiring the proxy even when policy allowed it. The blocking piece was not proxy logic but delivery: the VM guest init script runs as PID 1 and execs a fixed supervisor command line, and libkrun's krun_set_exec receives an empty argv, so there was no channel for driver-owned supervisor arguments. The supervisor's proxy flags deliberately have no environment fallback, and build_guest_environment merges user-supplied environment, so the guest env is not a safe transport either. Add a driver-authored argument file, mirroring the existing init.d manifest: the driver writes /opt/openshell/supervisor-args into the overlay upperdir on every launch and the guest reads it verbatim, one argument per line, appending it to every supervisor exec. It is written even when empty, which is what makes the channel unforgeable -- the upperdir always shadows the read-only image layer, so an image can neither supply its own arguments nor disable the operator's by omitting the file. Because both launch backends exec the same init script, this covers libkrun and QEMU without touching either. A microVM has no bind mounts or container secrets, so the credential and CA bundle are staged into the per-sandbox overlay the way the gateway JWT already is: credential root-only at 0600, CA at 0644, both rewritten every launch so a removed setting clears prior material, and both deleted with the sandbox state directory. This places the credential at rest in the overlay image on the gateway host, which differs from the Podman secret model and is documented as an explicit security consideration. Validation is fail-closed and shared: a new openshell_core::driver_utils::validate_upstream_proxy_settings holds the pairing rules the Podman driver established, and both the gateway and the driver call it so an invalid table names the offending key instead of surfacing as an opaque driver-readiness timeout. Guest egress leaves through gvproxy, so a proxy on the gateway host's loopback is reachable only through host.openshell.internal; the guest to gateway callback is unaffected. Closes #3088 Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(vm): bound the proxy CA read and scope the host-loopback recipe Two review findings on the corporate forward proxy support for microVM sandboxes. The driver read the operator's proxy_ca_bundle with an unbounded fs::read and accepted it on a substring match for the PEM BEGIN CERTIFICATE marker. A special file such as /dev/zero therefore grew driver memory without bound on every authorized sandbox create, and a PEM block holding invalid DER passed the host check but contributes no trust anchor in the guest, so every supervisor would fail after boot with an error attributed to the sandbox rather than to the setting. Move the read into openshell-core as read_upstream_proxy_ca_bundle_file: it reuses the credential reader's bounded-read path (non-regular files rejected on fstat, size capped, read bounded even if the file grows), then requires at least one anchor that RootCertStore::add_parsable_certificates accepts. The supervisor's own reader now delegates to it, so host acceptance and guest acceptance are the same function and cannot drift. The published host-loopback recipe was written for libkrun only. gvproxy NATs host.openshell.internal to the gateway host's 127.0.0.1, but GPU sandboxes run on the QEMU/TAP backend where that name resolves to the TAP host address and the driver's own nftables input chain accepts only the gateway port from the guest — no proxy on the gateway host is reachable there at any bind address, so an operator following the generic recipe lost all proxy-required egress while configuration validation succeeded. Scope the recipe to libkrun in every reference and reject a gateway-host proxy URL when a launch plan resolves to QEMU, naming the reason, instead of booting a sandbox whose policy-approved CONNECTs all time out. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(vm): match the QEMU proxy preflight to the selected TAP host The gateway-host proxy guard added for the QEMU/TAP backend classified the wrong set of addresses in both directions. It ran at the top of configure_qemu_launch_plan, before the subnet allocation that settles plan.host_ip, so it could not compare against the address the guest actually reaches the host on. An operator pointing https_proxy at the sandbox's own TAP host address, such as 10.0.128.1, passed the check, and the driver's nftables input chain — which accepts only the gateway port from the guest — then dropped every policy-approved CONNECT, which is exactly the silent timeout the guard exists to prevent. In the other direction it rejected 192.168.127.254 unconditionally. That address is special only to libkrun/gvproxy; on QEMU/TAP it is an ordinary address that may be routable through the guest's masqueraded egress, so the guard refused a working configuration. Run the check after the launch plan's network allocation, on both the freshly-allocated and already-complete paths, and compare IP literals with that sandbox's selected TAP host. Loopback literals, localhost, and the documented host aliases that write_host_gateway_aliases seeds to the TAP host still classify as the gateway host, and the failure names the address. The gvproxy host-loopback constant returns to being a documentation anchor. Signed-off-by: Philippe Martin <phmartin@redhat.com> --------- Signed-off-by: Philippe Martin <phmartin@redhat.com> |
||
|
|
9ca19e6c80 |
refactor(compute): decouple gateway driver composition (#2823)
* refactor(compute): decouple gateway driver composition Move first-party composition and VM process ownership into openshell-gateway, leaving openshell-server backend-independent. Update packaging and build references with the new crate, simplify the compiled-driver boundary, and keep the driver-free gateway path buildable with bundled Z3 tooling. Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(telemetry): bound compute driver categories Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(core): keep runtime transport generic Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(compute): complete server driver decoupling Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(compute): preserve driver integrations after rebase Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> * fix(compute): preserve docker tracing after decoupling Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> * fix(compute): preserve driver behavior after extraction Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(compute): remove MXC policy side channel Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(compute): separate policy delivery from readiness Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> |
||
|
|
b4afcd8a43 |
fix(cli): suppress ANSI color when stdout is not a terminal (#3026)
* fix(cli): suppress ANSI color when stdout is not a terminal
The CLI colorized output unconditionally. owo-colors is built without
its `supports-colors` feature, so `.green()` and friends emitted escape
sequences regardless of destination, and nothing in the CLI read
NO_COLOR. Piping any command through grep or awk matched against bytes
the caller could not see; `forward list` was the case that surfaced it,
where an escape sits immediately before the STATUS word and defeats a
pattern anchored on whitespace.
Add a `color` module holding a process-wide switch resolved once in
run_async, before any output. Command modules import its `Colorize`
trait in place of `OwoColorize`; the method names match, so the ~450
call sites are unchanged, but each consults the switch when it renders
and delegates to owo-colors so the escape bytes stay identical. The two
traits collide by design: importing both in one module is an ambiguity
error, which keeps unconditional coloring from returning.
owo-colors is not the only styled path, and the rest each carry their
own default, so the switch governs them too:
- tracing_subscriber formats with ANSI on, does no terminal detection,
and writes to stdout, so `openshell -v ... | ...` leaked escapes the
same way the tables did. It now takes the setting via with_ansi.
- indicatif and dialoguer both style through console, which has its
own detection but cannot learn about --color. Overriding console's
global switch covers every progress bar and prompt rather than the
specific ones constructed today. Both the stdout and stderr switches
are set, since prompts and progress bars draw to stderr.
- miette renders errors through its own handler, likewise unaware of
--color, so init installs one built from the setting.
Resolution order: `--color always|never`, then NO_COLOR, then
CLICOLOR_FORCE, then whether stdout is a terminal. The decision is made
against stdout even for stderr text, since stdout is what gets parsed;
`--color always` restores styling when redirecting.
Padding is unaffected — the format spec is forwarded to the inner
Display, so widths measure text rather than text plus escapes.
Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
* fix(cli): resolve color per output stream
Review feedback on #3026.
Resolving one answer from stdout and handing it to every library meant a
redirected stream inherited the other stream's terminal check. Running
`openshell ... 2> build.log` from a terminal wrote escapes into the log,
because console's stderr switch and miette's handler were both given
stdout's answer. That is worse than the behavior before this branch,
where both libraries did their own per-stream detection.
Resolve `auto` separately for stdout and stderr and hand each library
the answer for the stream it writes to: tracing and console's stdout
switch get stdout, miette and console's stderr switch get stderr. The
owo-colors wrapper is the exception, since its call sites are split
across println! and eprintln! and a Painted value cannot tell which
macro will consume it; it styles only when both streams accept escapes,
erring toward plain text rather than risking a redirected stream.
Existing tests could not catch this: Command::output gives both streams
pipes, so a per-stream decision and a single stdout-derived one look
identical. Add a test that puts stdout on a pty and stderr on a pipe,
which fails when stderr is handed stdout's answer.
Replace CLICOLOR_FORCE with FORCE_COLOR. The clicolors spec does not say
how to treat `0`, and implementations that special-case it disagree with
force-color.org, which keys on presence and non-emptiness only. Using
FORCE_COLOR gives it the same rule as NO_COLOR: set and non-empty means
yes, whatever the value. Nothing depended on CLICOLOR_FORCE, which was
introduced earlier on this branch and never released.
Carry the whole style in an owo_colors::Style rather than dispatching a
local enum through a six-arm match, and merge styles when chaining so
`x.green().bold()` emits one `\x1b[32;1m...\x1b[0m` instead of nesting
two wrappers. No call site styles already-styled text, so merging is
safe; the emitted bytes are shorter and there is a single reset.
Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
---------
Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
|
||
|
|
bb70461878 |
test(e2e): run conformance in gateway lanes (#2925)
* test(e2e): isolate VM-specific smoke assertions Signed-off-by: Evan Lezar <elezar@nvidia.com> * test(e2e): add portable CLI conformance baseline Signed-off-by: Evan Lezar <elezar@nvidia.com> * feat(conformance): add standalone CLI runner Signed-off-by: Evan Lezar <elezar@nvidia.com> * test(e2e): run conformance in gateway lanes Signed-off-by: Evan Lezar <elezar@nvidia.com> --------- Signed-off-by: Evan Lezar <elezar@nvidia.com> |
||
|
|
8a13bc1298 |
chore(deps): remove legacy rustls webpki path (#3013)
Signed-off-by: Evan Lezar <elezar@nvidia.com> |
||
|
|
23351771ca |
chore(deps): bump quinn-proto from 0.11.14 to 0.11.17 (#2986)
Bumps [quinn-proto](https://github.com/quinn-rs/quinn) from 0.11.14 to 0.11.17. - [Release notes](https://github.com/quinn-rs/quinn/releases) - [Commits](https://github.com/quinn-rs/quinn/compare/quinn-proto-0.11.14...quinn-proto-0.11.17) --- updated-dependencies: - dependency-name: quinn-proto dependency-version: 0.11.17 dependency-type: indirect ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
bcd517bbe0 |
feat(driver-mxc): native Windows MXC compute driver + server wiring (#2721)
* feat(driver): add MXC compute driver for Windows isolation sessions Introduces the openshell-driver-mxc crate implementing ComputeDriver backed by Microsoft MXC isolation sessions (Windows only). Wires the new driver into the server's build_compute_runtime dispatch and adds the Mxc variant to ComputeDriverKind. Also adds a local protobuf-src stub (tools/protobuf-src-local) to unblock Windows builds that lack MSYS2/MinGW, and pins the zig Windows x64 toolchain in mise.lock. (cherry picked from commit 4f7012224efb18fbfeb47aa87e0cfd3f036f32f0) Signed-off-by: Jamie King <jamiek@nvidia.com> * wip(mxc): checkpoint hung-agent work (recon, policy_map embed, A1 wiring, demo artifacts) Safety checkpoint of uncommitted work from the background agent run that stalled mid-Step-7. Includes: mxc-driver-recon.md (Step 0.5), policy_map.rs (~876L embedded mapper), A1 policy-threading edits across driver.rs/policy.rs/mxc.rs/compute/mod.rs, and examples/ (demo.yaml + mxc-gateway.toml). Not yet verified to compile end-to-end; to be reorganized into the skill's Step 11 commit sequence. (cherry picked from commit 38e42c03870be3d10e984a54f17a3b61122ff510) Signed-off-by: Jamie King <jamiek@nvidia.com> * test(mxc): fix lifecycle and policy unit-test compile drift - Bring futures::StreamExt into scope for the watch-stream `.next()` call in driver::lifecycle_tests so the negative policy proof test compiles. - Bind a local `mapper` and drop the unused/deprecated NetworkBinary in the embedded-mapper network-policy rejection test. Signed-off-by: Jamie King <jamiek@nvidia.com> (cherry picked from commit 039b0baf98735ca672dae52be8d3af2417dc0c1a) Signed-off-by: Jamie King <jamiek@nvidia.com> * fix(mxc): downgrade missing sandbox_token to debug log The gateway mints `sandbox_token` only when a sandbox-JWT issuer is configured. There is no in-sandbox supervisor on MXC (supervisor-removal design — D1/D4), so no component ever consumes the token; requiring it on the driver side blocks the demo's `--disable-tls` smoke gateway with a spurious `invalid_argument`. Log the absence and proceed instead. Signed-off-by: Jamie King <jamiek@nvidia.com> (cherry picked from commit cea209797d0edcb1d152251e748900b0a63cca62) Signed-off-by: Jamie King <jamiek@nvidia.com> * fix(mxc): keep sandbox Ready after a successful one-shot agent exec monitor_exec demoted Ready->Error on exit 0 (reason ExecCompleted), so the positive demo (write hello.txt + exit) landed in Error phase. Keep Ready=True (reason AgentCompleted) on success; only non-zero exits go to ExecFailed. Tighten the positive lifecycle test to assert the terminal condition stays Ready=True/AgentCompleted. Verified live via gateway mock round-trip: phase now Provisioning->Ready with no demotion. (cherry picked from commit 54ab030f03ca0f83d0050d8dd843b633651684ad) Signed-off-by: Jamie King <jamiek@nvidia.com> * feat(mxc): add processContainer backend for default-deny enforcement Add a backend selector to the MXC driver (isolation_session default | process_container). process_container drives a one-shot AppContainer that is genuinely default-deny: a write to any ungranted path is denied by the OS, unlike isolation_session which is grant-only and cannot deny. The lifecycle forks on the flag - isolation_session keeps provision/start/exec, process_container runs a single ephemeral container via run_oneshot. Also: run-demo.ps1 gains -Backend and hardens the CLI register/create calls; docs corrected to state isolation_session does NOT deny out-of-policy writes and that the negative proof requires process_container. Verified end-to-end on a real demo box (gateway -> CLI -> driver -> MXC): in-policy write succeeds, out-of-policy write denied (PermissionDenied), OVERALL: PASS. (cherry picked from commit c6cde3860bbe1b8edb3147d3e840f6bf0ece32d8) Signed-off-by: Jamie King <jamiek@nvidia.com> * refactor(driver-mxc): embed policy mapper as a module; remove standalone crate Adopt the proto-based mapper (map_to_mxc) as the single source of truth, embedded in openshell-driver-mxc as a Windows-gated `policy_map` module. Rewire EmbeddedPolicyMapper to call it directly on the typed SandboxPolicy, deleting the serde_yaml proto->YAML bridge. Move the CLI to a windows-gated example and the parity tests into the crate; delete openshell-policy-mapper. - gate policy_map + seam Windows-only (MXC is Windows-only) - drop serde_yaml; add dev-deps openshell-policy, clap, anyhow - normalize mapped paths to Windows form in the seam, in one place - docs: add driver-mxc to AGENTS.md table; correct design doc section 17 test lane Signed-off-by: Giedrius Burachas <gburachas@nvidia.com> (cherry picked from commit f22f9c7a25b9a651c5c5cc73f62fb01c4d6c1a8d) Signed-off-by: Jamie King <jamiek@nvidia.com> * feat(driver-mxc): implement lossless split_policy for proxy-delegated egress Signed-off-by: Giedrius Burachas <gburachas@nvidia.com> (cherry picked from commit 96d6afa0e2dc6a1d54edd12c34a0ceb0a30dadd0) Signed-off-by: Jamie King <jamiek@nvidia.com> * feat(driver-mxc): implement Pattern-C governed-egress split through the policy seam - split_policy: SocketAddr proxy_redirect (replaces bare port), processcontainer containment guard naming MXC M1, version preserved in the trimmed proxy_policy, delegation reported as an info loss item - seam: MappedConfig carries trimmed_policy + proxy_addr; MapCtx.egress selects the split path; coarse path unchanged when egress is disabled - driver: [openshell.drivers.mxc] egress_proxy / egress_proxy_addr config, validated at create (isolation_session rejected until M1); lifecycle threads the redirect into provision and stores the trimmed policy per sandbox, emitting an EgressRedirect platform event - mxc: optional MxcNetwork block (defaultPolicy=block + proxy) in provision and one-shot configs; mock records configs for test assertions - tests: lossless-invariant suite over all example policies (validate + serialize round-trip), split lifecycle proof, M1 rejection; example gains --split --proxy-addr writing mxc-config.json / trimmed-policy.yaml / loss-report.json Signed-off-by: Giedrius Burachas <gburachas@nvidia.com> (cherry picked from commit 34d54ad9f25dc6034c3ba15668555ff0d22cddd8) Signed-off-by: Jamie King <jamiek@nvidia.com> * fix(driver-mxc): emit MXC network.proxy as {localhost: port} Verified against the real wxc-exec 0.6.0-alpha via --dry-run: MXC accepts only the {localhost: N} proxy shape (the form the design doc specifies) and rejects {host, port} with a parse error. Schema 0.6.0-alpha can express only a loopback port, so non-127.0.0.1 redirect addresses are now rejected: split_policy emits an error loss (no proxy block) and the driver refuses egress_proxy_addr values off 127.0.0.1. Per-sandbox attribution must use per-sandbox ports until the schema widens. Signed-off-by: Giedrius Burachas <gburachas@nvidia.com> (cherry picked from commit edde8d5434571fd5398409204fcf6862672c0793) Signed-off-by: Jamie King <jamiek@nvidia.com> * fix(driver-mxc): serialize isolation_session stop/deprovision as unit variants Empirical contract finding from the real test lane (build 26300.8553, wxc-exec 2026-06-10): the stop and deprovision experimental blocks are unit variants in the wxc-exec schema and must serialize as null; sending {} is rejected with malformed_request (invalid type: map, expected unit), while provision/start accept maps. The production invoker, the real-lane test, the probe script, and the e2e runner all sent {} - the driver could provision and run an agent but never stop or delete an isolation-session sandbox against this build. Pinned by a unit test. Signed-off-by: Giedrius Burachas <gburachas@nvidia.com> (cherry picked from commit 0df39ca0b22ebc21eb965b2b567a5b2cac26af32) Signed-off-by: Jamie King <jamiek@nvidia.com> * test(driver-mxc): add Tier-0 mapper coverage matrix with schema drift guard Three-quadrant, table-driven matrix (38 tests): mappable fields assert exact MXC output; every OpenShell field MXC cannot express asserts a loss item with the expected severity (and seam rejection on error); an empty policy asserts the restrictive default-deny posture for every MXC knob OpenShell does not control. The handled_fields_inventory drift guard serializes a fully-populated policy and compares its YAML keys against the mapper-handled field lists, so a new openshell-policy field fails the suite until consciously mapped, delegated, or reported as loss. Re-exports the policy seam types for integration tests; adds serde_yml, base64, serde_json as dev-dependencies. Signed-off-by: Giedrius Burachas <gburachas@nvidia.com> (cherry picked from commit 91807f984a3b16846e35d6ca0d5ec41057aafa3a) Signed-off-by: Jamie King <jamiek@nvidia.com> * feat(driver-mxc): inject agent_env into sandbox process.env Add MxcComputeConfig.agent_env: each entry is either KEY=VALUE (verbatim) or a bare KEY resolved from the gateway host environment at launch, keeping secrets (e.g. inference API keys) out of the config file. Wire it into the agent process so gateway-launched agents can authenticate to cloud endpoints (process.env was previously hardcoded empty). Unit-tested via resolve_agent_env_passthrough_and_host_lookup. Also add a gateway-driven cloud-inference (T1) test harness: mxc-inference.toml (agent_env + curl agent), inference.yaml policy, and run-inference-test.ps1 which starts the gateway, creates an isolation_session sandbox, runs an authenticated Nemotron call, and bundles redacted results. Documented agent_env in mxc-gateway.toml. Validated end-to-end on the test box (chat HTTP 200 + completion via the gateway). (cherry picked from commit 94d9e827b8b77e0af9dc654943ea8e7d8400cc9f) Signed-off-by: Jamie King <jamiek@nvidia.com> * test(driver-mxc): avoid unsafe env mutation in resolve_agent_env test Replace std::env::{set,remove}_var (unsafe + racy under parallel test execution in edition 2024) with a read-only PATH lookup. Preserves all three behaviors under test and drops the #[allow(unsafe_code)]. (cherry picked from commit ac5766eeab2db4e8cc6fcd8d8a97809edaf3df30) Signed-off-by: Jamie King <jamiek@nvidia.com> * fix(mxc): adapt MXC driver to current GitHub OpenShell API The MXC driver crate was authored on GitLab against an earlier proto/core API. Adapt it to the API on GitHub main: - build_capabilities_response no longer takes supports_interactive_session - DriverSandboxSpec.gpu (bool) is now resource_requirements; detect GPU via effective_driver_gpu_count(driver_gpu_requirements(..)) - DriverSandbox gained a `workspace` field - SandboxPolicy gained `network_middlewares`: pass it through the proxy split, emit a loss item on the coarse MXC path, and account for it in the mapper drift-guard test Verified: cargo check + 75 mock-based tests pass (lib 27, examples 10, policy_mapper_matrix 38). Signed-off-by: Jamie King <jamiek@nvidia.com> * feat(server): wire the MXC compute driver into the gateway on Windows Register openshell-driver-mxc as the Windows-only in-process compute backend so compute_driver = "mxc" resolves to a working runtime: - ComputeRuntime::new_mxc, adapted to the current 11-arg from_driver - mxc_policy_sink A1 side channel, staged in create_sandbox before dispatch - mxc_config_from_context loader and the Mxc dispatch arm (Windows constructs; other targets return an explicit "Windows-only" error) - Windows-gated openshell-driver-mxc dependency - Mxc arms for the telemetry, config-file required-fields, and CLI reserved-builtin matches to keep them exhaustive/correct Verified with cargo check --workspace --features openshell-prover/bundled-z3 on x86_64-pc-windows-msvc, stacked on PR #2496. Signed-off-by: Jamie King <jamiek@nvidia.com> * fix(driver-mxc): implement GetGatewayListenerRequirements for #2496 base Signed-off-by: Jamie King <jamiek@nvidia.com> * test(driver-mxc): add probe-gated real wxc-exec test lane (no mocks) - tests/wxc_exec_real.rs: ignored-by-default integration tests against a real wxc-exec. Six --dry-run contract tests run wherever the binary exists (they caught the network.proxy shape mismatch); enforcement tests (processcontainer default-deny positive/negative, isolation session lifecycle round trip with a deprovision drop-guard) probe the backend and SKIP with a recorded reason where it is not live. - examples/probe-mxc-host.ps1: classifies a host (OS build, --probe, per-backend trial) and emits a JSON capability verdict. - examples/run-mxc-e2e.ps1 + e2e-policies/: scenario runner generalizing run-demo.ps1 (fs-rw, fs-readonly, fs-default-deny-empty, network-policy-rejected) with PASS/FAIL/SKIP gating and a stale OPENSHELL_MXC_MOCK_WXC guard in real mode. - tasks/windows.toml: windows:test:mxc-real:x64, windows:e2e:mxc, windows:e2e:mxc:mock. Signed-off-by: Giedrius Burachas <gburachas@nvidia.com> (cherry picked from commit 49afafe892caded59c4df50a9b652011ad97f41c) Signed-off-by: Jamie King <jamiek@nvidia.com> * fix(driver-mxc): address PR review feedback Signed-off-by: Shailendra Singh <shailendras@nvidia.com> * docs: defer public MXC documentation Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> * fix(driver-mxc): build Windows capabilities response Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> * fix(server): gate in-tree tracing on Windows Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> * test(windows): fix cross-platform test assumptions Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> * test(windows): verify process and PEM portably Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> * test(windows): keep lifecycle command in policy Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> --------- Signed-off-by: Jamie King <jamiek@nvidia.com> Signed-off-by: Giedrius Burachas <gburachas@nvidia.com> Signed-off-by: Shailendra Singh <shailendras@nvidia.com> Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> Co-authored-by: Prashant K <pkhodade@nvidia.com> Co-authored-by: Giedrius Burachas <gburachas@nvidia.com> Co-authored-by: Shailendra Singh <shailendras@nvidia.com> Co-authored-by: Drew Newberry <385+drew@users.noreply.github.com> |
||
|
|
d0dfb22baf |
feat(kubernetes): export driver traces over OTLP (#2958)
Mirror the VM, Podman, and Docker driver tracing setup for Kubernetes. Export standalone driver spans through OTLP/gRPC as the distinct openshell-driver-kubernetes service, preserve gateway trace context, record lifecycle operations and gRPC failures, and flush spans on shutdown. Kubernetes currently runs in-process when selected as a built-in gateway driver. Use the temporary server-boundary shim shared with Podman and Docker so traces retain the shape they will have when Kubernetes moves to a separate process. Move the common ComputeDriver RPC tracing layer into openshell-otel to keep all drivers aligned. Propagate the active W3C context through the controller-reserved Sandbox annotation and enable Agent Sandbox OTLP export in the local k3s workflow. This connects asynchronous controller reconciliation spans to the originating OpenShell create trace. Expose gateway OTLP configuration through Helm and add an Aspire collector to the local k3s workflow. Extend helm:k3s:forward with OTLP ingest and trace UI forwarding for Kubernetes and local container gateway development. Signed-off-by: Kris Hicks <khicks@nvidia.com> |
||
|
|
4e992093f8 |
fix(docker): trace standalone driver over OTLP (#2923)
Signed-off-by: Evan Lezar <elezar@nvidia.com> |
||
|
|
e457974a52 |
feat(docker): export driver traces over OTLP (#2851)
Mirror the VM and Podman driver tracing setup for Docker. Export Docker driver spans through OTLP/gRPC as the distinct openshell-driver-docker service, preserve gateway trace context, record lifecycle and asynchronous provisioning operations, and report gRPC failures. Docker currently runs in-process when selected as a built-in gateway driver. Add the same temporary server-boundary shim used by Podman so traces retain the shape they will have when Docker moves to a separate process. Generalize the gateway provider selection for both in-process drivers and share the OTLP collector fixture across Docker, Podman, and VM tracing tests. Signed-off-by: Kris Hicks <khicks@nvidia.com> |
||
|
|
905e99aa2a |
feat(build): add Nix-native Linux toolchains (#2875)
* feat(nix): add glibc 2.28 development shell * feat(nix): add musl development shell * feat(build): use mold in musl development shell * feat(flake): add nix remote cache * feat(build): use mold in default development shell |
||
|
|
40d1b48666 |
feat(provider): support for SPIFFE backed token exchange (#1970)
* feat(provider): add ability to request token exchange instead of client credentials as OAuth grant_type Signed-off-by: Gordon Sim <gsim@redhat.com> * test(proxy): add further tests for token exchange Signed-off-by: Gordon Sim <gsim@redhat.com> * test(provider): add runnable example for token exchange Signed-off-by: Gordon Sim <gsim@redhat.com> * test(e2e): cover Podman token exchange grants Signed-off-by: Gordon Sim <gsim@redhat.com> * refactor(oauth): extract duplicated functionality from server and supervisor Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(provider): evict nearest-to-expiry entry from intermediate token cache Signed-off-by: Gordon Sim <gsim@redhat.com> * doc(supervisor): add podman example for token exchange Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(provider): withhold token-exchange subject credentials Signed-off-by: Gordon Sim <gsim@redhat.com> --------- Signed-off-by: Gordon Sim <gsim@redhat.com> |