mirror of
https://github.com/NVIDIA/OpenShell.git
synced 2026-10-02 07:34:45 +08:00
main
193
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
021400be8a |
refactor(auth): separate sandbox identity from TLS (#3110)
* refactor(auth): separate sandbox identity from TLS Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(auth): clarify gateway mTLS behavior Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(auth): include workspace scope in TLS authorization checks Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(e2e): bound service auth sandbox names for large PIDs Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
252882f37f |
feat(providers): serve sandbox config files on demand (#3832)
* feat(providers): serve sandbox config files on demand Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(providers): defer managed file documentation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(providers): mark managed file api experimental Signed-off-by: Drew Newberry <anewberry@nvidia.com> * style(go): format provider profile fields Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(providers): preserve legacy environment with managed files Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
1358941b81 |
feat(mcp): inspect requests with Tower-selected protocol profiles (#3335)
* fix(sandbox-backend): sort boundary request objects before hashing Sort boundary request objects recursively before hashing so serde_json's preserve_order feature cannot change digest identity. Cover canonical bytes, envelope round trips, and rejection of modified provider values and operations. Signed-off-by: Shiju <shiju@nvidia.com> * feat(mcp): upgrade tower-mcp-types to 0.22.2 Upgrade tower-mcp-types from 0.12.0 to an exact-pinned 0.22.2 and use its inspection APIs to validate MCP requests against the selected revision. Carry inspection metadata into policy evaluation and validate requests after header rewriting, before forwarding. Add explicit support for the sessionless 2026-07-28 revision while keeping 2025-11-25 as the default. Validate per-request metadata and standard HTTP header mirrors, and support discovery, tools, and subscription requests. Delegate batch availability and parameter schemas to Tower. Share typed request names between policy and HTTP checks, retain the local batch resource cap, and centralize MCP policy version parsing and ordering. Keep supported MCP revisions and shared allowlist parsing in the canonical policy schema; core re-exports those types. Tower owns wire-profile semantics, and every supported policy revision must map to the matching inspector profile. Reject duplicate JSON keys, invalid known-method parameters, unavailable methods, and unsupported batches. Keep exact extension allow rules and deny precedence. Document request inspection boundaries and add unit, forwarding, and sandbox coverage. Refs #2174. Signed-off-by: Shiju <shiju@nvidia.com> * test(mcp): prove authorization at the forwarding boundary Cover March batch denial in both member orders, valid and malformed controls, and audit behavior across both relay entry paths. Exercise real middleware tool rewrites with matching metadata and assert the exact upstream representation or zero forwarded bytes. Verify legacy bodyless SSE GET remains usable while GET tool bodies and unsupported DELETE cleanup are rejected. Clarify request-selected profile and middleware mutation comments without changing production behavior. Signed-off-by: Shiju <shiju@nvidia.com> * test(mcp): exercise permitted profiles through the sandbox proxy Cover March and June singleton policies and select November and July separately under one endpoint allowlist. Capture upstream tool receipts to distinguish proxy policy denial from an upstream rejection. Extend middleware rewrite coverage to June and multi-version policies, and preserve the sessionless discovery and subscription checks through the shared fixture helpers. Signed-off-by: Shiju <shiju@nvidia.com> * test(kubernetes): box the admission check future Keep the admission test future below Clippy's size limit when the workspace dependency features are unified. Signed-off-by: Shiju <shiju@nvidia.com> * test(mcp): reuse the forwarding fixture identity cache Share the binary identity cache across protocol-profile cases, matching the proxy lifecycle and avoiding repeated hashes of the test executable. Keep procfs authorization and all forwarding assertions intact. Signed-off-by: Shiju <shiju@nvidia.com> --------- Signed-off-by: Shiju <shiju@nvidia.com> |
||
|
|
d4f5034d7f |
fix(gateway)!: make the WebSocket tunnel opt-in (#3727)
The /_ws_tunnel endpoint pipes a WebSocket into the full gRPC service. On a plaintext loopback gateway any web page could open it, since browsers do not apply CORS to WebSocket upgrades. Mount the tunnel only when enable_websocket_tunnel is set (config file, --enable-websocket-tunnel, or OPENSHELL_ENABLE_WEBSOCKET_TUNNEL; server.enableWebsocketTunnel in Helm). BREAKING CHANGE: gateways behind an authenticating edge proxy must set enable_websocket_tunnel = true for CLI edge-tunnel connections. Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
5448fb4f98 |
fix(auth): skip renewal for non-expiring sandbox JWTs (#3686)
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> |
||
|
|
0518bd4c83 |
fix(sandbox): preserve local sessions across host sleep (#3573)
* fix(cli): recover sandbox connect transport Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(auth): propagate non-expiring local sandbox sessions Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(cli): distinguish main exit from transport loss Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(cli): bound sandbox connect recovery Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
a8f98ec09d |
fix(auth): remove legacy sandbox JWT admission (#3562)
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> |
||
|
|
1e34e8c576 |
fix(drivers): require admission labels for external resources (#3538)
* fix(drivers): require admission labels for external resources Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(drivers): address resource admission review findings Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(core): reserve driver-owned admission labels Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(core): clarify workspace admission label Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(drivers): clarify resource admission failures Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(e2e): configure resource admission fixtures Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): retry forbidden admission lookups Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(e2e): preserve external driver admission defaults Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
35e0a68e4a |
feat(kubernetes): support corporate proxy CA bundle (#3447)
* feat(kubernetes): support corporate proxy CA bundle The Kubernetes driver had no way to supply a CA bundle for the corporate egress proxy, so an `https://` proxy with a private CA, or a TLS-intercepting proxy, could not be used. Podman and VM already expose `proxy_ca_bundle`. Add `proxy_ca_bundle` to `[openshell.drivers.kubernetes]` as a path the gateway Pod reads. The gateway stages the PEM into the existing per-generation supervisor bootstrap Secret and passes `--upstream-proxy-ca-bundle` on the supervisor argv. That Secret is already immutable, owner-referenced and garbage-collected, and its volume mounts every key at /.openshell/supervisor with no items filter, so this needs no new object kind, volume, mount, or RBAC verb, and works in shared, managed and operator workspace modes. The bundle is deliberately read from the gateway's filesystem rather than referenced as an object in the sandbox namespace. It becomes a trust anchor for every upstream the sandbox reaches, so it must stay in the gateway's trust domain; the immutable staging Secret also keeps the anchor from changing underneath a running sandbox. Bound the staged bundle at 256 KiB. The shared reader's limit is exactly the apiserver's own Secret limit and the bootstrap Secret carries four other keys, so a bundle between the two would pass gateway startup and then fail every sandbox create with an opaque `data: Too long`. Delegate the URL, no_proxy, connect_by_hostname and ca_bundle rules to the shared validate_upstream_proxy_settings, keeping the Secret-specific credential block local: this driver accepts an explicit `proxy_auth_allow_insecure = false` without credentials, which the shared rules reject. This also fixes the acknowledgement being demanded for an `https://` proxy, where the credential travels inside the verified TLS session. Add auth_setting_label so the inline-credential diagnostic names the Secret keys instead of proxy_auth_file, which this driver rejects as an unknown key. Document that the bundle should carry only the CA that signs the proxy's certificate, or that an intercepting proxy re-signs upstream certificates with. Public roots already reach the sandbox through the supervisor image and its TLS stack, and the bundle is concatenated with that system store into a single boundary control frame, so a full merged trust bundle spends the frame budget on duplicated roots. The frame, not the apiserver Secret limit, is the tighter of the two ceilings in practice; raising the staging bound requires checking it. Closes #3443 Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(helm): quote proxy CA ConfigMap references Signed-off-by: Philippe Martin <phmartin@redhat.com> --------- Signed-off-by: Philippe Martin <phmartin@redhat.com> |
||
|
|
50230616d5 |
refactor(runtime): retire Community image dependencies (#3386)
* feat(sandbox): default to official Alpine sandbox image default_sandbox_image() now returns docker.io/library/alpine:3.22, a generic version-qualified official image, so a fresh install no longer depends on the community sandbox image catalog. All compute drivers (docker, podman, kubernetes, vm) inherit this fallback. Part of #3116. Signed-off-by: Akram Signed-off-by: Akram <akram.benaissi@gmail.com> * feat(deploy): default deployment configs to the official Alpine sandbox image Update the shared gateway default_image, Helm chart values, the standalone Kubernetes manifest, and the dev gateway task scripts to use docker.io/library/alpine:3.22 instead of the community base image, consistent with default_sandbox_image(). GPU e2e image-build base is left unchanged (CUDA needs a glibc base). Part of #3116. Signed-off-by: Akram Signed-off-by: Akram <akram.benaissi@gmail.com> * feat(driver): default to numeric non-root identity for USER-less images With the default sandbox image now Alpine, images that declare no OCI USER must start instead of being rejected. When the image declares no USER and the policy requests none, the Podman and Docker drivers now supply a numeric non-root identity (DEFAULT_SANDBOX_UID/GID = 1000) instead of rejecting, matching the numeric-identity behavior of the Kubernetes and VM drivers. The supervisor's resolved-identity path runs the sandbox as a synthesized non-root account without the account existing in the image. Images that declare a USER keep the OCI resolution path unchanged. Part of #3116. Signed-off-by: Akram <akram.benaissi@gmail.com> Signed-off-by: Evan Lezar <elezar@nvidia.com> * test(conformance): use Alpine workload image Signed-off-by: Evan Lezar <elezar@nvidia.com> * refactor(policy): drop community image /app path from default policy The restrictive default policy granted read-only access to /app, a directory that only existed in the community base image. A generic Alpine default has no /app, so remove it. Landlock best-effort already ignores absent paths; this just stops advertising a community-specific layout in the default. Part of #3116. Signed-off-by: Akram Signed-off-by: Akram <akram.benaissi@gmail.com> * docs(config): document Alpine default images Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(podman): report early sandbox termination Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(podman): initialize rootless workspace ownership Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(sandbox): qualify NVIDIA Ubuntu default Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): initialize rootful default workspace Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(sftp): add native sandbox adapter Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sftp): gate runtime helper support to Linux Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sftp): support standard OpenSSH file operations Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sftp): harden rename and special file handling Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(runtime): remove community image dependencies Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(e2e): build provider readiness tool fixture Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(e2e): use a dedicated Noble fixture for Docker tests Signed-off-by: Evan Lezar <elezar@nvidia.com> --------- Signed-off-by: Akram Signed-off-by: Akram <akram.benaissi@gmail.com> Signed-off-by: Evan Lezar <elezar@nvidia.com> Signed-off-by: Drew Newberry <anewberry@nvidia.com> Co-authored-by: Evan Lezar <elezar@nvidia.com> Co-authored-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
2493d415c2 |
feat(extensions)!: normalize protocol negotiation (#3352)
* feat(extensions)!: normalize protocol negotiation Closes #3057 Introduce a shared extension handshake, enforce protocol and capability compatibility across extension families, and expose immutable negotiated snapshots through gateway info and the Go SDK. Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(credentials): fail fast on negotiation errors Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(extensions): validate gateway handshake metadata Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(go-sdk): re-export extension kind constants Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(extensions): fail fast on credential handshake rejection Signed-off-by: Seth Jennings <sjenning@redhat.com> --------- Signed-off-by: Seth Jennings <sjenning@redhat.com> |
||
|
|
17ce738bfb |
fix(ci)!: remove gateway callback listener dependency (#3365)
* fix(ci): repair post-merge release canary Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(packaging): bootstrap canary runtime prerequisites Signed-off-by: Drew Newberry <anewberry@nvidia.com> * ci(canary): collect macOS VM diagnostics Signed-off-by: Drew Newberry <anewberry@nvidia.com> * ci(canary): pin libkrun-compatible macOS runner Signed-off-by: Drew Newberry <anewberry@nvidia.com> * ci(canary): limit macOS smoke test to package startup Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(compute)!: remove gateway callback listeners Run Docker supervisors on host networking so they use the operator-configured primary gateway endpoint. Remove the unused compute-driver callback listener negotiation and listener-scoped routing machinery. BREAKING CHANGE: The ComputeDriver API no longer exposes GetGatewayListenerRequirements or GatewayListenerRequirement. External drivers must regenerate bindings and connect supervisors to the configured primary gateway endpoint. Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(docker): use sandbox runtime image in launcher Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(podman): exercise production endpoint selection Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(e2e): route supervisors to reachable gateways Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(e2e): align Podman endpoint fixtures Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): preserve host aliases for supervisors Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): align sandbox host gateway pin Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(e2e): address Docker fixtures by bridge IP Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(e2e): serialize sandbox lifecycle cases Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(e2e): host Docker TCP fixture with gateway Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(e2e): use loopback for host-network supervisor Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
d91b1999a0 |
feat(api)!: use sandbox names as canonical RPC references (#3272)
* feat(api)!: use sandbox names as canonical references Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(cli): update forward color fixture for workspace scope Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(supervisor): use sandbox names for settings lookup Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(e2e): use canonical sandbox request fields Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(e2e): use canonical sandbox receipt field Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(server): harden sandbox mutation handling Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(ci): update rebased sandbox references Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(api)!: standardize canonical entity references Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(api): codify protobuf API conventions Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(api): preserve workspace selector semantics Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(api): restore workspace selector parity Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(api): preserve descriptive name fields Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(api): update e2e request fixtures Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(core): omit workspace selector during bootstrap Signed-off-by: Drew Newberry <anewberry@nvidia.com> * chore(api): remove proto convention checker Signed-off-by: Drew Newberry <anewberry@nvidia.com> * chore(api): refresh schema fingerprints after rebase Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(cli): use canonical provider receipt field Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
1d010f4187 |
feat(sandbox): validate configuration before workload activation (#3259)
* feat(sandbox): validate configuration before workload activation Closes #3145 Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(sandbox): bound startup failures and preserve activation history Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(sandbox): enforce deadlines on startup RPC attempts Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(sandbox): isolate provider auto-create policy fixture Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(sandbox): remove unnecessary fixture string delimiters Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(sandbox): capture startup logs before ephemeral cleanup Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(sandbox): preserve credential revocation after rebase Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(server): retain provider revision helper for endpoint reports Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(server): import provider object trait in production Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(supervisor): preserve admission across boundary extraction Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(sandbox): use existing boundary discovery imports Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(sandbox): distinguish workload and supervisor containers Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(server): reconcile admission schema with timestamp migration Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(server): reconcile admission with typed deletion schema Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(server): reconcile admission with mutation request IDs Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(supervisor): adapt local startup fixture to admission state Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(supervisor): deduplicate startup quarantine diagnostics Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(tui): show configuration blockers in sandbox notes Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * feat(sandbox): expire provisioning repair attempts after five minutes Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(supervisor): reconcile admission with provider readiness Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(tui): separate configuration summaries from full diagnostics Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(tui): shorten invalid configuration note Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(server): refresh schema inventory after rebase Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> --------- Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> |
||
|
|
04146692d9 |
refactor(providers)!: make provider profiles import-only (#3383)
* chore(providers): remove dead provider plugin modules Twelve modules under crates/openshell-providers/src/providers/ were never declared in providers/mod.rs, so they have not been compiled since the plugin registry was narrowed to the two adapters it still registers. Four of them (generic, gitlab, opencode, outlook) key off provider type identifiers that normalize_provider_type already retired. Keep google_cloud and vertex, which are the only plugins ProviderRegistry::new registers. Signed-off-by: Philippe Martin <phmartin@redhat.com> * test(providers): load example profiles from providers/ at test time Add an example_profiles module that reads the YAML under providers/ from the source checkout at run time and parses it with the existing profile loader. It is gated behind a new non-default example-profiles feature so it is available to this crate's own tests and, once wired into dev-dependencies, to gateway and CLI tests, while never reaching a release binary. Nothing consumes it yet; later changes move the test fixtures off the compiled catalog and onto these files. Signed-off-by: Philippe Martin <phmartin@redhat.com> * test: source profile fixtures from providers/ instead of the compiled catalog Unit and integration fixtures reached into builtin_profiles() to get a profile to work with, which ties the tests to the compiled catalog rather than to the files an operator would import. Point them at the example_profiles loader instead. The profiles.rs tests keep their coverage unchanged and become explicit golden tests over providers/*.yaml: they are what keeps those files valid once nothing compiles them. builtin_profiles_are_sorted_by_id widens into a test that the whole example set parses, sorts, and lints clean as one catalog. The CLI fake gateways serve the example profiles the way a real gateway serves what an operator imported, through a shared helper. No behavior change: the gateway still loads the built-in source by default and serves the same profiles. Signed-off-by: Philippe Martin <phmartin@redhat.com> * docs(providers): document the example provider profiles Each file under providers/ now opens with a header naming its expected client binary identities, the image layout those paths assume, the credential scope, the endpoint access it grants, and a smoke test. Add a README covering the import commands and why a profile should be copied and edited rather than imported unchanged. Several of these profiles bind network access to paths that only exist in the OpenShell Community image — /sandbox/.venv, /app/.venv, /sandbox/.cursor-server, /usr/lib/node_modules. Imported unchanged into another image the profile matches nothing: the catalog still advertises it, but the credential is never injected and the traffic is denied. The headers say so where it applies. Comments only; the profile schema has no documentation fields and all fifteen files still lint clean. Signed-off-by: Philippe Martin <phmartin@redhat.com> * test(server): seed unit-test state with the example provider profiles Server unit tests inherited the builtin + user source default from ServerState, so around a hundred and forty assertions about github, openai and the rest resolved against the compiled catalog. Point the test state at the user source alone and import the example profiles from providers/ into its store first, the way an operator would. Every one of those assertions keeps passing unchanged, which is the point: it proves the gateway behaves identically with an imported catalog before the default moves. Four tests asserted builtin-source semantics specifically. Under import-only the only profiles a gateway cannot edit are the ones a non-user source vends, so they now exercise a source-managed profile composed with the user source; the read-only guard they cover is the one that still applies to interceptor catalogs. The list test asserts the imported profile's user/platform identity instead of builtin with an empty scope. test_server_state_with_user_only_github_profile is gone: the default test state now is a user-only gateway with github imported. Signed-off-by: Philippe Martin <phmartin@redhat.com> * refactor(cli): resolve credential suggestions from the gateway catalog The --env credential warning scanned a profile table compiled into the CLI, so its suggestions described the binary rather than the gateway the user is talking to: a profile the gateway does not serve was suggested anyway, and an imported custom profile never was. Fetch the catalog from the connected gateway instead. The warning moves out of argument parsing and into sandbox create and sandbox template create, where a client already exists. A catalog fetch failure is not fatal — the warning degrades to its generic form rather than blocking sandbox creation. Extract the ListProviderProfiles paging loop from provider list-profiles into a shared fetch_provider_profile_catalog, which the profile-driven paths now share. Signed-off-by: Philippe Martin <phmartin@redhat.com> * refactor(cli)!: infer providers from the gateway catalog Command-to-provider inference went through a hardcoded alias table in openshell-providers, a second copy of the built-in catalog's identifiers. A custom profile could never be inferred no matter what binaries it declared, and the table drifted from the profiles it mirrored. Infer from the connected gateway's catalog instead: match the command's basename against each profile's ID and against the basenames of the binaries the profile authorizes. A profile that names /usr/bin/claude is the profile for running claude. The match must be unique — where several profiles claim a command, the user names one with --provider — and an empty catalog infers nothing. No command is special-cased. `binaries` is the operator's authorization statement, so a profile that declares a binary claims the command that runs it, whatever that binary is; narrowing that belongs in the profile rather than in a list compiled into the CLI, which could never cover an unbounded catalog anyway. Breaking: the retired aliases stop resolving, and commands the old table never listed can now infer. git, pip and uv are declared by the github and pypi example profiles, so they infer where they previously did not. Signed-off-by: Philippe Martin <phmartin@redhat.com> * refactor(providers,server)!: resolve provider profiles by exact ID normalize_provider_type was a hardcoded alias map — gh to github, claude to claude-code, vertex to google-vertex-ai — and a second copy of the built-in catalog's identifiers, independent of the YAML it mirrored. It let a provider type resolve to a profile the operator never named, and it made the built-in IDs behave as a reserved namespace. Remove it, along with detect_provider_from_command and the alias-normalizing ProviderRegistry::inject_env. A provider type now names a profile exactly: - the effective catalog resolves an ID or reports it absent, with no alias retry - plugins activate only for the ID of a profile the gateway resolved; a provider with no resolvable profile gets no plugin projection, instead of falling back to an alias guess - the CLI surfaces the gateway's not-found instead of retrying under an alias - telemetry buckets by profile ID, and the gitlab, opencode and outlook buckets go with the aliases that were their only source Breaking: `--type gh`, `--type claude` and the other aliases no longer resolve. Use the profile's own ID. Signed-off-by: Philippe Martin <phmartin@redhat.com> * feat(server): fail closed when a provider's profile is absent A sandbox composed from a provider whose profile the gateway cannot resolve started anyway: the credential and policy builders warned and skipped, so the sandbox came up carrying none of that provider's credentials or network policy. The operator learned about it later, as a denied connection or a missing environment variable, rather than as the configuration error it is. Check at the two composition boundaries — CreateSandbox and AttachSandboxProvider — right beside the catalog snapshot already taken there, and reject with a bounded diagnostic naming the provider, the profile it refers to, and the import command that supplies it. Scope-aware: a platform-scoped provider is told to import with --global. Read paths are untouched. ListProviders, GetProvider and profile export keep working so an operator can see and recover an affected provider, and the shared policy builders keep their warn-and-skip for the diagnostic paths that also reach them. Signed-off-by: Philippe Martin <phmartin@redhat.com> * refactor(server): scope the vendor base-URL pin to declared endpoints provider_profile_endpoints_are_active withheld a credential and the provider's policy layer when an openai or anthropic provider pointed its client somewhere other than the public vendor endpoint. The guard was keyed on profile.source == "builtin" and on those two profile IDs, so it protected only profiles OpenShell shipped. Once profiles are import-only no profile is ever builtin, and the guard would silently stop applying — including to an operator who imported providers/openai.yaml verbatim. Key it on the profile instead of on where the profile came from. A profile's endpoints are the boundary its credential is bound to, so if the provider configures a *_BASE_URL pointing at a host the profile does not declare, the profile no longer describes where that credential goes and is treated as endpointless. Host matching reuses the DNS-label-aware matcher in openshell-core, so wildcard endpoints such as Vertex's *-aiplatform.googleapis.com resolve correctly. A profile with no declared endpoints has no boundary to contradict, and a config value that names no host is not a redirect this can reason about. Both keep the profile active. The control now covers every endpoint-bearing profile, including an operator's own, and the last "builtin" string leaves the gateway. Signed-off-by: Philippe Martin <phmartin@redhat.com> * test(e2e): import example provider profiles during gateway bring-up The e2e suites create providers from github, openai, nvidia, claude-code and google-cloud, and the GitHub lane runs a real git clone through the profile's binary attribution. A gateway serves only the profiles an operator imported, so the lanes have to import them. Add e2e_import_example_provider_profiles to the shared bring-up helpers and call it from the Docker, Podman, Kubernetes and VM wrappers once the gateway is healthy and registered. It runs the same command the upgrade notes give operators, against the repository's own providers/ directory, so the lanes exercise the documented path rather than a test-only shortcut. Lands before the default changes: a same-ID user profile already shadows a built-in, so importing works today. Signed-off-by: Philippe Martin <phmartin@redhat.com> * refactor(providers)!: make provider profiles import-only OpenShell compiled fifteen provider profile YAML files into every release binary and selected them by default, so a fresh gateway published a catalog it was never configured with. Those profiles are not image-neutral: their binary selectors name paths that exist in the OpenShell Community image, so changing the sandbox image could make a profile inert while the catalog still advertised it. Every endpoint, credential name and binary path in providers/ was also effectively part of the 0.1.0 public contract. A gateway's catalog is now exactly what an operator imported: - providers/*.yaml is no longer include_str!'d, and builtin_profiles() is gone from the openshell-providers API - the default provider_profile_sources is [{ type = "user" }], and a gateway with nothing imported reaches ready and serves an empty catalog — an empty catalog is a valid state, not a startup failure - the builtin source type is removed from the configuration schema, and a gateway.toml that still names it is rejected at parse time with the import command rather than an unknown-variant error - nothing reserves the canonical identifiers any more, so github, pypi, anthropic and the rest import at their own IDs and the imported profile is the only definition for that ID; the static_fallback that kept a shipped definition resident behind an imported one is gone with the source that produced it A collision between an interceptor-vended profile and an imported one still fails closed, which is the behavior that shadowing quietly bypassed for built-ins. Operators upgrading should export the profiles their deployment relies on before upgrading, or copy them from the providers/ directory of the matching release tag, then import them at the scope their providers use. Signed-off-by: Philippe Martin <phmartin@redhat.com> * docs(providers): document import-only provider profiles Rewrite the provider profile documentation around a catalog the operator builds rather than one the gateway ships. - gateway-config: the default is [{ type = "user" }], the builtin source type is gone and rejected at startup, and an empty catalog is a valid ready state - profiles: replace the built-in profile table and its shadowing semantics with the import workflow, and say plainly that a profile whose binary paths do not match the image is inert - manage-providers: replace the two fixed provider-type tables with `provider list-profiles`, and describe catalog-driven command inference - inference-routing: drop the "still loads built-in profiles for compatibility" transition text; its migration walkthrough is now the normal path - quickstart and the Docker Compose, GitHub, AWS, Google Cloud and Vertex tutorials: import the profile before creating the provider, since nothing resolves without it - release notes: a 0.1.0 migration section covering export-before-upgrade, importing from the release tag's providers/ directory, what happens to a provider whose profile is missing, and the removal of the legacy type aliases - README, architecture, the openshell-cli and debug-openshell-cluster skills, and the governance interceptor example follow the same change; the example's smoke assertion now imports a profile to prove an authoritative interceptor hides it Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(cli,e2e): distinguish catalog failures and authenticate OIDC seeding Addresses two findings from review (GATOR-c1bd9867-02 and -03). The provider profile catalog lookup collapsed its error into an empty catalog, so an authorization, availability or transport failure on ListProviderProfiles was indistinguishable from a gateway that genuinely has no profiles. Command inference then resolved nothing and the sandbox was created without the provider it needed, deferring the failure to the workload. Keep the lookup's outcome instead of discarding it. The credential warning is advisory and still degrades to its generic form, but inference now consults the catalog only when there is a command to resolve and surfaces the lookup failure when there is, naming the gateway and pointing at explicit --provider selection. Two regression tests cover it through a fake gateway whose ListProviderProfiles returns UNAVAILABLE while every other RPC succeeds: sandbox creation fails without sending a provider-less create request, and a sandbox with no trailing command still succeeds because it needs no catalog. The e2e profile seeding also ran unauthenticated in the OIDC lanes. Those lanes deliberately skip gateway registration and start the gateway without a TLS client CA, so no mTLS identity exists and no token has been acquired when the import runs; the wrapper exited during setup. Skipping the import is not sufficient because the provider tests now require the claude-code profile. Add e2e_register_oidc_admin_session, which mints an administrator token with Keycloak's password grant — the same grant the OIDC test helpers use — and writes the gateway metadata and token bundle that an interactive login would have stored, so the import runs as an authenticated administrator. The mTLS lanes keep the direct import unchanged. Both affected wrappers are covered: with-podman-gateway.sh had the same defect as with-docker-gateway.sh. Signed-off-by: Philippe Martin <phmartin@redhat.com> * refactor(cli)!: remove command-derived provider attachment Addresses GATOR-c1bd9867-01. A profile's `binaries` list authorizes a binary to reach that profile's endpoints. It is not a statement that running the binary asks for the provider, and reading it as attachment intent let a command silently gain provider authority: the aws-s3 example declares /bin/bash, so `sandbox create -- bash` resolved an existing aws-s3 provider and attached its credential-backed capability and network policy to the shell and its descendants. Attachment of an already-created provider never prompted, so the confirmation flow did not guard it. The distinction that would make inference sound — whether a declared binary is a profile's client or merely a permitted runtime — cannot be expressed: NetworkBinary carries only a path, and the field that encoded it was removed in 0.1.0. Any substitute is a guess. Restricting the guess to a unique claimant does not help, because uniqueness measures how sparse the catalog is rather than what the user intended, and a compiled list of "generic" commands could never cover an unbounded operator catalog. Remove trailing-command inference rather than approximate it. A provider is attached only when named with --provider, which still creates a missing provider from local discovery when the name matches an imported profile ID. Sandboxes with no providers remain a normal, fully supported state. Removing inference also settles GATOR-c1bd9867-02: with no consumer deriving authority from the catalog, its only remaining use is the advisory credential warning, so a failed lookup degrades that warning instead of blocking creation. The regression test now asserts that an unreachable catalog still creates the sandbox and attaches nothing. While repurposing the deduplication test, a pre-existing defect surfaced: repeating a name in --provider auto-created it twice and the second attempt failed with "provider already exists", because the explicit pass never consulted the set of names it had already handled. Guard it. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(e2e): load the OIDC token and trust the gateway when seeding profiles Addresses the carried finding GATOR-c1bd9867-03. e2e_register_oidc_admin_session established a session the CLI never used. The CLI decides whether to load a stored bearer token by matching on the auth_mode field of the gateway metadata alone; the helper omitted that field, so the metadata fell through to the default arm and oidc_token.json stayed on disk unread. The profile import went out unauthenticated exactly as it had before the helper existed. The helper also installed no trust anchor, leaving the CLI unable to verify the gateway's self-signed serving certificate. Write auth_mode = "oidc" so the stored token is loaded, and install the CA at <gateway>/mtls/ca.crt. Only the CA is installed: with no client certificate or key on disk the CLI falls back to CA-only server verification and authenticates with the bearer token, which is what these lanes need because they start the gateway without --tls-client-ca. Certificate verification stays on; no insecure transport override is introduced. Assert the session before anything depends on it. ListProviderProfiles is annotated auth_mode: "bearer", so it cannot succeed unless the token was loaded and accepted. The helper now fails at that point, naming the gateway config directory and echoing the CLI output, rather than letting the defect surface later as an opaque profile import error. Both affected wrappers pass the PKI directory and CLI binary the helper needs; the mTLS lanes are untouched. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(e2e): request the OpenShell scopes when minting the admin token The OIDC lanes failed at the session assertion added in b25dd0298 with PERMISSION_DENIED and "scope 'provider:read' required". Authentication was working -- the gateway logged ListProviderProfiles at gRPC status 7, which it can only reach once the bearer token has been loaded and accepted. The token simply carried no OpenShell scope. sandbox:*, provider:*, config:*, workspace:* and openshell:all are optional client scopes on the openshell-cli client in scripts/keycloak-realm.json, so Keycloak mints them only when the request asks for them. The password grant here asked for nothing, leaving the realm defaults (openid, profile, email, roles, web-origins, acr) and an access token that authorizes no RPC. Request "openid openshell:all", as e2e/python/oidc/oidc_auth_test.py already does for its administrator tokens. openshell:all is SCOPE_ALL in crates/openshell-server/src/auth/authz.rs, so one scope covers the setup calls without enumerating them. Verified against the realm: the token goes from "email profile openid" to "email profile openshell:all openid". Record the same scopes in metadata.json. The helper writes the bundle openshell gateway login would have stored, and oidc_scopes is the field that login path reads back, so leaving it out would misdescribe the stored token. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(e2e): seed provider profiles once per VM lane gateway The VM lanes failed on the second test target with "custom provider profile 'anthropic' already exists" for all fifteen example profiles, and the import exited non-zero. e2e_import_example_provider_profiles was called from run_e2e_test, so it ran once per target -- four times against one long-lived gateway. That was harmless while import overwrote silently, but profiles are now import-only: ImportProviderProfiles is create-only and reports an existing id as an error-severity diagnostic, with no overwrite flag on the request. The first import therefore succeeds and every later one fails. Hoist the call to just after the conformance run, which is where the docker, podman and kube lanes already seed their catalogs. The profiles persist for the gateway's lifetime, so every target still finds them, and the E2E_TEST_OVERRIDE path is covered by the same single call. Signed-off-by: Philippe Martin <phmartin@redhat.com> * test(e2e): name the provider explicitly in the OIDC workspace-user test user_can_create_sandbox_with_inferred_provider_command reached the gateway once the admin session was fixed, and then failed: it asserted a missing-provider error, but the sandbox was created and died provisioning with "failed to spawn sandbox entrypoint process 'claude-code'". The test drove provider resolution by passing claude-code as the trailing command and relying on the CLI to infer the provider type from it. That inference is what this branch removed, so the trailing word is now nothing but an entrypoint, and the image has no such binary. The regression the test guards is not inference itself: it is that a workspace user resolving a provider is not gated behind Platform Admin. Name the provider with --provider, the only remaining way to attach one. The CLI still has to fetch the claude-code profile before it can auto-create the provider, so the lookup a workspace user must be allowed to make still happens, and auto-creation still stops at the non-interactive branch with "missing required provider". Both assertions therefore keep their meaning. Rename the test and rework its comments to describe what it now exercises; the old name would otherwise outlive the behavior it was named for. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(providers): reconcile import-only profiles with main Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> --------- Signed-off-by: Philippe Martin <phmartin@redhat.com> Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> Co-authored-by: John Myers <9696606+johntmyers@users.noreply.github.com> |
||
|
|
9708ba9999 |
feat(providers): report applied sandbox provider changes (#3391)
* feat(providers): report applied sandbox provider changes Record exact provider mutation targets in shared configuration operations. Require authenticated evidence that credentials, effective policy, and the workload launch environment have been installed before reporting readiness. Add bounded CLI and Rust SDK status and wait support, preserving ordinary revision-scoped references for existing processes. Verify new-client rotation and acknowledged detach revocation without external-stable resolver changes. Signed-off-by: Shiju <shiju@nvidia.com> * fix(providers): align readiness times with protobuf contracts Represent readiness receipts, status, and operation times with Timestamp and report intervals with Duration. Reserve the scalar field tags, update all consumers and generated bindings, and preserve timestamp presence and nanosecond identity through storage and client validation. Qualify both empty-map constructors in the Linux boundary test so its module compiles while retaining the explicit default required by Clippy. Signed-off-by: Shiju <shiju@nvidia.com> * fix(cli): preserve provider mutation storage uncertainty Recognize the gateway's exact structured storage-uncertainty reason for provider attach, detach, and update. Explain that the change may already be saved and must be reconciled before retrying, without exposing server messages or metadata. Preserve uncertainty ahead of generic retry hints. Exercise saved mutations through the CLI and verify single submission, redaction, missing receipt handling, and untrusted error-detail rejection. Document the recovery guidance for users and the public CLI skill. Signed-off-by: Shiju <shiju@nvidia.com> * fix(cli): explain denied provider profile lookups Report exact and alias profile lookup denials with fixed permission and workspace guidance. Keep backend details redacted and stop before provider mutations. Cover denied create and update calls through the CLI. Verify the complete provider list independently in the cross-workspace OIDC regression, extracting its JSON object from surrounding startup diagnostics. Signed-off-by: Shiju <shiju@nvidia.com> --------- Signed-off-by: Shiju <shiju@nvidia.com> |
||
|
|
d68b7069c3 |
refactor(proto)!: use well-known time types (#3113)
* refactor(proto)!: use well-known time types Signed-off-by: Derek Carr <decarr@redhat.com> * fix(proto): preserve time migration behavior Signed-off-by: Derek Carr <decarr@redhat.com> * fix(proto): preserve timestamp boundary semantics Signed-off-by: Derek Carr <decarr@redhat.com> * fix(proto): convert sandbox token expiry to timestamp Signed-off-by: Derek Carr <decarr@redhat.com> * fix(proto): preserve time compatibility semantics Signed-off-by: Derek Carr <decarr@redhat.com> * fix(proto): preserve exact endpoint and profile times Signed-off-by: Derek Carr <decarr@redhat.com> * test(e2e): use duration for interactive exec timeout Signed-off-by: Derek Carr <decarr@redhat.com> * fix(sdk-go)!: remove legacy profile duration fields BREAKING CHANGE: Go provider profile callers must use RefreshBefore, MaxLifetime, and CacheTTL with ProfileDuration instead of the whole-second fields. Signed-off-by: Derek Carr <decarr@redhat.com> --------- Signed-off-by: Derek Carr <decarr@redhat.com> |
||
|
|
2ccef97769 |
feat(policy): establish one canonical authored policy representation (#3334)
* chore(policy): restart schema implementation Signed-off-by: Johnny Greco <jogreco@nvidia.com> * feat(policy-schema): add canonical authored policy model Signed-off-by: Johnny Greco <jogreco@nvidia.com> * refactor(policy): use canonical authored schema Signed-off-by: Johnny Greco <jogreco@nvidia.com> * refactor(prover): project canonical policy documents Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): describe shared schema boundary Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(policy): close reviewed parser gaps Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(policy-schema): fail closed on unsupported fields Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(policy): preserve partial process identities Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(policy-schema): harden authored policy inspection Signed-off-by: Johnny Greco <jogreco@nvidia.com> * refactor(policy): rename policy schema crate Signed-off-by: Johnny Greco <jogreco@nvidia.com> * revert(policy): restore policy schema crate Signed-off-by: Johnny Greco <jogreco@nvidia.com> * test(e2e): serialize OIDC PKCE scenarios Signed-off-by: Johnny Greco <jogreco@nvidia.com> --------- Signed-off-by: Johnny Greco <jogreco@nvidia.com> |
||
|
|
c1f2e7189f |
feat(isolation): implement the RFC 0012 sandbox architecture (#2942)
* feat(isolation): add RFC 0012 backend contract Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> * refactor(isolation): name the interface crate explicitly Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): expose trusted host gateway Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(agents): inventory the MXC driver Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): add mediated DNS transport Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): tighten interface error and digest contracts Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(isolation): remove unrelated driver inventory Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): define capability-free launch contract Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): seal confirmed boundary state Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): validate confirmation for external backend implementations Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): clarify mediated DNS identity Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(isolation): generalize loopback connector Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(isolation): unify typed network mediation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): bind launches to sandbox sessions Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(mxc): initialize extended sandbox status Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): add boundary protocol and Linux primitives Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): harden signals and separate process status from transport Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): validate remote confirmation through public contract Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): validate wire state and propagate snapshot failures Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(isolation): import owned agent specification explicitly Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(isolation): describe mediated DNS channel Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): bound mediation attach without nested retries Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(isolation): generalize loopback protocol Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): add transport-neutral session authentication Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(isolation): separate sandbox backend protocol Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): harden runtime boundary controls Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): add terminal boundary operation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): split supervisor and sandbox runtimes Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): harden boundary isolation and lifecycle ownership Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): reject private root redirects and adopt typed errors Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): preserve accept thread ownership on musl Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(sandbox): isolate credential probes from filtered threads Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): return retained exec exit status to independent waiters Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): bound network mediation and preserve socket authorization Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): bound control admission and retire stale mediation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * ci(e2e): select migrated drivers per stack layer Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(sandbox): implement loopback connector Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): authenticate the Sandbox Protocol Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(supervisor): rotate launch-scoped authentication Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(sandbox): consume dedicated backend crate Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(sandbox): align topology session fixture Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): align projected bootstrap bundle Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(auth): validate refreshed credentials before rotation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): fail closed across supervisor disconnects Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): repair rebased sandbox CI Signed-off-by: Drew Newberry <anewberry@nvidia.com> * build(runtime): publish separate sandbox and supervisor images Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(config): configure the sandbox runtime image Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(ci): validate sandbox binary linkage Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(isolation): use backend and runtime terminology Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(sandbox): use a scratch runtime image Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(ci): refresh schema and dependency policy Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): bind reconnects to supervisor process Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs: align runtime split operational guidance Signed-off-by: Drew Newberry <anewberry@nvidia.com> * chore(security): document Kubernetes runtime RBAC Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): enforce runtime lifecycle invariants Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(compute): identify sandbox start generations Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(server): restore sandbox launch sessions Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): support authenticated runtime replacement Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(auth): bind sandbox session successors Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(auth): retry pending sandbox successors Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(vm): run the supervisor outside the guest workload Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(vm): use sandbox backend protocol Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): repair rebase integration Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): use unified build toolchain Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(vm): use sandbox runtime terminology Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(vm): own guest network bootstrap Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): expose guest init version Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): select native supervisor artifacts Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): guard guest init Linux symbols Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): scope Linux test imports Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): avoid guest interface casts Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): reconcile admitted sandbox identity Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): share resolved sandbox identity Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): surface host supervisor failures Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): include guest logs on supervisor exit Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): rotate and clean runtime generations Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): make sandbox starts generation-aware Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): rotate restored sandbox sessions Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(vm): keep shared paths in the base layer Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): bind sandbox session lineage Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(docker): isolate workloads behind the host supervisor Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(docker): rotate launch-scoped authentication Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(docker): use sandbox backend protocol Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(docker): use host networking for supervisor Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): preserve host gateway alias resolution Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(docker): use separate sandbox and supervisor images Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): restore startup validation after rebase Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(docker): name the sandbox runtime directly Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): narrow supervisor CA runtime storage Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): close companion isolation gaps Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(docker): align mediated network expectations Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(docker): exercise mediated network paths Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): attach supervisor to managed network Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): defer supervisor recovery until gateway is ready Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): make sandbox starts generation-aware Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): rotate restored sandbox sessions Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): preserve workloads during session rotation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(docker): remove unrelated configuration RFC changes Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): bind sandbox session lineage Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(kubernetes): add proxy-pod isolation topology Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(kubernetes): use sandbox backend protocol Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): use stable sandbox service authority Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(kubernetes): split sandbox and supervisor images Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): adapt proxy pods to current runtime APIs Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(kubernetes): describe the single runtime placement Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(kubernetes): simplify sandbox orchestration Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): validate deployment prerequisites Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): update Trivy Helm profile inventory Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(kubernetes): update Trivy scan inventory count Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): reuse preloaded runtime images in e2e Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): type and clean runtime resources Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): make sandbox restarts recoverable Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): rotate restored sandbox sessions Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): preserve supervisor egress Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): bind sandbox session lineage Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(podman): adopt isolated sandbox and supervisor containers Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): stage bootstrap archives at named volume destinations Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(podman): rotate launch-scoped authentication Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(podman): use sandbox backend protocol Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(podman): use host networking for supervisor Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(podman): split sandbox and supervisor images Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): repair rebase integration Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(podman): name the sandbox runtime directly Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): provision supervisor CA runtime storage Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): address isolation review findings Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): inspect Debian supervisor provenance Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): use libpod-compatible tmpfs options Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): bind verified sandbox runtime binary Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): provide external driver data directory Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): start sandbox before joining user namespace Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): separate supervisor user namespace Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): make sandbox starts generation-aware Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): rotate restored sandbox sessions Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): bind sandbox session lineage Signed-off-by: Drew Newberry <anewberry@nvidia.com> * perf(isolation): add TCP and DNS benchmark harnesses Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(perf): align benchmark timing and supported protocols Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(perf): report TCP benchmark metrics accurately Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(perf): cancel failed worker startup Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): build matching local supervisor image Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): make local sandbox smoke test runnable Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): wire local sandbox runtime image Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): narrow sandbox service RBAC Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(ci): validate split runtime artifacts Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): harden runtime session handling Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(supervisor): add standalone network proxy role Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(rfc): remove implementation companion notes Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(vm): standardize runtime release name Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): pin renamed runtime artifacts Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(auth): persist sandbox runtime identity Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(runtime): restore branch validation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): reconcile main after rebase Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(network): close unframed HTTP 1.0 responses Signed-off-by: Drew Newberry <anewberry@nvidia.com> * chore(isolation): preserve upstream OCSF updates Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(security): close credential and TLS replay paths Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(auth): make sandbox refresh retries idempotent Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
39cf4823f7 |
feat(api): add structured gateway errors and SDK decoding (#3313)
* feat(api): expose structured gateway errors across SDKs Refs #3051. Add standard validation, conflict, and retry details; preserve raw transport status in Rust, Go, TypeScript, and Python; document status and recovery guidance. This is the structured-error foundation only. Mutation result shapes, allow_missing, durable request deduplication, and exec retry semantics remain follow-up work. Signed-off-by: Mrunal Patel <mrunalp@gmail.com> * fix(python): preserve wrapped RPC cleanup handling Inspect the original gRPC call when handling missing sandboxes during deletion waits and managed cleanup. Add intercepted cleanup regressions and clarify the error-wrapper migration contract. Addresses the cleanup review on #3313; part of #3051. Signed-off-by: Mrunal Patel <mrunalp@gmail.com> --------- Signed-off-by: Mrunal Patel <mrunalp@gmail.com> |
||
|
|
b799fccb8b |
fix(auth): harden OIDC trust root retrieval (#3332)
* fix(auth): harden OIDC trust root retrieval Signed-off-by: Mrunal Patel <mrunalp@gmail.com> * fix(e2e): pass OIDC HTTP acknowledgement value Signed-off-by: Mrunal Patel <mrunalp@gmail.com> --------- Signed-off-by: Mrunal Patel <mrunalp@gmail.com> |
||
|
|
fd3fd9cf74 |
feat(sandbox): explain failed calls to external tool servers (#3207)
Show configured tool server addresses and their last observed connection results together in sandbox status. Keep sandbox lifecycle readiness separate so an external connection failure does not mark the sandbox unready. Expose direct endpoint records through the CLI and SDKs, with plain-language failure explanations and gateway acceptance times. Keep observation tracking, runtime reporting, and gateway validation in dedicated endpoint status modules. Preserve bounded reporting, request attribution, retry ordering, and configuration and supervisor authority checks. Clear obsolete observations while retaining the configured addresses, and document the distinction between an observed HTTP response, current availability, and tool success. Signed-off-by: Shiju <shiju@nvidia.com> |
||
|
|
5b57f0d154 |
fix(ci): restore Windows test portability (#3288)
* fix(ci): restore Windows test portability Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(tasks): skip Unix lockfile check on Windows Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(tasks): use buf shim on Windows Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> --------- Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> |
||
|
|
0803c4aa4c |
refactor(policy)!: remove NetworkBinary harness field (#3222)
* refactor(policy)!: remove NetworkBinary harness field Closes #3054 Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> * test(policy): preserve unknown fields during migration Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> * fix(policy): preserve advisor provenance during merge Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> * fix(policy): remove redundant network binary default Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(policy): record harness schema migration Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> --------- Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> Co-authored-by: John Myers <9696606+johntmyers@users.noreply.github.com> |
||
|
|
02b664bb0d |
refactor(config): normalize and enforce gateway schema v2 (#2814)
* refactor(config): normalize compute driver field names Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * refactor(config): introduce canonical gateway fields Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * refactor(config): enforce gateway schema version 2 Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): preserve compute driver runtime guarantees Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): address schema v2 review regressions Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): complete schema v2 migration safeguards Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(config): expand schema v2 regression coverage Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(config): add schema v2 parity manifest Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): correct parity manifest inventory Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * docs(config): record schema v2 intentional changes Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * docs(config): disposition schema v2 parity gaps Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): add dual schema parity harness Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): establish compute lifecycle parity baseline Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): preserve gateway option compatibility Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): record gateway option parity Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * docs(config): close gateway-wide parity gaps Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(podman): apply configured pids limit Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): validate Podman option parity Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): add Kubernetes option parity harness Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): record Kubernetes option parity Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): disposition VM parity lanes Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): add external driver parity lane Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(e2e): preserve external driver pull policy Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): attest parity artifacts and launches Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): require clean parity build sources Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): bind parity runtime artifacts Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(e2e): use isolated supervisor tags Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(e2e): qualify parity image tags Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(e2e): serve parity supervisor locally Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): isolate parity podman services Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): harden parity evidence provenance Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): pin parity sandbox artifacts Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): attest parity runtime inputs Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): bind parity runtime evidence Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): record compute boundary parity Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): disposition cross-cutting parity lanes Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(packaging): preflight gateway config upgrades Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): preserve rebase integration guarantees Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(ci): isolate temporary git signing config Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): update remaining schema v2 consumers Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(ci): provide e2fs tools to VM tests Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): align preflight with gateway startup Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(vm): preserve rootfs tar configuration Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * chore(config): adopt duration unit constructors Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(packaging): preflight RPM gateway config Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): address driver review findings Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(e2e): require fresh semantic parity evidence Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(docker): update tests for renamed sandbox label Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(gateway): preserve selective driver coverage after rebase Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> Signed-off-by: Drew Newberry <anewberry@nvidia.com> Co-authored-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
33bbda3d33 |
refactor(persistence): adopt continuation-token pagination (#3249)
* refactor(persistence): adopt continuation-token pagination Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(pagination): address continuation review findings Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(tui): recover completed list refreshes Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(pagination): address review scalability findings Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(pagination): repair branch validation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(go): use page size in template example Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
226a83b323 |
fix(supervisor): classify credential placeholders in request bodies (#3246)
* fix(supervisor): classify credential placeholders in request bodies Closes #2904 Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * fix(supervisor): preserve same-provider placeholders in request bodies Signed-off-by: John Myers <johntmyers@users.noreply.github.com> --------- Signed-off-by: John Myers <johntmyers@users.noreply.github.com> Co-authored-by: John Myers <johntmyers@users.noreply.github.com> |
||
|
|
90dbe5454b |
feat(api): add typed workspace selectors (#3245)
* feat(api)!: add typed workspace selectors Signed-off-by: Mrunal Patel <mrunalp@gmail.com> * fix(cli): preserve template workspace metadata Signed-off-by: Mrunal Patel <mrunalp@gmail.com> * test(e2e): migrate workspace request selectors Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(api): update public schema inventory Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> --------- Signed-off-by: Mrunal Patel <mrunalp@gmail.com> Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> Co-authored-by: John Myers <9696606+johntmyers@users.noreply.github.com> |
||
|
|
0357daee31 |
refactor(proto): isolate gateway storage messages (#3169)
* refactor(proto): isolate gateway storage messages Move persistence-only protobufs into a server-private versioned package, remove them from generated public SDKs, and gate durable/public schema compatibility with legacy database fixtures. Closes #3053 Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> * docs(gateway): sync protobuf schema inventory Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> --------- Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> |
||
|
|
f4dc6be4b2 |
refactor(inference): remove managed inference routes (#3195)
* refactor(inference): remove managed inference routes Closes #3172 Remove the inference route control plane, inference.local data path, built-in router crate, and SDK surface. Move inference workloads to explicitly imported provider profiles and native endpoints, with migration cleanup and updated tests and documentation. Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(policy): preserve alternate upstream isolation Restore the provider policy activation guard so legacy OpenAI and Anthropic providers configured for alternate base URLs do not grant egress to the built-in public vendor endpoints. Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> --------- Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> |
||
|
|
7f4bd49a47 |
fix(policy): harden landlock.compatibility validation (#2541)
* fix(policy): reject invalid landlock.compatibility values at parse time Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * fix(policy): abort sandbox startup when hard_requirement has no filesystem paths Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * fix(policy): validate landlock.compatibility at gateway and fix zero-path logging Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * fix(policy): reject invalid landlock.compatibility on serialization Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * fix: fixed linting error Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> --------- Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> |
||
|
|
48c449d8c8 |
chore(deps): replace ring with AWS-LC (#3243)
* chore(deps): replace ring with AWS-LC Signed-off-by: Simon Scatton <sscatton@nvidia.com> * fix(lint): address warnings after dependency upgrades Signed-off-by: Simon Scatton <sscatton@nvidia.com> * fix(tls): limit provider initialization to reqwest clients Signed-off-by: Simon Scatton <sscatton@nvidia.com> --------- Signed-off-by: Simon Scatton <sscatton@nvidia.com> |
||
|
|
6e6b3c8905 |
refactor(cli): remove local Dockerfile image builds (#3214)
Signed-off-by: Evie Howard <evhoward@redhat.com> |
||
|
|
592df3e014 |
feat(policy): preserve exact MCP revision allowlists (#3027)
* feat(mcp): add version-aware wire profile metadata Signed-off-by: Shiju <shiju@nvidia.com> * feat(policy): canonicalize MCP version allowlists Signed-off-by: Shiju <shiju@nvidia.com> * fix(policy): align MCP policy tests with current main Signed-off-by: Shiju <shiju@nvidia.com> * fix(policy): canonicalize supervisor protobuf ingress Materialize defaultable MCP revisions before ambiguity checks, OPA construction, and sidecar delivery. Reject invalid sidecar policies with bounded errors. Signed-off-by: Shiju <shiju@nvidia.com> --------- Signed-off-by: Shiju <shiju@nvidia.com> |
||
|
|
039b265096 |
feat(middleware): define HTTP response pre-return interface (#3073)
* feat(middleware): define HTTP response pre-return interface Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(middleware): clarify HTTP response interface Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * refactor(middleware): align response result actions Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * refactor(middleware): expose response reason codes Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * refactor(middleware): share session end reasons Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * feat(middleware)!: finalize HTTP response pre-return contract Replace the separate body_end event with HttpResponseBodyUnit.end_of_stream. Every body-inspecting stage receives exactly one flagged unit, which may be empty; a zero-byte body is one empty flagged unit and OpenShell never reads ahead to set the flag. Defer response trailers from V1 and reserve their field numbers. HTTP/1.0 clients and Content-Length bodies cannot carry trailers and that behavior was undefined. Add HttpResponsePreflight.permitted_body_modes, computed once from the original upstream head so every stage sees the same list, and make an unlisted selection a failure rather than a downgrade. Add the block_delivery preflight action as a successful decision enforced regardless of on_error. Expose Content-Length, Content-Encoding, and Content-Range read-only in preflight. Cap STREAM_BYTES input units at half of max_payload_bytes and permit deferring bytes across replacements only for fail_closed bindings, surfaced as deferral_permitted. Split PEER_DISCONNECT into DOWNSTREAM_DISCONNECT and UPSTREAM_DISCONNECT and attribute WebSocket relay failures by direction instead of a generic peer error. Compile the content-guard example in lint and branch checks so proto renames cannot break it silently. BREAKING CHANGE: WebSocketSessionEndReason and WebSocketSessionEnd are replaced by the shared MiddlewareSessionEndReason and MiddlewareSessionEnd. NORMAL_CLOSE is now NORMAL, UPSTREAM_REJECTED is now UPSTREAM_FAILURE, and PEER_DISCONNECT is split into DOWNSTREAM_DISCONNECT and UPSTREAM_DISCONNECT. Enum numbers are unchanged so binary wire compatibility is preserved; generated symbols and JSON names change. Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(middleware): describe skip as opting out of inspection Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * feat(middleware): add body-phase block_delivery and skip_remaining actions Body results may now stop delivery or opt out of inspecting the rest of the response after a prefix. One HttpResponseBlockDelivery message is shared by preflight and body results and documents the difference between blocking before and after head commitment. Drop the field reservations, since nothing in this contract has shipped, and renumber session_end to close the gap. Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * refactor(middleware): share HTTP body leaf messages across directions HttpBodyUnit, HttpBodyPassThrough, HttpBodyTransform, HttpBodySkipRemaining, and HttpBodyMode carry no response-specific semantics, so name them for reuse by the streaming request hook. Envelopes, results, preflight, and block_delivery stay response-specific because commitment semantics differ. Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * refactor(middleware): keep HTTP body leaf messages response-specific Reverts the shared HttpBody* naming. A direction-specific payload such as a response-only semantic mode would otherwise add unreachable variants to the other direction or force a source-breaking fork after 0.1.0. The streaming request hook defines its own HttpRequestBody* messages and copies the shape; SDKs present a direction-neutral body handler over both. Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(middleware): simplify response proto comments Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(middleware): reject undispatched response bindings Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(middleware): simplify phase field comment Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * refactor(middleware): rename HTTP response preflight result Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * feat(middleware): add HTTP response trailer results Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(middleware): define response block delivery Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(middleware): define stage-local response body modes Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(middleware): define final response body units Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(middleware): define streaming response deferral Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(middleware): defer whole-body accumulation timeout Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(middleware): align response result diagnostics Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * test(middleware): cover upstream WebSocket disconnect Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(middleware): keep response streams unit-local Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(middleware): trim disconnect compatibility note Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> --------- Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> |
||
|
|
48a8a4bf09 |
feat(ocsf): configurable schema version for SIEM backward compatibility (#2717)
* feat(ocsf): configurable schema version for SIEM backward compatibility Add a gateway-configurable OCSF schema version target that downgrades JSONL output for SIEMs that only support older schema versions. AWS Security Lake requires v1.1.0, Splunk CIM Add-On targets v1.1-v1.3. The downgrade filter strips profile-gated fields (ai_model, container, observation_point_id), removes unknown profiles from metadata.profiles, rewrites metadata.version, and adds an unmapped.downgraded_from breadcrumb so auditors can distinguish "no model involved" from "model attribution stripped." Supported target versions (1.1, 1.3) are enforced by an allow-list in the settings registry. Invalid values are rejected with a clear error. The setting flows to sandboxes via the settings bundle and takes effect on the next poll cycle. The shorthand log output is unaffected. Closes #2662 Signed-off-by: Adel Zaalouk <zanetworker@gmail.com> Signed-off-by: Adel Zaalouk <azaalouk@redhat.com> * fix(ocsf): align downgrade with schema version Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> --------- Signed-off-by: Adel Zaalouk <zanetworker@gmail.com> Signed-off-by: Adel Zaalouk <azaalouk@redhat.com> Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> Co-authored-by: John Myers <9696606+johntmyers@users.noreply.github.com> |
||
|
|
08eac8c46d |
fix(sandbox): detect an available login shell instead of hardcoding /bin/bash (#3147)
* fix(sandbox): detect an available login shell instead of hardcoding /bin/bash The built-in default sandbox command and the interactive SSH session hardcoded /bin/bash. Minimal images such as Alpine ship only /bin/sh (BusyBox ash), so sandbox startup failed with an opaque "No such file or directory (os error 2)" that never named the missing binary. Add openshell-core::shell with shell-path constants and a runtime detect_login_shell() that resolves a shell present in the sandbox image ($SHELL if executable, then bash, then /bin/sh). Use it for: - the built-in default command (only the default is remapped; explicit user commands are never rewritten), resolved in the supervisor so it inspects the sandbox filesystem rather than the gateway's - the SSH interactive shell - the SHELL environment variable Also name the program in the spawn error so a missing shell/binary is diagnosable instead of a bare ENOENT. Refs #3146 Signed-off-by: Akram <akram.benaissi@gmail.com> * fix(sandbox): drop $SHELL preference in shell detection $SHELL is image/user-controlled and the detected shell is later invoked with `-lc`, so an executable that is not a compatible shell (e.g. SHELL=/bin/false) would pass the executable check and then break command execution even when /bin/sh is available. Resolve only from known shell paths instead. Also add a USR_BASH constant for /usr/bin/bash rather than a string literal in SHELL_CANDIDATES. Refs #3146 Signed-off-by: Akram <akram.benaissi@gmail.com> * fix(sandbox): resolve the default login shell in the supervisor (empty command = default) Addresses review: interactive PTY SSH now uses the detected shell, the shell tests are portable across the Windows lane, and default-shell provenance is carried without a new spec field. An omitted command is left empty end to end and resolved in the supervisor, which is the only place that sees the sandbox image: - The CLI forwards the command as-is; the gateway persists an omitted command as empty (no baked /bin/bash -l) and requests a TTY. - MainProcessConfig carries the command empty (the transport now allows it); the supervisor resolves a login shell that exists in the sandbox image (bash when present, otherwise /bin/sh on minimal images like Alpine) and logs the resolved shell. - Interactive PTY SSH (spawn_pty_shell) uses the detected shell; a shared build_ssh_shell_command helper covers the PTY and non-PTY paths, with a deterministic sh-only regression test. - Unix-only shell tests are gated with cfg(unix). An explicit command is always run verbatim. Refs #3146 Signed-off-by: Akram <akram.benaissi@gmail.com> --------- Signed-off-by: Akram <akram.benaissi@gmail.com> |
||
|
|
e64b0352e8 |
feat(middleware): broaden HTTP header mutation authority (#3072)
* feat(middleware): broaden HTTP header mutation authority Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(middleware): protect response credential headers from mutation Response middleware could write or remove Set-Cookie, WWW-Authenticate, Authentication-Info, and Proxy-Authentication-Info, letting a stage plant or strip credentials the sandbox client acts on. Protect them in both directions, matching the request profile's treatment of Authorization and Cookie, and reserve the x-openshell-credential prefix for responses too. Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(middleware): reject credential placeholder writes Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> --------- Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> |
||
|
|
5021f23b06 |
fix(tui): replace alpha badges with version (#3114)
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> |
||
|
|
857af42a16 |
feat(vm): support corporate HTTP forward proxy egress for microVM sandboxes (#3090)
* feat(vm): support corporate HTTP forward proxy egress for microVM sandboxes The corporate forward proxy machinery from #1792 is driver-agnostic and already merged: openshell-supervisor-network implements CONNECT chaining, NO_PROXY matching, credentials, https:// proxies and corporate CA trust, and openshell-sandbox exposes it as six argv-only flags. Podman gained the driver half in #2245/#2512 and Kubernetes in #2633; the VM driver had none of it, so VM sandboxes on proxy-only networks could not reach any destination requiring the proxy even when policy allowed it. The blocking piece was not proxy logic but delivery: the VM guest init script runs as PID 1 and execs a fixed supervisor command line, and libkrun's krun_set_exec receives an empty argv, so there was no channel for driver-owned supervisor arguments. The supervisor's proxy flags deliberately have no environment fallback, and build_guest_environment merges user-supplied environment, so the guest env is not a safe transport either. Add a driver-authored argument file, mirroring the existing init.d manifest: the driver writes /opt/openshell/supervisor-args into the overlay upperdir on every launch and the guest reads it verbatim, one argument per line, appending it to every supervisor exec. It is written even when empty, which is what makes the channel unforgeable -- the upperdir always shadows the read-only image layer, so an image can neither supply its own arguments nor disable the operator's by omitting the file. Because both launch backends exec the same init script, this covers libkrun and QEMU without touching either. A microVM has no bind mounts or container secrets, so the credential and CA bundle are staged into the per-sandbox overlay the way the gateway JWT already is: credential root-only at 0600, CA at 0644, both rewritten every launch so a removed setting clears prior material, and both deleted with the sandbox state directory. This places the credential at rest in the overlay image on the gateway host, which differs from the Podman secret model and is documented as an explicit security consideration. Validation is fail-closed and shared: a new openshell_core::driver_utils::validate_upstream_proxy_settings holds the pairing rules the Podman driver established, and both the gateway and the driver call it so an invalid table names the offending key instead of surfacing as an opaque driver-readiness timeout. Guest egress leaves through gvproxy, so a proxy on the gateway host's loopback is reachable only through host.openshell.internal; the guest to gateway callback is unaffected. Closes #3088 Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(vm): bound the proxy CA read and scope the host-loopback recipe Two review findings on the corporate forward proxy support for microVM sandboxes. The driver read the operator's proxy_ca_bundle with an unbounded fs::read and accepted it on a substring match for the PEM BEGIN CERTIFICATE marker. A special file such as /dev/zero therefore grew driver memory without bound on every authorized sandbox create, and a PEM block holding invalid DER passed the host check but contributes no trust anchor in the guest, so every supervisor would fail after boot with an error attributed to the sandbox rather than to the setting. Move the read into openshell-core as read_upstream_proxy_ca_bundle_file: it reuses the credential reader's bounded-read path (non-regular files rejected on fstat, size capped, read bounded even if the file grows), then requires at least one anchor that RootCertStore::add_parsable_certificates accepts. The supervisor's own reader now delegates to it, so host acceptance and guest acceptance are the same function and cannot drift. The published host-loopback recipe was written for libkrun only. gvproxy NATs host.openshell.internal to the gateway host's 127.0.0.1, but GPU sandboxes run on the QEMU/TAP backend where that name resolves to the TAP host address and the driver's own nftables input chain accepts only the gateway port from the guest — no proxy on the gateway host is reachable there at any bind address, so an operator following the generic recipe lost all proxy-required egress while configuration validation succeeded. Scope the recipe to libkrun in every reference and reject a gateway-host proxy URL when a launch plan resolves to QEMU, naming the reason, instead of booting a sandbox whose policy-approved CONNECTs all time out. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(vm): match the QEMU proxy preflight to the selected TAP host The gateway-host proxy guard added for the QEMU/TAP backend classified the wrong set of addresses in both directions. It ran at the top of configure_qemu_launch_plan, before the subnet allocation that settles plan.host_ip, so it could not compare against the address the guest actually reaches the host on. An operator pointing https_proxy at the sandbox's own TAP host address, such as 10.0.128.1, passed the check, and the driver's nftables input chain — which accepts only the gateway port from the guest — then dropped every policy-approved CONNECT, which is exactly the silent timeout the guard exists to prevent. In the other direction it rejected 192.168.127.254 unconditionally. That address is special only to libkrun/gvproxy; on QEMU/TAP it is an ordinary address that may be routable through the guest's masqueraded egress, so the guard refused a working configuration. Run the check after the launch plan's network allocation, on both the freshly-allocated and already-complete paths, and compare IP literals with that sandbox's selected TAP host. Loopback literals, localhost, and the documented host aliases that write_host_gateway_aliases seeds to the TAP host still classify as the gateway host, and the failure names the address. The gvproxy host-loopback constant returns to being a documentation anchor. Signed-off-by: Philippe Martin <phmartin@redhat.com> --------- Signed-off-by: Philippe Martin <phmartin@redhat.com> |
||
|
|
74960ebfae |
feat(server): add sandbox templates (#2833)
* feat(server): add sandbox workload templates Signed-off-by: Gordon Sim <gsim@redhat.com> * feat(go-sdk): add sandbox workload template support Signed-off-by: Gordon Sim <gsim@redhat.com> * feat(rust-sdk): add sandbox workload template support Signed-off-by: Gordon Sim <gsim@redhat.com> * feat(python-sdk): add sandbox workload template support Signed-off-by: Gordon Sim <gsim@redhat.com> * feat(typescript-sdk): add sandbox workload template support Signed-off-by: Gordon Sim <gsim@redhat.com> * docs(agents): document sandbox workload templates Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(cli): support default GPU requests in sandbox templates Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(server): cap sandbox templates per workspace Signed-off-by: Gordon Sim <gsim@redhat.com> * docs(architecture): document sandbox workload template boundaries Signed-off-by: Gordon Sim <gsim@redhat.com> * feat(cli+sdk): expose sandbox workload template provenance Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(cli): include sandbox template annotations in output Signed-off-by: Gordon Sim <gsim@redhat.com> * feat(sandbox): add label selectors to template listing Signed-off-by: Gordon Sim <gsim@redhat.com> * test(e2e): cover sandbox template failure paths Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(server): preserve command and ttl when creating sandbox from template Signed-off-by: Gordon Sim <gsim@redhat.com> * test(server): add field coverage test for template merge Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(go-sdk): add pagination support to fake client Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(cli): warn on env vars that looks like secrets Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(sdk-ts): propagate sandbox workspace through lifecycle calls Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(docs): update workspace management docs Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(ts-sdk): support command and tty when creating from template Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(go-sdk): support command and tty when creating from template Signed-off-by: Gordon Sim <gsim@redhat.com> * test(python-sdk): verify command and tty handling when creating from template Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(go-sdk): update docs and ClientInterface Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(server): validate sandbox create specs before I/O Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(go-sdk): guard empty DNS-1123 label validation Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(python-sdk): allow empty template builder mappings Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(cli): align template GPU JSON default output Signed-off-by: Gordon Sim <gsim@redhat.com> --------- Signed-off-by: Gordon Sim <gsim@redhat.com> |
||
|
|
9ca19e6c80 |
refactor(compute): decouple gateway driver composition (#2823)
* refactor(compute): decouple gateway driver composition Move first-party composition and VM process ownership into openshell-gateway, leaving openshell-server backend-independent. Update packaging and build references with the new crate, simplify the compiled-driver boundary, and keep the driver-free gateway path buildable with bundled Z3 tooling. Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(telemetry): bound compute driver categories Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(core): keep runtime transport generic Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(compute): complete server driver decoupling Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(compute): preserve driver integrations after rebase Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> * fix(compute): preserve docker tracing after decoupling Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> * fix(compute): preserve driver behavior after extraction Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(compute): remove MXC policy side channel Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(compute): separate policy delivery from readiness Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> |
||
|
|
07df822090 |
feat(providers): make profiles authoritative (#2962)
* feat(providers): make profiles authoritative Closes #1988 Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * docs(providers): move profiles into provider navigation Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * docs(providers): clarify provider attachment lifecycle Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * fix(tui): scroll provider profile picker Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * fix(providers): honor profile credential semantics Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * fix(providers): prefer exact profile IDs Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * fix(providers): harden authoritative profile adoption Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * test(oidc): align provider fixtures with profiles Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * fix(providers): preserve authoritative profile lifecycle Signed-off-by: John Myers <johntmyers@users.noreply.github.com> --------- Signed-off-by: John Myers <johntmyers@users.noreply.github.com> Co-authored-by: John Myers <johntmyers@users.noreply.github.com> |
||
|
|
4eaa1051f8 |
docs: correct some comments and references (#3059)
OCSF_VERSION has been 1.8.0; two descriptions still said 1.7.0. The default_gateway_id doc comment described a hostname fallback the implementation never had. A per-replica default would be wrong here: gateway_id is embedded in the JWT iss/aud, so it has to be stable across replicas. Signed-off-by: Kris Hicks <khicks@nvidia.com> |
||
|
|
c27a3a3ce7 |
fix(compute): recover Error-phase sandboxes on gateway startup (#2269)
* fix(compute): recover Error-phase sandboxes on gateway startup Sandboxes whose container exited on its own (e.g. SIGTERM from a Podman or Docker machine restart) were left stuck in the Error phase even though the container could be restarted. The startup sweep skipped every Error-phase sandbox unconditionally. Include Error-phase sandboxes in the startup sweep when their Ready condition reason marks a container exit or stop. If the driver restarts the container, move the sandbox back to Provisioning with a Resumed condition; if the container is gone or the start fails, leave the Error state untouched. Genuine (non-container-exit) errors are still skipped. Container-exit restart is handled by the drivers' existing idempotent start_sandbox path, which already restarts stopped/exited containers. The ContainerExited/ContainerStopped condition reasons are promoted to shared constants in openshell-core so the gateway and drivers agree on the recovery signal. Fixes #2179 Signed-off-by: Ian Miller <milleryan2003@gmail.com> * fix(compute): keep ordinary container exits terminal during startup recovery Startup recovery treated every non-OOM container exit as recoverable, lumping application crashes and non-zero exits together with machine/daemon-restart signal kills under the shared `ContainerExited` Ready-condition reason. On the next gateway startup it restarted those crashed workloads and replaced the terminal error with `Resumed`, erasing the failure signal and repeatedly reviving crash-prone sandboxes. Introduce a distinct `ContainerRuntimeRestart` reason for containers terminated by an external signal (exit 137/143 = SIGKILL/SIGTERM), which is the signature of a machine/daemon restart. The Podman inspect-based classification emits it; the startup recovery gate now recovers only `ContainerRuntimeRestart` and `ContainerStopped`. Ordinary `ContainerExited` records stay terminal, so genuine failures keep their error signal. Because only the new reason is recoverable and older gateways never persisted it, exits recorded before this change also stay terminal after an upgrade. Add regression coverage: a generic `ContainerExited` error is not relaunched, a signal-kill classifies as `ContainerRuntimeRestart`, and the startup sweep recovers only the runtime-restart record. Signed-off-by: Ian Miller <milleryan2003@gmail.com> * fix(docker): reclassify signal-killed containers as runtime restart Mirror the Podman classification in the Docker driver so that externally signal-killed containers (exit 137/143, non-OOM) are persisted with the ContainerRuntimeRestart ready reason and become eligible for startup recovery, while ordinary exits and OOM kills stay terminal. current_snapshots now inspects EXITED containers to obtain the exit code and OOM flag, then applies the shared exit classification. Without the inspect step the list-summary path lacks an exit code, so the recovery gate could never fire for Docker. Adds unit tests covering signal-kill reclassification (137/143), ordinary exit staying terminal (exit 1), and OOM staying terminal despite exit 137. Signed-off-by: Ian Miller <milleryan2003@gmail.com> --------- Signed-off-by: Ian Miller <milleryan2003@gmail.com> |
||
|
|
69a05ebb3b | fix(sandbox): complete successful main processes (#2884) | ||
|
|
f68867b869 |
feat(gateway): identify gateways in exported traces (#2647)
* feat(gateway): add installation name configuration Add a first-class operator-assigned gateway name with TOML, CLI, environment, and Helm configuration surfaces. Local gateways default to openshell, while Helm defaults to the chart fullname; operators sharing a collector across namespaces or clusters can set a globally distinct name. Signed-off-by: Kris Hicks <khicks@nvidia.com> * feat(gateway): identify gateways in exported traces Attach the configured gateway installation name and compute driver to the gateway OpenTelemetry resource so operators can filter traces from multiple installations that share a collector. Forward the gateway name and OTLP endpoint to managed external drivers so their distinct service resources carry the same installation identity. Keep service.name stable per process type, omit blank resource values, and leave per-span operation names and request attributes unchanged. Refs #2507 Signed-off-by: Kris Hicks <khicks@nvidia.com> --------- Signed-off-by: Kris Hicks <khicks@nvidia.com> |
||
|
|
bcd517bbe0 |
feat(driver-mxc): native Windows MXC compute driver + server wiring (#2721)
* feat(driver): add MXC compute driver for Windows isolation sessions Introduces the openshell-driver-mxc crate implementing ComputeDriver backed by Microsoft MXC isolation sessions (Windows only). Wires the new driver into the server's build_compute_runtime dispatch and adds the Mxc variant to ComputeDriverKind. Also adds a local protobuf-src stub (tools/protobuf-src-local) to unblock Windows builds that lack MSYS2/MinGW, and pins the zig Windows x64 toolchain in mise.lock. (cherry picked from commit 4f7012224efb18fbfeb47aa87e0cfd3f036f32f0) Signed-off-by: Jamie King <jamiek@nvidia.com> * wip(mxc): checkpoint hung-agent work (recon, policy_map embed, A1 wiring, demo artifacts) Safety checkpoint of uncommitted work from the background agent run that stalled mid-Step-7. Includes: mxc-driver-recon.md (Step 0.5), policy_map.rs (~876L embedded mapper), A1 policy-threading edits across driver.rs/policy.rs/mxc.rs/compute/mod.rs, and examples/ (demo.yaml + mxc-gateway.toml). Not yet verified to compile end-to-end; to be reorganized into the skill's Step 11 commit sequence. (cherry picked from commit 38e42c03870be3d10e984a54f17a3b61122ff510) Signed-off-by: Jamie King <jamiek@nvidia.com> * test(mxc): fix lifecycle and policy unit-test compile drift - Bring futures::StreamExt into scope for the watch-stream `.next()` call in driver::lifecycle_tests so the negative policy proof test compiles. - Bind a local `mapper` and drop the unused/deprecated NetworkBinary in the embedded-mapper network-policy rejection test. Signed-off-by: Jamie King <jamiek@nvidia.com> (cherry picked from commit 039b0baf98735ca672dae52be8d3af2417dc0c1a) Signed-off-by: Jamie King <jamiek@nvidia.com> * fix(mxc): downgrade missing sandbox_token to debug log The gateway mints `sandbox_token` only when a sandbox-JWT issuer is configured. There is no in-sandbox supervisor on MXC (supervisor-removal design — D1/D4), so no component ever consumes the token; requiring it on the driver side blocks the demo's `--disable-tls` smoke gateway with a spurious `invalid_argument`. Log the absence and proceed instead. Signed-off-by: Jamie King <jamiek@nvidia.com> (cherry picked from commit cea209797d0edcb1d152251e748900b0a63cca62) Signed-off-by: Jamie King <jamiek@nvidia.com> * fix(mxc): keep sandbox Ready after a successful one-shot agent exec monitor_exec demoted Ready->Error on exit 0 (reason ExecCompleted), so the positive demo (write hello.txt + exit) landed in Error phase. Keep Ready=True (reason AgentCompleted) on success; only non-zero exits go to ExecFailed. Tighten the positive lifecycle test to assert the terminal condition stays Ready=True/AgentCompleted. Verified live via gateway mock round-trip: phase now Provisioning->Ready with no demotion. (cherry picked from commit 54ab030f03ca0f83d0050d8dd843b633651684ad) Signed-off-by: Jamie King <jamiek@nvidia.com> * feat(mxc): add processContainer backend for default-deny enforcement Add a backend selector to the MXC driver (isolation_session default | process_container). process_container drives a one-shot AppContainer that is genuinely default-deny: a write to any ungranted path is denied by the OS, unlike isolation_session which is grant-only and cannot deny. The lifecycle forks on the flag - isolation_session keeps provision/start/exec, process_container runs a single ephemeral container via run_oneshot. Also: run-demo.ps1 gains -Backend and hardens the CLI register/create calls; docs corrected to state isolation_session does NOT deny out-of-policy writes and that the negative proof requires process_container. Verified end-to-end on a real demo box (gateway -> CLI -> driver -> MXC): in-policy write succeeds, out-of-policy write denied (PermissionDenied), OVERALL: PASS. (cherry picked from commit c6cde3860bbe1b8edb3147d3e840f6bf0ece32d8) Signed-off-by: Jamie King <jamiek@nvidia.com> * refactor(driver-mxc): embed policy mapper as a module; remove standalone crate Adopt the proto-based mapper (map_to_mxc) as the single source of truth, embedded in openshell-driver-mxc as a Windows-gated `policy_map` module. Rewire EmbeddedPolicyMapper to call it directly on the typed SandboxPolicy, deleting the serde_yaml proto->YAML bridge. Move the CLI to a windows-gated example and the parity tests into the crate; delete openshell-policy-mapper. - gate policy_map + seam Windows-only (MXC is Windows-only) - drop serde_yaml; add dev-deps openshell-policy, clap, anyhow - normalize mapped paths to Windows form in the seam, in one place - docs: add driver-mxc to AGENTS.md table; correct design doc section 17 test lane Signed-off-by: Giedrius Burachas <gburachas@nvidia.com> (cherry picked from commit f22f9c7a25b9a651c5c5cc73f62fb01c4d6c1a8d) Signed-off-by: Jamie King <jamiek@nvidia.com> * feat(driver-mxc): implement lossless split_policy for proxy-delegated egress Signed-off-by: Giedrius Burachas <gburachas@nvidia.com> (cherry picked from commit 96d6afa0e2dc6a1d54edd12c34a0ceb0a30dadd0) Signed-off-by: Jamie King <jamiek@nvidia.com> * feat(driver-mxc): implement Pattern-C governed-egress split through the policy seam - split_policy: SocketAddr proxy_redirect (replaces bare port), processcontainer containment guard naming MXC M1, version preserved in the trimmed proxy_policy, delegation reported as an info loss item - seam: MappedConfig carries trimmed_policy + proxy_addr; MapCtx.egress selects the split path; coarse path unchanged when egress is disabled - driver: [openshell.drivers.mxc] egress_proxy / egress_proxy_addr config, validated at create (isolation_session rejected until M1); lifecycle threads the redirect into provision and stores the trimmed policy per sandbox, emitting an EgressRedirect platform event - mxc: optional MxcNetwork block (defaultPolicy=block + proxy) in provision and one-shot configs; mock records configs for test assertions - tests: lossless-invariant suite over all example policies (validate + serialize round-trip), split lifecycle proof, M1 rejection; example gains --split --proxy-addr writing mxc-config.json / trimmed-policy.yaml / loss-report.json Signed-off-by: Giedrius Burachas <gburachas@nvidia.com> (cherry picked from commit 34d54ad9f25dc6034c3ba15668555ff0d22cddd8) Signed-off-by: Jamie King <jamiek@nvidia.com> * fix(driver-mxc): emit MXC network.proxy as {localhost: port} Verified against the real wxc-exec 0.6.0-alpha via --dry-run: MXC accepts only the {localhost: N} proxy shape (the form the design doc specifies) and rejects {host, port} with a parse error. Schema 0.6.0-alpha can express only a loopback port, so non-127.0.0.1 redirect addresses are now rejected: split_policy emits an error loss (no proxy block) and the driver refuses egress_proxy_addr values off 127.0.0.1. Per-sandbox attribution must use per-sandbox ports until the schema widens. Signed-off-by: Giedrius Burachas <gburachas@nvidia.com> (cherry picked from commit edde8d5434571fd5398409204fcf6862672c0793) Signed-off-by: Jamie King <jamiek@nvidia.com> * fix(driver-mxc): serialize isolation_session stop/deprovision as unit variants Empirical contract finding from the real test lane (build 26300.8553, wxc-exec 2026-06-10): the stop and deprovision experimental blocks are unit variants in the wxc-exec schema and must serialize as null; sending {} is rejected with malformed_request (invalid type: map, expected unit), while provision/start accept maps. The production invoker, the real-lane test, the probe script, and the e2e runner all sent {} - the driver could provision and run an agent but never stop or delete an isolation-session sandbox against this build. Pinned by a unit test. Signed-off-by: Giedrius Burachas <gburachas@nvidia.com> (cherry picked from commit 0df39ca0b22ebc21eb965b2b567a5b2cac26af32) Signed-off-by: Jamie King <jamiek@nvidia.com> * test(driver-mxc): add Tier-0 mapper coverage matrix with schema drift guard Three-quadrant, table-driven matrix (38 tests): mappable fields assert exact MXC output; every OpenShell field MXC cannot express asserts a loss item with the expected severity (and seam rejection on error); an empty policy asserts the restrictive default-deny posture for every MXC knob OpenShell does not control. The handled_fields_inventory drift guard serializes a fully-populated policy and compares its YAML keys against the mapper-handled field lists, so a new openshell-policy field fails the suite until consciously mapped, delegated, or reported as loss. Re-exports the policy seam types for integration tests; adds serde_yml, base64, serde_json as dev-dependencies. Signed-off-by: Giedrius Burachas <gburachas@nvidia.com> (cherry picked from commit 91807f984a3b16846e35d6ca0d5ec41057aafa3a) Signed-off-by: Jamie King <jamiek@nvidia.com> * feat(driver-mxc): inject agent_env into sandbox process.env Add MxcComputeConfig.agent_env: each entry is either KEY=VALUE (verbatim) or a bare KEY resolved from the gateway host environment at launch, keeping secrets (e.g. inference API keys) out of the config file. Wire it into the agent process so gateway-launched agents can authenticate to cloud endpoints (process.env was previously hardcoded empty). Unit-tested via resolve_agent_env_passthrough_and_host_lookup. Also add a gateway-driven cloud-inference (T1) test harness: mxc-inference.toml (agent_env + curl agent), inference.yaml policy, and run-inference-test.ps1 which starts the gateway, creates an isolation_session sandbox, runs an authenticated Nemotron call, and bundles redacted results. Documented agent_env in mxc-gateway.toml. Validated end-to-end on the test box (chat HTTP 200 + completion via the gateway). (cherry picked from commit 94d9e827b8b77e0af9dc654943ea8e7d8400cc9f) Signed-off-by: Jamie King <jamiek@nvidia.com> * test(driver-mxc): avoid unsafe env mutation in resolve_agent_env test Replace std::env::{set,remove}_var (unsafe + racy under parallel test execution in edition 2024) with a read-only PATH lookup. Preserves all three behaviors under test and drops the #[allow(unsafe_code)]. (cherry picked from commit ac5766eeab2db4e8cc6fcd8d8a97809edaf3df30) Signed-off-by: Jamie King <jamiek@nvidia.com> * fix(mxc): adapt MXC driver to current GitHub OpenShell API The MXC driver crate was authored on GitLab against an earlier proto/core API. Adapt it to the API on GitHub main: - build_capabilities_response no longer takes supports_interactive_session - DriverSandboxSpec.gpu (bool) is now resource_requirements; detect GPU via effective_driver_gpu_count(driver_gpu_requirements(..)) - DriverSandbox gained a `workspace` field - SandboxPolicy gained `network_middlewares`: pass it through the proxy split, emit a loss item on the coarse MXC path, and account for it in the mapper drift-guard test Verified: cargo check + 75 mock-based tests pass (lib 27, examples 10, policy_mapper_matrix 38). Signed-off-by: Jamie King <jamiek@nvidia.com> * feat(server): wire the MXC compute driver into the gateway on Windows Register openshell-driver-mxc as the Windows-only in-process compute backend so compute_driver = "mxc" resolves to a working runtime: - ComputeRuntime::new_mxc, adapted to the current 11-arg from_driver - mxc_policy_sink A1 side channel, staged in create_sandbox before dispatch - mxc_config_from_context loader and the Mxc dispatch arm (Windows constructs; other targets return an explicit "Windows-only" error) - Windows-gated openshell-driver-mxc dependency - Mxc arms for the telemetry, config-file required-fields, and CLI reserved-builtin matches to keep them exhaustive/correct Verified with cargo check --workspace --features openshell-prover/bundled-z3 on x86_64-pc-windows-msvc, stacked on PR #2496. Signed-off-by: Jamie King <jamiek@nvidia.com> * fix(driver-mxc): implement GetGatewayListenerRequirements for #2496 base Signed-off-by: Jamie King <jamiek@nvidia.com> * test(driver-mxc): add probe-gated real wxc-exec test lane (no mocks) - tests/wxc_exec_real.rs: ignored-by-default integration tests against a real wxc-exec. Six --dry-run contract tests run wherever the binary exists (they caught the network.proxy shape mismatch); enforcement tests (processcontainer default-deny positive/negative, isolation session lifecycle round trip with a deprovision drop-guard) probe the backend and SKIP with a recorded reason where it is not live. - examples/probe-mxc-host.ps1: classifies a host (OS build, --probe, per-backend trial) and emits a JSON capability verdict. - examples/run-mxc-e2e.ps1 + e2e-policies/: scenario runner generalizing run-demo.ps1 (fs-rw, fs-readonly, fs-default-deny-empty, network-policy-rejected) with PASS/FAIL/SKIP gating and a stale OPENSHELL_MXC_MOCK_WXC guard in real mode. - tasks/windows.toml: windows:test:mxc-real:x64, windows:e2e:mxc, windows:e2e:mxc:mock. Signed-off-by: Giedrius Burachas <gburachas@nvidia.com> (cherry picked from commit 49afafe892caded59c4df50a9b652011ad97f41c) Signed-off-by: Jamie King <jamiek@nvidia.com> * fix(driver-mxc): address PR review feedback Signed-off-by: Shailendra Singh <shailendras@nvidia.com> * docs: defer public MXC documentation Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> * fix(driver-mxc): build Windows capabilities response Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> * fix(server): gate in-tree tracing on Windows Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> * test(windows): fix cross-platform test assumptions Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> * test(windows): verify process and PEM portably Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> * test(windows): keep lifecycle command in policy Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> --------- Signed-off-by: Jamie King <jamiek@nvidia.com> Signed-off-by: Giedrius Burachas <gburachas@nvidia.com> Signed-off-by: Shailendra Singh <shailendras@nvidia.com> Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> Co-authored-by: Prashant K <pkhodade@nvidia.com> Co-authored-by: Giedrius Burachas <gburachas@nvidia.com> Co-authored-by: Shailendra Singh <shailendras@nvidia.com> Co-authored-by: Drew Newberry <385+drew@users.noreply.github.com> |
||
|
|
aa848f164d |
fix(core): fall back to podman CLI when no API socket is found (#1858)
* fix(core): fall back to podman CLI when no API socket responds Auto-detection only checked well-known Podman socket paths, so a Podman machine exposing its API socket at a non-standard location went undetected. The symlink at a well-known path is not always present; it varies by Podman version, machine provider, and platform. Extend detect_podman_socket() to fall back to podman CLI discovery when no well-known candidate responds: podman info --format json determines whether the service is local or remote, and podman machine inspect resolves the host-side forwarded socket for VM-backed machines. All existing callers (driver auto-detection, the Podman driver, and the VM driver's container-engine fallback) pick this up without change. Select the machine backing the active Podman connection instead of the first entry from podman machine inspect, honoring podman's connection precedence: CONTAINER_CONNECTION, then CONTAINER_HOST (mapped to a connection by URI), then the containers.conf default. An explicit endpoint that maps to no known machine is left unresolved rather than guessing an unrelated machine. When CONTAINER_HOST is an explicit unix:// socket, use that path directly since podman info connects through it. Update the gateway config reference and the Podman driver README, which described probe-only detection. Signed-off-by: Russell Bryant <rbryant@redhat.com> * fix(core): inspect podman machine by name and bound discovery probes Address two startup defects in Podman socket auto-detection: - Remote discovery ran `podman machine inspect` with no arguments, which inspects only `podman-machine-default`. On a host whose default connection is a different machine, selection fell back to the first (wrong) entry, so the gateway could operate against the wrong backend. Resolve the active machine first and inspect it by name, trying the `-root`-stripped machine for rootful connections and returning None when no machine can be mapped instead of substituting another. - `podman info`, `podman machine inspect`, and `podman system connection list` used unbounded `Command::output()`, so a stalled machine, SSH connection, or helper could hang gateway startup indefinitely. Route all three through a bounded runner that kills and reaps the child on a documented deadline and returns None so detection can continue. Signed-off-by: Russell Bryant <rbryant@redhat.com> * fix(core): bound podman probe stdout drainage and kill descendants The previous timeout only bounded waiting for the direct child to exit. Draining stdout still called read_to_end, which waits for pipe EOF — a Podman probe can leave a daemonized descendant (SSH multiplexer, gvproxy) that inherited the stdout pipe, so drainage could wait indefinitely even after the direct child exited, defeating the deadline. Run each probe as its own process-group leader, cap stdout drainage by the same deadline via a channel, and on expiry kill the whole process group so descendants holding the pipe are terminated and EOF is reached. Add a regression test where the direct child exits but a backgrounded descendant keeps stdout open. Signed-off-by: Russell Bryant <rbryant@redhat.com> * fix(core): make podman probe deadline absolute against escaped descendants The prior drainage timeout still ended in an untimed rx.recv() after killing the process group. A descendant that escaped the probe's process group (e.g. via setsid) while holding the stdout pipe would not be killed, so read_to_end never saw EOF, the reader never sent, and that recv() could block gateway startup indefinitely. On drainage timeout, best-effort kill the process group and return None immediately, abandoning the reader thread instead of waiting on it again. The deadline now bounds the whole call regardless of what descendants do. Add a regression test whose stdout-holding descendant escapes the process group (`set -m`) and assert the call still returns promptly. Signed-off-by: Russell Bryant <rbryant@redhat.com> --------- Signed-off-by: Russell Bryant <rbryant@redhat.com> |
||
|
|
0a1f246587 |
feat(sandbox,podman): trust corporate CA for https:// proxies and intercepted TLS (#2512)
* feat(sandbox,podman): trust corporate CA for https:// proxies and intercepted TLS The corporate proxy chaining only accepted plain http:// proxy URLs, so operators whose forward proxy terminates TLS with a private corporate CA had no way to reach it, and TLS-intercepting proxies (mitmproxy, squid ssl-bump) that re-sign tunneled server certificates broke every upstream handshake after CONNECT. The supervisor now accepts https:// proxy URLs: it wraps the connection to the proxy in TLS before the CONNECT handshake, verifying the proxy certificate against the built-in Mozilla roots, the system CA bundle, and an optional operator corporate CA bundle. The upstream dial returns a Plain/Tls stream enum consumed generically by the relay paths. The corporate CA is delivered as a driver-supplied command-line argument (--upstream-proxy-ca-bundle), never an environment variable, matching the hardened proxy-config model where a sandbox image cannot influence the operator's egress boundary. It is folded into the sandbox combined trust bundle (write_ca_files) and the L7 upstream verification store (build_upstream_client_config) at startup, so intercepted upstream handshakes succeed and sandbox workloads trust the re-signed certificates. Configuration is fail-closed: a CA bundle set without a proxy, or an unreadable or certificate-free file, is fatal rather than silently weakening the trust boundary. The shared parse_upstream_proxy_url validator accepts https:// (recording the scheme so the driver and supervisor agree), keeping the explicit-port requirement. The Podman driver gains a proxy_ca_bundle operator setting (TOML, --sandbox-proxy-ca-bundle, OPENSHELL_SANDBOX_PROXY_CA_BUNDLE) that bind-mounts the host PEM read-only into the sandbox (a CA certificate is not secret) and points --upstream-proxy-ca-bundle at it, with a create-time readability check. The standalone dev gateway task passes OPENSHELL_SANDBOX_PROXY_CA_BUNDLE through to the generated podman config, so a local gateway can be pointed at a TLS-intercepting proxy without hand-editing the regenerated TOML. Refs #1792 Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(sandbox): reject CA bundles with valid PEM framing but invalid X.509 DER The proxy CA bundle validation counted PEM blocks that base64-decoded successfully, but did not verify the decoded bytes were accepted as trust anchors by RootCertStore. A bundle with syntactically valid PEM framing but invalid DER would pass the startup check while contributing zero usable anchors, causing opaque TLS failures at runtime instead of a fail-closed startup error. Validate decoded certificates through RootCertStore::add_parsable_certificates and reject the bundle unless at least one is accepted. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(sandbox,podman): allow proxy auth without insecure acknowledgement for https:// proxies For an https:// proxy the Proxy-Authorization credential travels inside the verified TLS session, so the proxy_auth_allow_insecure acknowledgement is unnecessary. Previously both http:// and https:// proxies required it, producing a misleading cleartext-risk diagnostic for a path that is already encrypted. Skip the requirement when the proxy URL uses https://; the acknowledgement is still tolerated if set. Updated in both the supervisor and Podman driver validation paths, with docs and tests. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix: format Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(kubernetes): use truly unsupported scheme in proxy validation test https:// is now a supported proxy scheme after a13c4dce, so the unsupported-scheme test must use a genuinely unsupported scheme. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(e2e): sign the https proxy fixture listener cert with a CA The corporate-proxy E2E fixture served a single `openssl req -x509` certificate as its TLS listener identity. OpenSSL marks that certificate `basicConstraints: critical, CA:TRUE`, and rustls refuses a CA certificate presented as an end-entity certificate (CaUsedAsEndEntity). The supervisor's TLS handshake with the proxy therefore failed, the upstream dial errored, and the workload's CONNECT was dropped without a response, so podman_corporate_proxy_trusts_ca_bundle_for_https_proxy failed on the approved destination while policy denial still worked. Generate a corporate CA and a separate listener leaf signed by it, serve the leaf chain, and publish only the CA as the bundle the supervisor trusts. This is what an intercepting proxy actually presents, and it exercises the corporate-CA trust path rather than pinning the listener certificate itself. Refs #1792 Signed-off-by: Philippe Martin <phmartin@redhat.com> --------- Signed-off-by: Philippe Martin <phmartin@redhat.com> |