mirror of
https://github.com/NVIDIA/OpenShell.git
synced 2026-10-02 07:34:45 +08:00
main
298
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
2263685cf3 |
test(tmachine): migrate Keycloak provider refresh coverage (#3404)
* test(tmachine): add Keycloak provider refresh suite Signed-off-by: Evan Lezar <elezar@nvidia.com> * refactor(tmachine): share container runtime detection Signed-off-by: Evan Lezar <elezar@nvidia.com> * ci(tmachine): run feature suites in GitHub Actions Signed-off-by: Evan Lezar <elezar@nvidia.com> * ci(tmachine): run conformance with Podman tests Signed-off-by: Evan Lezar <elezar@nvidia.com> * ci(tmachine): cover provider refresh with Podman Signed-off-by: Evan Lezar <elezar@nvidia.com> * ci(integration): split input preparation from runners Signed-off-by: Evan Lezar <elezar@nvidia.com> --------- Signed-off-by: Evan Lezar <elezar@nvidia.com> |
||
|
|
1d010f4187 |
feat(sandbox): validate configuration before workload activation (#3259)
* feat(sandbox): validate configuration before workload activation Closes #3145 Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(sandbox): bound startup failures and preserve activation history Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(sandbox): enforce deadlines on startup RPC attempts Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(sandbox): isolate provider auto-create policy fixture Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(sandbox): remove unnecessary fixture string delimiters Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(sandbox): capture startup logs before ephemeral cleanup Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(sandbox): preserve credential revocation after rebase Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(server): retain provider revision helper for endpoint reports Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(server): import provider object trait in production Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(supervisor): preserve admission across boundary extraction Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(sandbox): use existing boundary discovery imports Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(sandbox): distinguish workload and supervisor containers Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(server): reconcile admission schema with timestamp migration Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(server): reconcile admission with typed deletion schema Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(server): reconcile admission with mutation request IDs Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(supervisor): adapt local startup fixture to admission state Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(supervisor): deduplicate startup quarantine diagnostics Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(tui): show configuration blockers in sandbox notes Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * feat(sandbox): expire provisioning repair attempts after five minutes Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(supervisor): reconcile admission with provider readiness Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(tui): separate configuration summaries from full diagnostics Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(tui): shorten invalid configuration note Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(server): refresh schema inventory after rebase Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> --------- Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> |
||
|
|
04146692d9 |
refactor(providers)!: make provider profiles import-only (#3383)
* chore(providers): remove dead provider plugin modules Twelve modules under crates/openshell-providers/src/providers/ were never declared in providers/mod.rs, so they have not been compiled since the plugin registry was narrowed to the two adapters it still registers. Four of them (generic, gitlab, opencode, outlook) key off provider type identifiers that normalize_provider_type already retired. Keep google_cloud and vertex, which are the only plugins ProviderRegistry::new registers. Signed-off-by: Philippe Martin <phmartin@redhat.com> * test(providers): load example profiles from providers/ at test time Add an example_profiles module that reads the YAML under providers/ from the source checkout at run time and parses it with the existing profile loader. It is gated behind a new non-default example-profiles feature so it is available to this crate's own tests and, once wired into dev-dependencies, to gateway and CLI tests, while never reaching a release binary. Nothing consumes it yet; later changes move the test fixtures off the compiled catalog and onto these files. Signed-off-by: Philippe Martin <phmartin@redhat.com> * test: source profile fixtures from providers/ instead of the compiled catalog Unit and integration fixtures reached into builtin_profiles() to get a profile to work with, which ties the tests to the compiled catalog rather than to the files an operator would import. Point them at the example_profiles loader instead. The profiles.rs tests keep their coverage unchanged and become explicit golden tests over providers/*.yaml: they are what keeps those files valid once nothing compiles them. builtin_profiles_are_sorted_by_id widens into a test that the whole example set parses, sorts, and lints clean as one catalog. The CLI fake gateways serve the example profiles the way a real gateway serves what an operator imported, through a shared helper. No behavior change: the gateway still loads the built-in source by default and serves the same profiles. Signed-off-by: Philippe Martin <phmartin@redhat.com> * docs(providers): document the example provider profiles Each file under providers/ now opens with a header naming its expected client binary identities, the image layout those paths assume, the credential scope, the endpoint access it grants, and a smoke test. Add a README covering the import commands and why a profile should be copied and edited rather than imported unchanged. Several of these profiles bind network access to paths that only exist in the OpenShell Community image — /sandbox/.venv, /app/.venv, /sandbox/.cursor-server, /usr/lib/node_modules. Imported unchanged into another image the profile matches nothing: the catalog still advertises it, but the credential is never injected and the traffic is denied. The headers say so where it applies. Comments only; the profile schema has no documentation fields and all fifteen files still lint clean. Signed-off-by: Philippe Martin <phmartin@redhat.com> * test(server): seed unit-test state with the example provider profiles Server unit tests inherited the builtin + user source default from ServerState, so around a hundred and forty assertions about github, openai and the rest resolved against the compiled catalog. Point the test state at the user source alone and import the example profiles from providers/ into its store first, the way an operator would. Every one of those assertions keeps passing unchanged, which is the point: it proves the gateway behaves identically with an imported catalog before the default moves. Four tests asserted builtin-source semantics specifically. Under import-only the only profiles a gateway cannot edit are the ones a non-user source vends, so they now exercise a source-managed profile composed with the user source; the read-only guard they cover is the one that still applies to interceptor catalogs. The list test asserts the imported profile's user/platform identity instead of builtin with an empty scope. test_server_state_with_user_only_github_profile is gone: the default test state now is a user-only gateway with github imported. Signed-off-by: Philippe Martin <phmartin@redhat.com> * refactor(cli): resolve credential suggestions from the gateway catalog The --env credential warning scanned a profile table compiled into the CLI, so its suggestions described the binary rather than the gateway the user is talking to: a profile the gateway does not serve was suggested anyway, and an imported custom profile never was. Fetch the catalog from the connected gateway instead. The warning moves out of argument parsing and into sandbox create and sandbox template create, where a client already exists. A catalog fetch failure is not fatal — the warning degrades to its generic form rather than blocking sandbox creation. Extract the ListProviderProfiles paging loop from provider list-profiles into a shared fetch_provider_profile_catalog, which the profile-driven paths now share. Signed-off-by: Philippe Martin <phmartin@redhat.com> * refactor(cli)!: infer providers from the gateway catalog Command-to-provider inference went through a hardcoded alias table in openshell-providers, a second copy of the built-in catalog's identifiers. A custom profile could never be inferred no matter what binaries it declared, and the table drifted from the profiles it mirrored. Infer from the connected gateway's catalog instead: match the command's basename against each profile's ID and against the basenames of the binaries the profile authorizes. A profile that names /usr/bin/claude is the profile for running claude. The match must be unique — where several profiles claim a command, the user names one with --provider — and an empty catalog infers nothing. No command is special-cased. `binaries` is the operator's authorization statement, so a profile that declares a binary claims the command that runs it, whatever that binary is; narrowing that belongs in the profile rather than in a list compiled into the CLI, which could never cover an unbounded catalog anyway. Breaking: the retired aliases stop resolving, and commands the old table never listed can now infer. git, pip and uv are declared by the github and pypi example profiles, so they infer where they previously did not. Signed-off-by: Philippe Martin <phmartin@redhat.com> * refactor(providers,server)!: resolve provider profiles by exact ID normalize_provider_type was a hardcoded alias map — gh to github, claude to claude-code, vertex to google-vertex-ai — and a second copy of the built-in catalog's identifiers, independent of the YAML it mirrored. It let a provider type resolve to a profile the operator never named, and it made the built-in IDs behave as a reserved namespace. Remove it, along with detect_provider_from_command and the alias-normalizing ProviderRegistry::inject_env. A provider type now names a profile exactly: - the effective catalog resolves an ID or reports it absent, with no alias retry - plugins activate only for the ID of a profile the gateway resolved; a provider with no resolvable profile gets no plugin projection, instead of falling back to an alias guess - the CLI surfaces the gateway's not-found instead of retrying under an alias - telemetry buckets by profile ID, and the gitlab, opencode and outlook buckets go with the aliases that were their only source Breaking: `--type gh`, `--type claude` and the other aliases no longer resolve. Use the profile's own ID. Signed-off-by: Philippe Martin <phmartin@redhat.com> * feat(server): fail closed when a provider's profile is absent A sandbox composed from a provider whose profile the gateway cannot resolve started anyway: the credential and policy builders warned and skipped, so the sandbox came up carrying none of that provider's credentials or network policy. The operator learned about it later, as a denied connection or a missing environment variable, rather than as the configuration error it is. Check at the two composition boundaries — CreateSandbox and AttachSandboxProvider — right beside the catalog snapshot already taken there, and reject with a bounded diagnostic naming the provider, the profile it refers to, and the import command that supplies it. Scope-aware: a platform-scoped provider is told to import with --global. Read paths are untouched. ListProviders, GetProvider and profile export keep working so an operator can see and recover an affected provider, and the shared policy builders keep their warn-and-skip for the diagnostic paths that also reach them. Signed-off-by: Philippe Martin <phmartin@redhat.com> * refactor(server): scope the vendor base-URL pin to declared endpoints provider_profile_endpoints_are_active withheld a credential and the provider's policy layer when an openai or anthropic provider pointed its client somewhere other than the public vendor endpoint. The guard was keyed on profile.source == "builtin" and on those two profile IDs, so it protected only profiles OpenShell shipped. Once profiles are import-only no profile is ever builtin, and the guard would silently stop applying — including to an operator who imported providers/openai.yaml verbatim. Key it on the profile instead of on where the profile came from. A profile's endpoints are the boundary its credential is bound to, so if the provider configures a *_BASE_URL pointing at a host the profile does not declare, the profile no longer describes where that credential goes and is treated as endpointless. Host matching reuses the DNS-label-aware matcher in openshell-core, so wildcard endpoints such as Vertex's *-aiplatform.googleapis.com resolve correctly. A profile with no declared endpoints has no boundary to contradict, and a config value that names no host is not a redirect this can reason about. Both keep the profile active. The control now covers every endpoint-bearing profile, including an operator's own, and the last "builtin" string leaves the gateway. Signed-off-by: Philippe Martin <phmartin@redhat.com> * test(e2e): import example provider profiles during gateway bring-up The e2e suites create providers from github, openai, nvidia, claude-code and google-cloud, and the GitHub lane runs a real git clone through the profile's binary attribution. A gateway serves only the profiles an operator imported, so the lanes have to import them. Add e2e_import_example_provider_profiles to the shared bring-up helpers and call it from the Docker, Podman, Kubernetes and VM wrappers once the gateway is healthy and registered. It runs the same command the upgrade notes give operators, against the repository's own providers/ directory, so the lanes exercise the documented path rather than a test-only shortcut. Lands before the default changes: a same-ID user profile already shadows a built-in, so importing works today. Signed-off-by: Philippe Martin <phmartin@redhat.com> * refactor(providers)!: make provider profiles import-only OpenShell compiled fifteen provider profile YAML files into every release binary and selected them by default, so a fresh gateway published a catalog it was never configured with. Those profiles are not image-neutral: their binary selectors name paths that exist in the OpenShell Community image, so changing the sandbox image could make a profile inert while the catalog still advertised it. Every endpoint, credential name and binary path in providers/ was also effectively part of the 0.1.0 public contract. A gateway's catalog is now exactly what an operator imported: - providers/*.yaml is no longer include_str!'d, and builtin_profiles() is gone from the openshell-providers API - the default provider_profile_sources is [{ type = "user" }], and a gateway with nothing imported reaches ready and serves an empty catalog — an empty catalog is a valid state, not a startup failure - the builtin source type is removed from the configuration schema, and a gateway.toml that still names it is rejected at parse time with the import command rather than an unknown-variant error - nothing reserves the canonical identifiers any more, so github, pypi, anthropic and the rest import at their own IDs and the imported profile is the only definition for that ID; the static_fallback that kept a shipped definition resident behind an imported one is gone with the source that produced it A collision between an interceptor-vended profile and an imported one still fails closed, which is the behavior that shadowing quietly bypassed for built-ins. Operators upgrading should export the profiles their deployment relies on before upgrading, or copy them from the providers/ directory of the matching release tag, then import them at the scope their providers use. Signed-off-by: Philippe Martin <phmartin@redhat.com> * docs(providers): document import-only provider profiles Rewrite the provider profile documentation around a catalog the operator builds rather than one the gateway ships. - gateway-config: the default is [{ type = "user" }], the builtin source type is gone and rejected at startup, and an empty catalog is a valid ready state - profiles: replace the built-in profile table and its shadowing semantics with the import workflow, and say plainly that a profile whose binary paths do not match the image is inert - manage-providers: replace the two fixed provider-type tables with `provider list-profiles`, and describe catalog-driven command inference - inference-routing: drop the "still loads built-in profiles for compatibility" transition text; its migration walkthrough is now the normal path - quickstart and the Docker Compose, GitHub, AWS, Google Cloud and Vertex tutorials: import the profile before creating the provider, since nothing resolves without it - release notes: a 0.1.0 migration section covering export-before-upgrade, importing from the release tag's providers/ directory, what happens to a provider whose profile is missing, and the removal of the legacy type aliases - README, architecture, the openshell-cli and debug-openshell-cluster skills, and the governance interceptor example follow the same change; the example's smoke assertion now imports a profile to prove an authoritative interceptor hides it Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(cli,e2e): distinguish catalog failures and authenticate OIDC seeding Addresses two findings from review (GATOR-c1bd9867-02 and -03). The provider profile catalog lookup collapsed its error into an empty catalog, so an authorization, availability or transport failure on ListProviderProfiles was indistinguishable from a gateway that genuinely has no profiles. Command inference then resolved nothing and the sandbox was created without the provider it needed, deferring the failure to the workload. Keep the lookup's outcome instead of discarding it. The credential warning is advisory and still degrades to its generic form, but inference now consults the catalog only when there is a command to resolve and surfaces the lookup failure when there is, naming the gateway and pointing at explicit --provider selection. Two regression tests cover it through a fake gateway whose ListProviderProfiles returns UNAVAILABLE while every other RPC succeeds: sandbox creation fails without sending a provider-less create request, and a sandbox with no trailing command still succeeds because it needs no catalog. The e2e profile seeding also ran unauthenticated in the OIDC lanes. Those lanes deliberately skip gateway registration and start the gateway without a TLS client CA, so no mTLS identity exists and no token has been acquired when the import runs; the wrapper exited during setup. Skipping the import is not sufficient because the provider tests now require the claude-code profile. Add e2e_register_oidc_admin_session, which mints an administrator token with Keycloak's password grant — the same grant the OIDC test helpers use — and writes the gateway metadata and token bundle that an interactive login would have stored, so the import runs as an authenticated administrator. The mTLS lanes keep the direct import unchanged. Both affected wrappers are covered: with-podman-gateway.sh had the same defect as with-docker-gateway.sh. Signed-off-by: Philippe Martin <phmartin@redhat.com> * refactor(cli)!: remove command-derived provider attachment Addresses GATOR-c1bd9867-01. A profile's `binaries` list authorizes a binary to reach that profile's endpoints. It is not a statement that running the binary asks for the provider, and reading it as attachment intent let a command silently gain provider authority: the aws-s3 example declares /bin/bash, so `sandbox create -- bash` resolved an existing aws-s3 provider and attached its credential-backed capability and network policy to the shell and its descendants. Attachment of an already-created provider never prompted, so the confirmation flow did not guard it. The distinction that would make inference sound — whether a declared binary is a profile's client or merely a permitted runtime — cannot be expressed: NetworkBinary carries only a path, and the field that encoded it was removed in 0.1.0. Any substitute is a guess. Restricting the guess to a unique claimant does not help, because uniqueness measures how sparse the catalog is rather than what the user intended, and a compiled list of "generic" commands could never cover an unbounded operator catalog. Remove trailing-command inference rather than approximate it. A provider is attached only when named with --provider, which still creates a missing provider from local discovery when the name matches an imported profile ID. Sandboxes with no providers remain a normal, fully supported state. Removing inference also settles GATOR-c1bd9867-02: with no consumer deriving authority from the catalog, its only remaining use is the advisory credential warning, so a failed lookup degrades that warning instead of blocking creation. The regression test now asserts that an unreachable catalog still creates the sandbox and attaches nothing. While repurposing the deduplication test, a pre-existing defect surfaced: repeating a name in --provider auto-created it twice and the second attempt failed with "provider already exists", because the explicit pass never consulted the set of names it had already handled. Guard it. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(e2e): load the OIDC token and trust the gateway when seeding profiles Addresses the carried finding GATOR-c1bd9867-03. e2e_register_oidc_admin_session established a session the CLI never used. The CLI decides whether to load a stored bearer token by matching on the auth_mode field of the gateway metadata alone; the helper omitted that field, so the metadata fell through to the default arm and oidc_token.json stayed on disk unread. The profile import went out unauthenticated exactly as it had before the helper existed. The helper also installed no trust anchor, leaving the CLI unable to verify the gateway's self-signed serving certificate. Write auth_mode = "oidc" so the stored token is loaded, and install the CA at <gateway>/mtls/ca.crt. Only the CA is installed: with no client certificate or key on disk the CLI falls back to CA-only server verification and authenticates with the bearer token, which is what these lanes need because they start the gateway without --tls-client-ca. Certificate verification stays on; no insecure transport override is introduced. Assert the session before anything depends on it. ListProviderProfiles is annotated auth_mode: "bearer", so it cannot succeed unless the token was loaded and accepted. The helper now fails at that point, naming the gateway config directory and echoing the CLI output, rather than letting the defect surface later as an opaque profile import error. Both affected wrappers pass the PKI directory and CLI binary the helper needs; the mTLS lanes are untouched. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(e2e): request the OpenShell scopes when minting the admin token The OIDC lanes failed at the session assertion added in b25dd0298 with PERMISSION_DENIED and "scope 'provider:read' required". Authentication was working -- the gateway logged ListProviderProfiles at gRPC status 7, which it can only reach once the bearer token has been loaded and accepted. The token simply carried no OpenShell scope. sandbox:*, provider:*, config:*, workspace:* and openshell:all are optional client scopes on the openshell-cli client in scripts/keycloak-realm.json, so Keycloak mints them only when the request asks for them. The password grant here asked for nothing, leaving the realm defaults (openid, profile, email, roles, web-origins, acr) and an access token that authorizes no RPC. Request "openid openshell:all", as e2e/python/oidc/oidc_auth_test.py already does for its administrator tokens. openshell:all is SCOPE_ALL in crates/openshell-server/src/auth/authz.rs, so one scope covers the setup calls without enumerating them. Verified against the realm: the token goes from "email profile openid" to "email profile openshell:all openid". Record the same scopes in metadata.json. The helper writes the bundle openshell gateway login would have stored, and oidc_scopes is the field that login path reads back, so leaving it out would misdescribe the stored token. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(e2e): seed provider profiles once per VM lane gateway The VM lanes failed on the second test target with "custom provider profile 'anthropic' already exists" for all fifteen example profiles, and the import exited non-zero. e2e_import_example_provider_profiles was called from run_e2e_test, so it ran once per target -- four times against one long-lived gateway. That was harmless while import overwrote silently, but profiles are now import-only: ImportProviderProfiles is create-only and reports an existing id as an error-severity diagnostic, with no overwrite flag on the request. The first import therefore succeeds and every later one fails. Hoist the call to just after the conformance run, which is where the docker, podman and kube lanes already seed their catalogs. The profiles persist for the gateway's lifetime, so every target still finds them, and the E2E_TEST_OVERRIDE path is covered by the same single call. Signed-off-by: Philippe Martin <phmartin@redhat.com> * test(e2e): name the provider explicitly in the OIDC workspace-user test user_can_create_sandbox_with_inferred_provider_command reached the gateway once the admin session was fixed, and then failed: it asserted a missing-provider error, but the sandbox was created and died provisioning with "failed to spawn sandbox entrypoint process 'claude-code'". The test drove provider resolution by passing claude-code as the trailing command and relying on the CLI to infer the provider type from it. That inference is what this branch removed, so the trailing word is now nothing but an entrypoint, and the image has no such binary. The regression the test guards is not inference itself: it is that a workspace user resolving a provider is not gated behind Platform Admin. Name the provider with --provider, the only remaining way to attach one. The CLI still has to fetch the claude-code profile before it can auto-create the provider, so the lookup a workspace user must be allowed to make still happens, and auto-creation still stops at the non-interactive branch with "missing required provider". Both assertions therefore keep their meaning. Rename the test and rework its comments to describe what it now exercises; the old name would otherwise outlive the behavior it was named for. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(providers): reconcile import-only profiles with main Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> --------- Signed-off-by: Philippe Martin <phmartin@redhat.com> Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> Co-authored-by: John Myers <9696606+johntmyers@users.noreply.github.com> |
||
|
|
b7a932a4d0 |
docs(inference): remove stale managed endpoint references (#3428)
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> |
||
|
|
a72351d370 |
fix(policy)!: require explicit L7 append targets and scope (#3380)
Make allow and deny appends identify a rule and endpoint and declare every affected binary and port. Reject incomplete, stale, ambiguous, and provider targets atomically so a small append cannot silently change a broader scope. Update CLI previews, wire requests, Go types and generated bindings, SDK regressions, operator documentation, and live policy-update coverage. BREAKING CHANGE: AddAllowRules and AddDenyRules require L7RuleTarget instead of host and port. CLI appends require a rule name and explicit binary scope. Signed-off-by: Shiju <shiju@nvidia.com> |
||
|
|
cb93f62bfe |
fix(e2e): support distroless supervisor fixture (#3431)
* fix(e2e): support distroless supervisor fixture Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * test(e2e): preserve provider fixture trust bundle Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> Signed-off-by: Drew Newberry <anewberry@nvidia.com> Co-authored-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
9708ba9999 |
feat(providers): report applied sandbox provider changes (#3391)
* feat(providers): report applied sandbox provider changes Record exact provider mutation targets in shared configuration operations. Require authenticated evidence that credentials, effective policy, and the workload launch environment have been installed before reporting readiness. Add bounded CLI and Rust SDK status and wait support, preserving ordinary revision-scoped references for existing processes. Verify new-client rotation and acknowledged detach revocation without external-stable resolver changes. Signed-off-by: Shiju <shiju@nvidia.com> * fix(providers): align readiness times with protobuf contracts Represent readiness receipts, status, and operation times with Timestamp and report intervals with Duration. Reserve the scalar field tags, update all consumers and generated bindings, and preserve timestamp presence and nanosecond identity through storage and client validation. Qualify both empty-map constructors in the Linux boundary test so its module compiles while retaining the explicit default required by Clippy. Signed-off-by: Shiju <shiju@nvidia.com> * fix(cli): preserve provider mutation storage uncertainty Recognize the gateway's exact structured storage-uncertainty reason for provider attach, detach, and update. Explain that the change may already be saved and must be reconciled before retrying, without exposing server messages or metadata. Preserve uncertainty ahead of generic retry hints. Exercise saved mutations through the CLI and verify single submission, redaction, missing receipt handling, and untrusted error-detail rejection. Document the recovery guidance for users and the public CLI skill. Signed-off-by: Shiju <shiju@nvidia.com> * fix(cli): explain denied provider profile lookups Report exact and alias profile lookup denials with fixed permission and workspace guidance. Keep backend details redacted and stop before provider mutations. Cover denied create and update calls through the CLI. Verify the complete provider list independently in the cross-workspace OIDC regression, extracting its JSON object from surrounding startup diagnostics. Signed-off-by: Shiju <shiju@nvidia.com> --------- Signed-off-by: Shiju <shiju@nvidia.com> |
||
|
|
8de26878f9 |
fix(supervisor): serialize child registration with reaping (#3142)
* test(supervisor): reproduce fast child reaping race Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(supervisor): serialize child registration with reaping Signed-off-by: Drew Newberry <anewberry@nvidia.com> Signed-off-by: Evan Lezar <elezar@nvidia.com> * test(supervisor): cover child registration reaper race Signed-off-by: Evan Lezar <elezar@nvidia.com> --------- Signed-off-by: Evan Lezar <elezar@nvidia.com> Signed-off-by: Drew Newberry <anewberry@nvidia.com> Co-authored-by: Evan Lezar <elezar@nvidia.com> |
||
|
|
3dd5ce3314 |
test(parity): align fixture supervisor provenance (#3387)
The deterministic fixture still modeled Alpine after the supervisor image switched to Debian, causing the wrapper to exit before producing parity evidence. Update its image expectations and launch evidence to match the current harness contract. Force the intentional replacement of the read-only staged artifact so BSD mv does not prompt when the deterministic parity test runs from a terminal. Signed-off-by: Kris Hicks <khicks@nvidia.com> |
||
|
|
a316fd7832 |
feat(api): add durable workspace mutation admission and replay (#3321)
* feat(api): add durable workspace mutation admission and replay Part of #3051 (phase 3a). Preserve unresolved admissions, reauthorize replay, and protect same-name replacements. Signed-off-by: Mrunal Patel <mrunalp@gmail.com> * feat(api): extend mutation replay through gateway interceptors (#3323) Add typed durable replay receipts for 24 ordinary unary mutations, protect sensitive payload fingerprints, and revalidate intercepted retries without repeating post-commit observation. Part of #3051 (phase 3b). Signed-off-by: Mrunal Patel <mrunalp@gmail.com> * fix(api): scope workspace request IDs by target Include the requested workspace name in create/delete admission keys while leaving workspace UUID guards unset. Cover cross-target UUID reuse, replay, and missing targets with server and live gateway regressions. Merge the latest phase-two SDK fixes and preserve the approved interceptor replay changes. Refs #3051. Signed-off-by: Mrunal Patel <mrunalp@gmail.com> --------- Signed-off-by: Mrunal Patel <mrunalp@gmail.com> |
||
|
|
7e7a8d5610 |
fix(sandbox): pass declared environment to the initial process (#3392)
Apply the workload environment before stripping supervisor-only values and injecting provider placeholders in both canonical launch paths. The RFC 0012 bootstrap carries these values separately from inherited supervisor environment. Add child-process coverage for declared values, explicit HOME, stripped identity material, and provider precedence. Add a Docker-backed lifecycle regression for initial and exec environments in TTY and non-TTY modes. Validation: the child-process test fails without env application; the Docker regression fails with the pinned unpatched runtime and passes with the rebuilt runtime. Sandbox library tests: 201 passed, one ignored on Rust 1.95.0. mise run pre-commit passed. Full mise run ci is blocked by the host linker missing libz3. Fixes #3377. Signed-off-by: Carlos Villela <cvillela@nvidia.com> |
||
|
|
0fd3385726 |
fix(container): use distroless Debian 13 for supervisor (#3393)
* fix(container): use distroless Debian 13 for supervisor Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(docker): preserve supervisor bootstrap ownership Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> --------- Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> |
||
|
|
c502be9fd7 |
feat(api): return typed deletion outcomes with explicit missing-target semantics (#3317)
* feat(api): return typed deletion outcomes with explicit missing-target semantics Implement phase 2 of #3051 across the public gateway API and first-party SDKs. Preserve asynchronous sandbox deletion and observed resource identities, reserve legacy wire fields, and document the coordinated migration. Signed-off-by: Mrunal Patel <mrunalp@gmail.com> * fix(sdk): bind deletion waits to sandbox identity Track accepted deletions by original sandbox identity in Rust and TypeScript. Preserve name-only waits and cover replacement races and lookup failures. Merge the phase-one cleanup fix and adapt its interceptor regressions to typed deletion outcomes. Refs #3051. Signed-off-by: Mrunal Patel <mrunalp@gmail.com> * test(e2e): accept asynchronous stopped sandbox deletion Allow either completed or accepted deletion output, then continue polling for actual sandbox absence. This preserves the lifecycle assertion under the phase-two deletion contract. Refs #3051. Signed-off-by: Mrunal Patel <mrunalp@gmail.com> --------- Signed-off-by: Mrunal Patel <mrunalp@gmail.com> |
||
|
|
d68b7069c3 |
refactor(proto)!: use well-known time types (#3113)
* refactor(proto)!: use well-known time types Signed-off-by: Derek Carr <decarr@redhat.com> * fix(proto): preserve time migration behavior Signed-off-by: Derek Carr <decarr@redhat.com> * fix(proto): preserve timestamp boundary semantics Signed-off-by: Derek Carr <decarr@redhat.com> * fix(proto): convert sandbox token expiry to timestamp Signed-off-by: Derek Carr <decarr@redhat.com> * fix(proto): preserve time compatibility semantics Signed-off-by: Derek Carr <decarr@redhat.com> * fix(proto): preserve exact endpoint and profile times Signed-off-by: Derek Carr <decarr@redhat.com> * test(e2e): use duration for interactive exec timeout Signed-off-by: Derek Carr <decarr@redhat.com> * fix(sdk-go)!: remove legacy profile duration fields BREAKING CHANGE: Go provider profile callers must use RefreshBefore, MaxLifetime, and CacheTTL with ProfileDuration instead of the whole-second fields. Signed-off-by: Derek Carr <decarr@redhat.com> --------- Signed-off-by: Derek Carr <decarr@redhat.com> |
||
|
|
2ccef97769 |
feat(policy): establish one canonical authored policy representation (#3334)
* chore(policy): restart schema implementation Signed-off-by: Johnny Greco <jogreco@nvidia.com> * feat(policy-schema): add canonical authored policy model Signed-off-by: Johnny Greco <jogreco@nvidia.com> * refactor(policy): use canonical authored schema Signed-off-by: Johnny Greco <jogreco@nvidia.com> * refactor(prover): project canonical policy documents Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): describe shared schema boundary Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(policy): close reviewed parser gaps Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(policy-schema): fail closed on unsupported fields Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(policy): preserve partial process identities Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(policy-schema): harden authored policy inspection Signed-off-by: Johnny Greco <jogreco@nvidia.com> * refactor(policy): rename policy schema crate Signed-off-by: Johnny Greco <jogreco@nvidia.com> * revert(policy): restore policy schema crate Signed-off-by: Johnny Greco <jogreco@nvidia.com> * test(e2e): serialize OIDC PKCE scenarios Signed-off-by: Johnny Greco <jogreco@nvidia.com> --------- Signed-off-by: Johnny Greco <jogreco@nvidia.com> |
||
|
|
b3e4ad4579 |
fix(ci): repair RFC 0012 post-merge checks (#3360)
* fix(ci): lock nextest for Windows ARM64 Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): preserve GPU access in sandbox runtime Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(windows): restore host proxy identity binding Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(windows): satisfy cross-platform network lint Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(windows): satisfy MXC lint Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(windows): exclude Unix supervisor runtime Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(mxc): satisfy Windows test lint Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(windows): escape proxy test policy paths Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): accept admitted GPU runtime groups Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(docker): serialize proxy pipeline scenarios Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
c1f2e7189f |
feat(isolation): implement the RFC 0012 sandbox architecture (#2942)
* feat(isolation): add RFC 0012 backend contract Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> * refactor(isolation): name the interface crate explicitly Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): expose trusted host gateway Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(agents): inventory the MXC driver Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): add mediated DNS transport Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): tighten interface error and digest contracts Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(isolation): remove unrelated driver inventory Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): define capability-free launch contract Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): seal confirmed boundary state Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): validate confirmation for external backend implementations Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): clarify mediated DNS identity Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(isolation): generalize loopback connector Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(isolation): unify typed network mediation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): bind launches to sandbox sessions Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(mxc): initialize extended sandbox status Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): add boundary protocol and Linux primitives Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): harden signals and separate process status from transport Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): validate remote confirmation through public contract Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): validate wire state and propagate snapshot failures Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(isolation): import owned agent specification explicitly Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(isolation): describe mediated DNS channel Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): bound mediation attach without nested retries Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(isolation): generalize loopback protocol Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): add transport-neutral session authentication Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(isolation): separate sandbox backend protocol Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): harden runtime boundary controls Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): add terminal boundary operation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): split supervisor and sandbox runtimes Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): harden boundary isolation and lifecycle ownership Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): reject private root redirects and adopt typed errors Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): preserve accept thread ownership on musl Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(sandbox): isolate credential probes from filtered threads Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): return retained exec exit status to independent waiters Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): bound network mediation and preserve socket authorization Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): bound control admission and retire stale mediation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * ci(e2e): select migrated drivers per stack layer Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(sandbox): implement loopback connector Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): authenticate the Sandbox Protocol Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(supervisor): rotate launch-scoped authentication Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(sandbox): consume dedicated backend crate Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(sandbox): align topology session fixture Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): align projected bootstrap bundle Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(auth): validate refreshed credentials before rotation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): fail closed across supervisor disconnects Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): repair rebased sandbox CI Signed-off-by: Drew Newberry <anewberry@nvidia.com> * build(runtime): publish separate sandbox and supervisor images Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(config): configure the sandbox runtime image Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(ci): validate sandbox binary linkage Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(isolation): use backend and runtime terminology Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(sandbox): use a scratch runtime image Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(ci): refresh schema and dependency policy Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): bind reconnects to supervisor process Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs: align runtime split operational guidance Signed-off-by: Drew Newberry <anewberry@nvidia.com> * chore(security): document Kubernetes runtime RBAC Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): enforce runtime lifecycle invariants Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(compute): identify sandbox start generations Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(server): restore sandbox launch sessions Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): support authenticated runtime replacement Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(auth): bind sandbox session successors Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(auth): retry pending sandbox successors Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(vm): run the supervisor outside the guest workload Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(vm): use sandbox backend protocol Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): repair rebase integration Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): use unified build toolchain Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(vm): use sandbox runtime terminology Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(vm): own guest network bootstrap Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): expose guest init version Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): select native supervisor artifacts Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): guard guest init Linux symbols Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): scope Linux test imports Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): avoid guest interface casts Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): reconcile admitted sandbox identity Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): share resolved sandbox identity Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): surface host supervisor failures Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): include guest logs on supervisor exit Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): rotate and clean runtime generations Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): make sandbox starts generation-aware Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): rotate restored sandbox sessions Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(vm): keep shared paths in the base layer Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): bind sandbox session lineage Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(docker): isolate workloads behind the host supervisor Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(docker): rotate launch-scoped authentication Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(docker): use sandbox backend protocol Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(docker): use host networking for supervisor Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): preserve host gateway alias resolution Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(docker): use separate sandbox and supervisor images Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): restore startup validation after rebase Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(docker): name the sandbox runtime directly Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): narrow supervisor CA runtime storage Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): close companion isolation gaps Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(docker): align mediated network expectations Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(docker): exercise mediated network paths Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): attach supervisor to managed network Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): defer supervisor recovery until gateway is ready Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): make sandbox starts generation-aware Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): rotate restored sandbox sessions Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): preserve workloads during session rotation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(docker): remove unrelated configuration RFC changes Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): bind sandbox session lineage Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(kubernetes): add proxy-pod isolation topology Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(kubernetes): use sandbox backend protocol Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): use stable sandbox service authority Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(kubernetes): split sandbox and supervisor images Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): adapt proxy pods to current runtime APIs Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(kubernetes): describe the single runtime placement Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(kubernetes): simplify sandbox orchestration Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): validate deployment prerequisites Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): update Trivy Helm profile inventory Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(kubernetes): update Trivy scan inventory count Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): reuse preloaded runtime images in e2e Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): type and clean runtime resources Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): make sandbox restarts recoverable Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): rotate restored sandbox sessions Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): preserve supervisor egress Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): bind sandbox session lineage Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(podman): adopt isolated sandbox and supervisor containers Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): stage bootstrap archives at named volume destinations Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(podman): rotate launch-scoped authentication Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(podman): use sandbox backend protocol Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(podman): use host networking for supervisor Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(podman): split sandbox and supervisor images Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): repair rebase integration Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(podman): name the sandbox runtime directly Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): provision supervisor CA runtime storage Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): address isolation review findings Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): inspect Debian supervisor provenance Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): use libpod-compatible tmpfs options Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): bind verified sandbox runtime binary Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): provide external driver data directory Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): start sandbox before joining user namespace Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): separate supervisor user namespace Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): make sandbox starts generation-aware Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): rotate restored sandbox sessions Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): bind sandbox session lineage Signed-off-by: Drew Newberry <anewberry@nvidia.com> * perf(isolation): add TCP and DNS benchmark harnesses Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(perf): align benchmark timing and supported protocols Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(perf): report TCP benchmark metrics accurately Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(perf): cancel failed worker startup Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): build matching local supervisor image Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): make local sandbox smoke test runnable Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): wire local sandbox runtime image Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): narrow sandbox service RBAC Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(ci): validate split runtime artifacts Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): harden runtime session handling Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(supervisor): add standalone network proxy role Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(rfc): remove implementation companion notes Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(vm): standardize runtime release name Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): pin renamed runtime artifacts Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(auth): persist sandbox runtime identity Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(runtime): restore branch validation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): reconcile main after rebase Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(network): close unframed HTTP 1.0 responses Signed-off-by: Drew Newberry <anewberry@nvidia.com> * chore(isolation): preserve upstream OCSF updates Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(security): close credential and TLS replay paths Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(auth): make sandbox refresh retries idempotent Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
dbe36eaf85 |
fix(security): harden Vault credential transport (#3329)
Reject non-loopback plaintext Vault endpoints, disable redirects, and support private CA bundles without weakening hostname verification. Update Helm configuration, documentation, operator skills, and regression coverage for OSSR-002. Signed-off-by: Seth Jennings <sjenning@redhat.com> |
||
|
|
b799fccb8b |
fix(auth): harden OIDC trust root retrieval (#3332)
* fix(auth): harden OIDC trust root retrieval Signed-off-by: Mrunal Patel <mrunalp@gmail.com> * fix(e2e): pass OIDC HTTP acknowledgement value Signed-off-by: Mrunal Patel <mrunalp@gmail.com> --------- Signed-off-by: Mrunal Patel <mrunalp@gmail.com> |
||
|
|
42e9bcf2b0 |
feat(e2e): support the Vault credential-driver lane on OpenShift (#3312)
Implements https://github.com/NVIDIA/OpenShell/issues/3212 Running e2e:kubernetes with the Vault credential driver failed on OpenShift in two ways: the OpenBao fixture pod was rejected by the restricted-v2 SCC, and provider-creating tests hit HTTP 403 from auth/kubernetes/login because OpenBao's Kubernetes auth was provisioned only inside a single feature-gated test. - Deploy OpenBao with the chart's OpenShift mode (global.openshift=true) when OpenShift is detected, so the pod inherits a namespace-assigned, SCC-compliant security context with no manual SCC grant. Hoist OpenShift detection ahead of the credential-driver fixtures so the flag is set before the fixture is deployed. - Provision the OpenBao KV store, Kubernetes auth method, storage policy, and gateway login role in the harness (deploy_vault_fixture), making a Vault-backed gateway usable by the whole suite instead of only the credential_drivers test. Remove the now-redundant configure_vault_storage helper from the test. - Harden openbao_exec so it tolerates only the idempotent "path is already in use" error on reruns and fails fast with output on any other error, instead of a blanket `|| true` that masked genuine failures (e.g. an unresponsive pod) until a later cryptic write. - Document the OpenShift Vault credential-store SCC and Kubernetes-auth 403 troubleshooting in the debug-openshell-cluster skill. Signed-off-by: Jorge Garcia Oncins <jgarciao@redhat.com> |
||
|
|
5b9daab935 |
fix(ci): restore mise run ci on macOS (#3294)
- Replace BSD-incompatible in-place sed calls with portable temp-file rewrites. - Remove test-only shell interception and capture generated gateway config directly. - Allow parity tests to use supplied supervisor binaries without resolving a Linux target. - Normalize temporary-directory paths and use portable RPM config installation. - Set a valid setuptools-scm version for Python protobuf generation in Jujutsu checkouts. Signed-off-by: Kris Hicks <khicks@nvidia.com> |
||
|
|
02b664bb0d |
refactor(config): normalize and enforce gateway schema v2 (#2814)
* refactor(config): normalize compute driver field names Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * refactor(config): introduce canonical gateway fields Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * refactor(config): enforce gateway schema version 2 Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): preserve compute driver runtime guarantees Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): address schema v2 review regressions Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): complete schema v2 migration safeguards Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(config): expand schema v2 regression coverage Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(config): add schema v2 parity manifest Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): correct parity manifest inventory Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * docs(config): record schema v2 intentional changes Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * docs(config): disposition schema v2 parity gaps Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): add dual schema parity harness Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): establish compute lifecycle parity baseline Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): preserve gateway option compatibility Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): record gateway option parity Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * docs(config): close gateway-wide parity gaps Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(podman): apply configured pids limit Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): validate Podman option parity Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): add Kubernetes option parity harness Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): record Kubernetes option parity Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): disposition VM parity lanes Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): add external driver parity lane Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(e2e): preserve external driver pull policy Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): attest parity artifacts and launches Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): require clean parity build sources Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): bind parity runtime artifacts Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(e2e): use isolated supervisor tags Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(e2e): qualify parity image tags Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(e2e): serve parity supervisor locally Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): isolate parity podman services Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): harden parity evidence provenance Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): pin parity sandbox artifacts Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): attest parity runtime inputs Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): bind parity runtime evidence Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): record compute boundary parity Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): disposition cross-cutting parity lanes Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(packaging): preflight gateway config upgrades Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): preserve rebase integration guarantees Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(ci): isolate temporary git signing config Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): update remaining schema v2 consumers Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(ci): provide e2fs tools to VM tests Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): align preflight with gateway startup Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(vm): preserve rootfs tar configuration Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * chore(config): adopt duration unit constructors Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(packaging): preflight RPM gateway config Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): address driver review findings Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(e2e): require fresh semantic parity evidence Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(docker): update tests for renamed sandbox label Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(gateway): preserve selective driver coverage after rebase Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> Signed-off-by: Drew Newberry <anewberry@nvidia.com> Co-authored-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
1860010850 |
feat(sdk): add lazy pagination pagers (#3256)
* feat(sdk): add lazy pagination pagers Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(sdk): harden pager edge cases Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(sdk): cover initial resume token Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(go): fix all-workspaces pager examples Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
33bbda3d33 |
refactor(persistence): adopt continuation-token pagination (#3249)
* refactor(persistence): adopt continuation-token pagination Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(pagination): address continuation review findings Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(tui): recover completed list refreshes Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(pagination): address review scalability findings Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(pagination): repair branch validation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(go): use page size in template example Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
226a83b323 |
fix(supervisor): classify credential placeholders in request bodies (#3246)
* fix(supervisor): classify credential placeholders in request bodies Closes #2904 Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * fix(supervisor): preserve same-provider placeholders in request bodies Signed-off-by: John Myers <johntmyers@users.noreply.github.com> --------- Signed-off-by: John Myers <johntmyers@users.noreply.github.com> Co-authored-by: John Myers <johntmyers@users.noreply.github.com> |
||
|
|
90dbe5454b |
feat(api): add typed workspace selectors (#3245)
* feat(api)!: add typed workspace selectors Signed-off-by: Mrunal Patel <mrunalp@gmail.com> * fix(cli): preserve template workspace metadata Signed-off-by: Mrunal Patel <mrunalp@gmail.com> * test(e2e): migrate workspace request selectors Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(api): update public schema inventory Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> --------- Signed-off-by: Mrunal Patel <mrunalp@gmail.com> Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> Co-authored-by: John Myers <9696606+johntmyers@users.noreply.github.com> |
||
|
|
8211274bfb |
fix(e2e): follow credential storage identity (#3202)
Signed-off-by: Drew Newberry <anewberry@nvidia.com> Signed-off-by: Shiju <shiju@nvidia.com> Co-authored-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
ea8eda6d5b |
feat(supervisor): enforce MCP request protocol versions (#3241)
* feat(supervisor): enforce MCP request protocol versions Signed-off-by: Shiju <shiju@nvidia.com> * fix(supervisor): enforce MCP versions across HTTP forwarding Apply shared request-version guards before authorization and after forward-request rewriting. Require version metadata to survive HTTP header cleanup, and cover valid initialization and selected-revision forwarding through middleware. Signed-off-by: Shiju <shiju@nvidia.com> --------- Signed-off-by: Shiju <shiju@nvidia.com> |
||
|
|
f4dc6be4b2 |
refactor(inference): remove managed inference routes (#3195)
* refactor(inference): remove managed inference routes Closes #3172 Remove the inference route control plane, inference.local data path, built-in router crate, and SDK surface. Move inference workloads to explicitly imported provider profiles and native endpoints, with migration cleanup and updated tests and documentation. Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(policy): preserve alternate upstream isolation Restore the provider policy activation guard so legacy OpenAI and Anthropic providers configured for alternate base URLs do not grant egress to the built-in public vendor endpoints. Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> --------- Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> |
||
|
|
48c449d8c8 |
chore(deps): replace ring with AWS-LC (#3243)
* chore(deps): replace ring with AWS-LC Signed-off-by: Simon Scatton <sscatton@nvidia.com> * fix(lint): address warnings after dependency upgrades Signed-off-by: Simon Scatton <sscatton@nvidia.com> * fix(tls): limit provider initialization to reqwest clients Signed-off-by: Simon Scatton <sscatton@nvidia.com> --------- Signed-off-by: Simon Scatton <sscatton@nvidia.com> |
||
|
|
8af79a7f4b |
fix(podman): resolve macOS Podman socket dynamically (#3135)
* docs(podman): document macOS socket path mismatch and dynamic lookup On macOS, Homebrew-installed Podman does not create the default socket path that the Podman driver probes. Document the OPENSHELL_PODMAN_SOCKET override and the podman machine inspect lookup in both the compute drivers reference and the debug-openshell-cluster skill. Fixes #1690 Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com> * fix(podman): resolve macOS Podman socket dynamically Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com> * chore: restore debug skill file Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com> * chore: drop legacy debug skill path Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com> * fix(podman): trim unrelated e2e changes Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com> * fix(e2e): harden shell array expansion Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com> * chore: remove unrelated skill note * ci: retrigger checks Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com> --------- Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com> |
||
|
|
6e6b3c8905 |
refactor(cli): remove local Dockerfile image builds (#3214)
Signed-off-by: Evie Howard <evhoward@redhat.com> |
||
|
|
118b250f01 |
feat(sandbox): support rootfs tar as --from source for VM driver (#2863)
* feat(sandbox): support rootfs tar as --from source for VM driver Accept flat rootfs tar archives (.tar, .tar.gz, .tgz) via the --from flag for VM-backed gateways. The CLI detects the archive extension, validates that the gateway uses the VM compute driver, and passes the tar path through driver_config. The VM driver copies the tar into its staging area and feeds it into the existing rootfs extraction and ext4 disk creation pipeline, skipping the container image pull/export steps. Closes #2175 Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(sandbox): validate rootfs tar path at the VM driver boundary The rootfs_tar_path field in driver_config was passed from the API caller directly to tokio::fs::copy without validation. An authenticated user bypassing the CLI could supply arbitrary host paths (e.g. /dev/zero for disk exhaustion, or readable host files for data exfiltration). Introduce a trusted staging directory that the VM driver creates on startup and advertises via GetCapabilities. The CLI now copies the tar into the staging directory before creating the sandbox, and the driver validates that the received path is a regular file inside the staging root and within a configurable size limit (default 10 GiB) before any I/O. New VmDriverConfig options: - rootfs_tar_staging_dir: override the staging directory (default: <state_dir>/rootfs-tar-staging) - rootfs_tar_max_bytes: override the size limit (default: 10 GiB) Addresses GATOR-28b5152e-01. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(sandbox): request-scoped staging, size pre-check, and cleanup for rootfs tar Tighten the rootfs tar staging flow to address the remaining GATOR-01 obligations: - Request-scoped staging: the CLI creates a unique per-request subdirectory (req-<pid>) under the staging root instead of placing files directly in the shared directory. The driver enforces that the tar path is at depth 2 (staging_root/<subdir>/<file>), preventing cross-request path selection. - Size pre-check: the driver advertises rootfs_tar_max_bytes via GetCapabilities. The CLI reads this limit and rejects oversized files before copying, avoiding disk exhaustion in the staging directory. - Cleanup: the driver removes the request staging subdirectory after consuming the tar (on cache hit, copy success, or copy failure), ensuring staged data does not persist beyond the request. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(vm): restore rootfs-tar sandboxes from persisted image identity On restore or restart, the one-shot staged tar archive has already been cleaned up. Reading the persisted image identity from the sandbox state directory and resolving the cached disk path directly avoids re-accessing the deleted staging path. Addresses GATOR-168b9210-01. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(cli): use random staging dirs and enforce byte limit during rootfs tar copy Replace PID-based request staging directories with tempfile-generated random names to prevent collisions and make paths unpredictable. Replace bare tokio::fs::copy with a streaming copy loop that enforces the advertised max_bytes limit during transfer, closing the TOCTOU gap between the pre-copy size check and the actual copy. Signed-off-by: Philippe Martin <phmartin@nvidia.com> Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(sandbox): issue rootfs tar staging slots from the gateway A caller could name any host path in `driver_config.vm.rootfs_tar_path`, which the privileged VM driver then read. The CLI-side locality check did not apply to direct API requests. The gateway now owns staging. `BeginRootfsTarStaging` allocates a request-scoped directory under the driver-advertised staging root and returns an opaque single-use token; `CreateSandbox` carries the token, and the gateway substitutes the path it allocated before dispatching to the driver. `template.driver_config.<driver>.rootfs_tar_path` is rejected outright in request validation, so a caller-supplied path never reaches privileged I/O. Tokens are bound to the issuing workspace and subject, consumed once, and expire after 30 minutes. Outstanding slots are capped per caller and overall, so one caller can neither exhaust the staging filesystem nor starve others. An RAII guard reclaims the directory on every failure path after consumption, and an age-gated sweep runs at startup and on each reconcile pass for directories whose driver died before its own cleanup. The token is stripped from the public sandbox before persistence: the stored copy is returned verbatim by GetSandbox, ListSandboxes and WatchSandbox to every member of the workspace. Also fixes two defects this exposed: - The CLI wrote `rootfs_tar_path` at the top level of `driver_config`, but the gateway forwards only `driver_config.<driver_name>`, silently dropping unmatched keys. The archive never reached the VM driver, so the documented `--from ./rootfs.tar` flow did not work at all. Config is now nested under `vm` and deep-merged, so a caller's existing VM settings survive instead of being clobbered by a shallow extend. - Staging previously required `GetGatewayInfo`, which is restricted to `platform_admin`, making the feature unusable for ordinary users on any RBAC-enabled gateway. The new RPC matches CreateSandbox at `sandbox:write` / `workspace_role: user`. `compute_driver.proto` is unchanged; the gateway reads the staging root from the capabilities it already stores. Refs #2175 Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(vm): derive rootfs tar cache identity from archive contents The prepared-disk cache key combined the archive's full path with an mtime truncated to seconds, then mapped punctuation to `-`. Distinct paths such as `/tmp/a/b.tar` and `/tmp/a-b.tar` collapsed onto the same key and reused each other's disk, a rewrite within the same second kept stale contents, and a long path could exceed filesystem component limits. Identity is now a SHA-256 of the archive contents. This is also what makes the cache work at all now that the gateway allocates a fresh staging directory per request: a path-derived key would miss on every create. The archive is hashed, the cache checked, and only on a miss copied — so a hit skips writing a multi-gigabyte file. The copy is hashed as it is written and rejected if the digest differs from the first pass, which closes the window where the source changes during staging rather than approximating it with a re-stat. Refs #2175 Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(vm): decompress gzip rootfs tar archives during staging `--from` accepts `.tar.gz` and `.tgz`, but the driver staged whatever bytes it was given and the guest image-prep VM extracts the staged file with a plain `tar -xpf`. Compressed sources therefore depended on the guest tar auto-detecting gzip, and the prepared disk was sized from the compressed length, which is far too small for the expanded rootfs. Staging now detects gzip from the archive's magic bytes -- the driver only ever sees a gateway-issued staging path, never the caller's file name -- and writes an uncompressed tar. The digest still covers the source bytes, so the "archive changed while staging" check is unaffected, and expansion is bounded by `rootfs_tar_max_bytes` so a compression bomb cannot fill the host disk. `extract_rootfs_archive_to` sniffs gzip as well, so the host-side extraction path matches. Adds unit coverage for gzip staging, bounded expansion, and gzip extraction, plus an e2e sandbox created from a gzip-compressed export. Signed-off-by: Philippe Martin <phmartin@redhat.com> --------- Signed-off-by: Philippe Martin <phmartin@redhat.com> Signed-off-by: Philippe Martin <phmartin@nvidia.com> |
||
|
|
519e5eb35f |
feat(e2e): make e2e:kubernetes work transparently on OpenShift (#3183)
* feat(e2e): make e2e:kubernetes work transparently on OpenShift
Running `mise run e2e:kubernetes` on OpenShift required manual namespace
creation, SCC grants, Helm value overrides, and cleanup. A separate
`e2e:openshift` task existed but only checked pod readiness without
running the Rust e2e test suite, and even with the suite wired up the
SSH-relay `sandbox connect` path stalled to the ready timeout because
`kubectl port-forward` cannot carry round-trip-heavy SSH over the
internet.
The harness now auto-detects OpenShift via the `route.openshift.io` API
group and, on OpenShift, both configures the cluster and switches the
gateway transport automatically:
- Drives the gateway through a passthrough OpenShift Route secured with
mandatory mTLS instead of port-forward, so the connect suites
(live_policy_update, port_forward, sync, connect-based
sandbox_lifecycle, settings_management) actually pass. Computes the
Route host from the cluster ingress domain, extracts client mTLS
material from the openshell-client-tls secret, waits for the Route to
serve mTLS, asserts a certless caller is rejected at the TLS
handshake, and registers an mTLS CLI gateway pointing at the Route.
- Applies an SCC-compatible Helm values overlay that removes hardcoded
runAsUser/fsGroup, letting OpenShift assign UIDs from the namespace
range.
- Grants the privileged SCC to openshell-sandbox before Helm install
and removes it during cleanup.
- Grants the anyuid SCC to the PostgreSQL fixture service account in
DB scenarios and removes it during cleanup.
- All oc commands use --context to target the correct cluster.
The OpenShift e2e overlay (ci/values-openshift-e2e.yaml) turns TLS back
on, enables the Route, promotes the cert-verified caller to a dev
principal, and forces `image.pullPolicy`/`supervisor.image.pullPolicy`
to Always so runs against the `latest` upstream image use it instead of
a stale copy cached on the cluster nodes. Every OpenShift branch is
gated on OPENSHIFT_DETECTED, so the vanilla-Kubernetes port-forward path
is unchanged.
The Helm template for podSecurityContext is wrapped with {{- with }} so
null values omit the block instead of rendering invalid YAML.
The separate e2e:openshift task and e2e-openshift.sh script are removed
since e2e:kubernetes now covers OpenShift.
TESTING.md is updated with Kubernetes e2e documentation including
OpenShift auto-detection, dropping the e2e-host-gateway feature on
remote clusters, pinning IMAGE_TAG when the CLI and image versions
differ, task variants, and environment variables.
The debug-openshell-cluster skill gains an OpenShift platform row and
two SCC failure patterns (gateway rejected over hardcoded runAsUser,
sandbox missing the privileged SCC) covering the SCC handling and
podSecurityContext behavior this change introduces.
Signed-off-by: Jorge Garcia Oncins <jgarciao@redhat.com>
* fix(e2e): harden OpenShift SCC cleanup, mTLS gate, and Route timeout
Track the anyuid SCC grant for the PostgreSQL fixture with a dedicated
OPENSHIFT_POSTGRES_SCC_GRANTED flag set before the fixture apply, so a
failed apply no longer leaks the binding; cleanup now revokes it whenever
the grant succeeded, independent of deploy state.
Validate the Route server cert in the certless security gate (curl
--cacert instead of -k) and classify curl's exit code so only a TLS
client-auth rejection (35/56) counts as the expected certless rejection;
an unrelated DNS/timeout/TLS failure now fails loudly instead of masking
a potential mTLS hole.
Raise the OpenShift Route timeout in the e2e overlay. The default HAProxy
Route timeout is 30s, which severed long-lived transfers (large sandbox
upload/download, SSH-relay `sandbox connect`) mid-stream and failed the
sync e2e tests. Set both haproxy.router.openshift.io/timeout and
timeout-tunnel to 300s: a passthrough Route proxies in TCP mode, so
timeout-tunnel governs the established tunnel while timeout covers the
pre-tunnel phase.
Document the OpenShift transport exception, oc prerequisites and SCC
grants, and make the skopeo tag-check example copy-safe in TESTING.md.
Signed-off-by: Jorge Garcia Oncins <jgarciao@redhat.com>
---------
Signed-off-by: Jorge Garcia Oncins <jgarciao@redhat.com>
|
||
|
|
e4369adcd0 |
chore(deps): replace serde_yml with noyalib (#3031)
* chore(deps): replace serde_yml with noyalib Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(providers): annotate generic YAML test values Signed-off-by: Evan Lezar <elezar@nvidia.com> --------- Signed-off-by: Evan Lezar <elezar@nvidia.com> |
||
|
|
592df3e014 |
feat(policy): preserve exact MCP revision allowlists (#3027)
* feat(mcp): add version-aware wire profile metadata Signed-off-by: Shiju <shiju@nvidia.com> * feat(policy): canonicalize MCP version allowlists Signed-off-by: Shiju <shiju@nvidia.com> * fix(policy): align MCP policy tests with current main Signed-off-by: Shiju <shiju@nvidia.com> * fix(policy): canonicalize supervisor protobuf ingress Materialize defaultable MCP revisions before ambiguity checks, OPA construction, and sidecar delivery. Reject invalid sidecar policies with bounded errors. Signed-off-by: Shiju <shiju@nvidia.com> --------- Signed-off-by: Shiju <shiju@nvidia.com> |
||
|
|
1e1a8b5810 |
test(e2e): keep lifecycle sandboxes running (#3128)
* test(e2e): keep lifecycle sandboxes running Signed-off-by: Evan Lezar <elezar@nvidia.com> * test(e2e): cover sandbox terminal lifecycle semantics Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(e2e): synchronize detached main exit checks Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Evan Lezar <elezar@nvidia.com> Signed-off-by: Drew Newberry <anewberry@nvidia.com> Co-authored-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
857af42a16 |
feat(vm): support corporate HTTP forward proxy egress for microVM sandboxes (#3090)
* feat(vm): support corporate HTTP forward proxy egress for microVM sandboxes The corporate forward proxy machinery from #1792 is driver-agnostic and already merged: openshell-supervisor-network implements CONNECT chaining, NO_PROXY matching, credentials, https:// proxies and corporate CA trust, and openshell-sandbox exposes it as six argv-only flags. Podman gained the driver half in #2245/#2512 and Kubernetes in #2633; the VM driver had none of it, so VM sandboxes on proxy-only networks could not reach any destination requiring the proxy even when policy allowed it. The blocking piece was not proxy logic but delivery: the VM guest init script runs as PID 1 and execs a fixed supervisor command line, and libkrun's krun_set_exec receives an empty argv, so there was no channel for driver-owned supervisor arguments. The supervisor's proxy flags deliberately have no environment fallback, and build_guest_environment merges user-supplied environment, so the guest env is not a safe transport either. Add a driver-authored argument file, mirroring the existing init.d manifest: the driver writes /opt/openshell/supervisor-args into the overlay upperdir on every launch and the guest reads it verbatim, one argument per line, appending it to every supervisor exec. It is written even when empty, which is what makes the channel unforgeable -- the upperdir always shadows the read-only image layer, so an image can neither supply its own arguments nor disable the operator's by omitting the file. Because both launch backends exec the same init script, this covers libkrun and QEMU without touching either. A microVM has no bind mounts or container secrets, so the credential and CA bundle are staged into the per-sandbox overlay the way the gateway JWT already is: credential root-only at 0600, CA at 0644, both rewritten every launch so a removed setting clears prior material, and both deleted with the sandbox state directory. This places the credential at rest in the overlay image on the gateway host, which differs from the Podman secret model and is documented as an explicit security consideration. Validation is fail-closed and shared: a new openshell_core::driver_utils::validate_upstream_proxy_settings holds the pairing rules the Podman driver established, and both the gateway and the driver call it so an invalid table names the offending key instead of surfacing as an opaque driver-readiness timeout. Guest egress leaves through gvproxy, so a proxy on the gateway host's loopback is reachable only through host.openshell.internal; the guest to gateway callback is unaffected. Closes #3088 Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(vm): bound the proxy CA read and scope the host-loopback recipe Two review findings on the corporate forward proxy support for microVM sandboxes. The driver read the operator's proxy_ca_bundle with an unbounded fs::read and accepted it on a substring match for the PEM BEGIN CERTIFICATE marker. A special file such as /dev/zero therefore grew driver memory without bound on every authorized sandbox create, and a PEM block holding invalid DER passed the host check but contributes no trust anchor in the guest, so every supervisor would fail after boot with an error attributed to the sandbox rather than to the setting. Move the read into openshell-core as read_upstream_proxy_ca_bundle_file: it reuses the credential reader's bounded-read path (non-regular files rejected on fstat, size capped, read bounded even if the file grows), then requires at least one anchor that RootCertStore::add_parsable_certificates accepts. The supervisor's own reader now delegates to it, so host acceptance and guest acceptance are the same function and cannot drift. The published host-loopback recipe was written for libkrun only. gvproxy NATs host.openshell.internal to the gateway host's 127.0.0.1, but GPU sandboxes run on the QEMU/TAP backend where that name resolves to the TAP host address and the driver's own nftables input chain accepts only the gateway port from the guest — no proxy on the gateway host is reachable there at any bind address, so an operator following the generic recipe lost all proxy-required egress while configuration validation succeeded. Scope the recipe to libkrun in every reference and reject a gateway-host proxy URL when a launch plan resolves to QEMU, naming the reason, instead of booting a sandbox whose policy-approved CONNECTs all time out. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(vm): match the QEMU proxy preflight to the selected TAP host The gateway-host proxy guard added for the QEMU/TAP backend classified the wrong set of addresses in both directions. It ran at the top of configure_qemu_launch_plan, before the subnet allocation that settles plan.host_ip, so it could not compare against the address the guest actually reaches the host on. An operator pointing https_proxy at the sandbox's own TAP host address, such as 10.0.128.1, passed the check, and the driver's nftables input chain — which accepts only the gateway port from the guest — then dropped every policy-approved CONNECT, which is exactly the silent timeout the guard exists to prevent. In the other direction it rejected 192.168.127.254 unconditionally. That address is special only to libkrun/gvproxy; on QEMU/TAP it is an ordinary address that may be routable through the guest's masqueraded egress, so the guard refused a working configuration. Run the check after the launch plan's network allocation, on both the freshly-allocated and already-complete paths, and compare IP literals with that sandbox's selected TAP host. Loopback literals, localhost, and the documented host aliases that write_host_gateway_aliases seeds to the TAP host still classify as the gateway host, and the failure names the address. The gvproxy host-loopback constant returns to being a documentation anchor. Signed-off-by: Philippe Martin <phmartin@redhat.com> --------- Signed-off-by: Philippe Martin <phmartin@redhat.com> |
||
|
|
74960ebfae |
feat(server): add sandbox templates (#2833)
* feat(server): add sandbox workload templates Signed-off-by: Gordon Sim <gsim@redhat.com> * feat(go-sdk): add sandbox workload template support Signed-off-by: Gordon Sim <gsim@redhat.com> * feat(rust-sdk): add sandbox workload template support Signed-off-by: Gordon Sim <gsim@redhat.com> * feat(python-sdk): add sandbox workload template support Signed-off-by: Gordon Sim <gsim@redhat.com> * feat(typescript-sdk): add sandbox workload template support Signed-off-by: Gordon Sim <gsim@redhat.com> * docs(agents): document sandbox workload templates Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(cli): support default GPU requests in sandbox templates Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(server): cap sandbox templates per workspace Signed-off-by: Gordon Sim <gsim@redhat.com> * docs(architecture): document sandbox workload template boundaries Signed-off-by: Gordon Sim <gsim@redhat.com> * feat(cli+sdk): expose sandbox workload template provenance Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(cli): include sandbox template annotations in output Signed-off-by: Gordon Sim <gsim@redhat.com> * feat(sandbox): add label selectors to template listing Signed-off-by: Gordon Sim <gsim@redhat.com> * test(e2e): cover sandbox template failure paths Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(server): preserve command and ttl when creating sandbox from template Signed-off-by: Gordon Sim <gsim@redhat.com> * test(server): add field coverage test for template merge Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(go-sdk): add pagination support to fake client Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(cli): warn on env vars that looks like secrets Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(sdk-ts): propagate sandbox workspace through lifecycle calls Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(docs): update workspace management docs Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(ts-sdk): support command and tty when creating from template Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(go-sdk): support command and tty when creating from template Signed-off-by: Gordon Sim <gsim@redhat.com> * test(python-sdk): verify command and tty handling when creating from template Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(go-sdk): update docs and ClientInterface Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(server): validate sandbox create specs before I/O Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(go-sdk): guard empty DNS-1123 label validation Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(python-sdk): allow empty template builder mappings Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(cli): align template GPU JSON default output Signed-off-by: Gordon Sim <gsim@redhat.com> --------- Signed-off-by: Gordon Sim <gsim@redhat.com> |
||
|
|
9ca19e6c80 |
refactor(compute): decouple gateway driver composition (#2823)
* refactor(compute): decouple gateway driver composition Move first-party composition and VM process ownership into openshell-gateway, leaving openshell-server backend-independent. Update packaging and build references with the new crate, simplify the compiled-driver boundary, and keep the driver-free gateway path buildable with bundled Z3 tooling. Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(telemetry): bound compute driver categories Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(core): keep runtime transport generic Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(compute): complete server driver decoupling Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(compute): preserve driver integrations after rebase Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> * fix(compute): preserve docker tracing after decoupling Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> * fix(compute): preserve driver behavior after extraction Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(compute): remove MXC policy side channel Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(compute): separate policy delivery from readiness Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> |
||
|
|
07df822090 |
feat(providers): make profiles authoritative (#2962)
* feat(providers): make profiles authoritative Closes #1988 Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * docs(providers): move profiles into provider navigation Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * docs(providers): clarify provider attachment lifecycle Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * fix(tui): scroll provider profile picker Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * fix(providers): honor profile credential semantics Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * fix(providers): prefer exact profile IDs Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * fix(providers): harden authoritative profile adoption Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * test(oidc): align provider fixtures with profiles Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * fix(providers): preserve authoritative profile lifecycle Signed-off-by: John Myers <johntmyers@users.noreply.github.com> --------- Signed-off-by: John Myers <johntmyers@users.noreply.github.com> Co-authored-by: John Myers <johntmyers@users.noreply.github.com> |
||
|
|
a547dc9f4d |
fix(sandbox): reconcile early container exits (#3101)
* fix(sandbox): reconcile early container exits Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * test(oidc): isolate RBAC checks from lifecycle exits Signed-off-by: John Myers <johntmyers@users.noreply.github.com> --------- Signed-off-by: John Myers <johntmyers@users.noreply.github.com> Co-authored-by: John Myers <johntmyers@users.noreply.github.com> |
||
|
|
bb70461878 |
test(e2e): run conformance in gateway lanes (#2925)
* test(e2e): isolate VM-specific smoke assertions Signed-off-by: Evan Lezar <elezar@nvidia.com> * test(e2e): add portable CLI conformance baseline Signed-off-by: Evan Lezar <elezar@nvidia.com> * feat(conformance): add standalone CLI runner Signed-off-by: Evan Lezar <elezar@nvidia.com> * test(e2e): run conformance in gateway lanes Signed-off-by: Evan Lezar <elezar@nvidia.com> --------- Signed-off-by: Evan Lezar <elezar@nvidia.com> |
||
|
|
69a05ebb3b | fix(sandbox): complete successful main processes (#2884) | ||
|
|
981606d2f8 |
ci: build release binaries with Nix (#2977)
* ci: build release binaries with Nix Refs #1683 Signed-off-by: Simon Scatton <sscatton@nvidia.com> * ci: build VM artifacts with Nix Refs #1683 Signed-off-by: Simon Scatton <sscatton@nvidia.com> * ci: build images from Nix artifacts Refs #1683 Signed-off-by: Simon Scatton <sscatton@nvidia.com> * fix(nix): prevent host header leakage Signed-off-by: Simon Scatton <sscatton@nvidia.com> * ci: parallelize artifact builds Refs #1683 Signed-off-by: Simon Scatton <sscatton@nvidia.com> * ci: build external driver test artifacts Refs #1683 Signed-off-by: Simon Scatton <sscatton@nvidia.com> * fix(nix): disable mold in musl shells Signed-off-by: Simon Scatton <sscatton@nvidia.com> * ci: key Rust cache by Nix shell derivation Signed-off-by: Simon Scatton <sscatton@nvidia.com> * ci: refactor end-to-end workflows Signed-off-by: Simon Scatton <sscatton@nvidia.com> * ci: split platform binary workflows Signed-off-by: Simon Scatton <sscatton@nvidia.com> * ci: remove obsolete native build workflows Signed-off-by: Simon Scatton <sscatton@nvidia.com> * ci: replace disallowed mise action Signed-off-by: Simon Scatton <sscatton@nvidia.com> * ci: fix refactored e2e lanes Signed-off-by: Simon Scatton <sscatton@nvidia.com> * ci: check out local result action Signed-off-by: Simon Scatton <sscatton@nvidia.com> * ci: cache mise installations Signed-off-by: Simon Scatton <sscatton@nvidia.com> * ci: run docker builds on host runners Signed-off-by: Simon Scatton <sscatton@nvidia.com> * ci: disable unstable kubernetes e2e lanes Signed-off-by: Simon Scatton <sscatton@nvidia.com> * fix(ci): scope binary builds to cargo packages Signed-off-by: Simon Scatton <sscatton@nvidia.com> * fix(ci): address zizmor template injection findings Signed-off-by: Simon Scatton <sscatton@nvidia.com> * fix(ci): resolve remaining zizmor annotations Signed-off-by: Simon Scatton <sscatton@nvidia.com> --------- Signed-off-by: Simon Scatton <sscatton@nvidia.com> |
||
|
|
9f88f8ff9b |
ci: remove rootless podman e2e lane (#2981)
Signed-off-by: Evan Lezar <elezar@nvidia.com> |
||
|
|
18ce13b9b1 |
feat(providers): expose actionable OAuth refresh failures (#2887)
* fix(providers): classify OAuth refresh failures Signed-off-by: Mrunal Patel <mrunalp@gmail.com> * test(providers): add Keycloak refresh e2e lane Signed-off-by: Mrunal Patel <mrunalp@gmail.com> * fix(providers): harden OAuth refresh recovery Signed-off-by: Mrunal Patel <mrunalp@gmail.com> * fix(providers): classify post-mint refresh failures Signed-off-by: Mrunal Patel <mrunalp@gmail.com> * fix(providers): tolerate malformed OAuth subtypes Signed-off-by: Mrunal Patel <mrunalp@gmail.com> --------- Signed-off-by: Mrunal Patel <mrunalp@gmail.com> |
||
|
|
5206bc51b2 |
feat(sdk): add OAuth Client Credentials support to SDKs (#2907)
* feat(sdk): add renewable client credentials auth Implement lazy OAuth client-credentials acquisition and renewal for the Python, TypeScript, and Go SDK clients, with shared security conformance coverage and service-account documentation.\n\nCloses #2803 Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(sdk): address client credentials review Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(sdk): address client credentials review Signed-off-by: Seth Jennings <sjenning@redhat.com> --------- Signed-off-by: Seth Jennings <sjenning@redhat.com> |
||
|
|
72b9c4acef |
test(e2e): pin direct podman calls to harness socket on macOS (#2909)
The podman_gateway_start e2e test shells out to `podman` directly (unlike every other Podman e2e test, which routes through the gateway). The harness `e2e/with-podman-gateway.sh` clobbers `XDG_CONFIG_HOME` with an empty dir to isolate CLI/SDK gateway metadata. On macOS the podman client resolves its VM connection config from `XDG_CONFIG_HOME`, so a bare `podman ps` falls back to a nonexistent native rootless socket and fails with "unable to connect to Podman socket". On Linux the socket resolves via `XDG_RUNTIME_DIR`, so the test passes there and this is a macOS-only false failure. Target the same API socket the gateway uses by passing `--url unix://$SOCKET` when the harness-exported `OPENSHELL_PODMAN_SOCKET` is set, mirroring the shell's `podman_cmd` helper. When the var is unset (running the test outside the harness), fall back to plain `podman`, so Linux behavior is unchanged. Signed-off-by: Russell Bryant <rbryant@redhat.com> |
||
|
|
0a1f246587 |
feat(sandbox,podman): trust corporate CA for https:// proxies and intercepted TLS (#2512)
* feat(sandbox,podman): trust corporate CA for https:// proxies and intercepted TLS The corporate proxy chaining only accepted plain http:// proxy URLs, so operators whose forward proxy terminates TLS with a private corporate CA had no way to reach it, and TLS-intercepting proxies (mitmproxy, squid ssl-bump) that re-sign tunneled server certificates broke every upstream handshake after CONNECT. The supervisor now accepts https:// proxy URLs: it wraps the connection to the proxy in TLS before the CONNECT handshake, verifying the proxy certificate against the built-in Mozilla roots, the system CA bundle, and an optional operator corporate CA bundle. The upstream dial returns a Plain/Tls stream enum consumed generically by the relay paths. The corporate CA is delivered as a driver-supplied command-line argument (--upstream-proxy-ca-bundle), never an environment variable, matching the hardened proxy-config model where a sandbox image cannot influence the operator's egress boundary. It is folded into the sandbox combined trust bundle (write_ca_files) and the L7 upstream verification store (build_upstream_client_config) at startup, so intercepted upstream handshakes succeed and sandbox workloads trust the re-signed certificates. Configuration is fail-closed: a CA bundle set without a proxy, or an unreadable or certificate-free file, is fatal rather than silently weakening the trust boundary. The shared parse_upstream_proxy_url validator accepts https:// (recording the scheme so the driver and supervisor agree), keeping the explicit-port requirement. The Podman driver gains a proxy_ca_bundle operator setting (TOML, --sandbox-proxy-ca-bundle, OPENSHELL_SANDBOX_PROXY_CA_BUNDLE) that bind-mounts the host PEM read-only into the sandbox (a CA certificate is not secret) and points --upstream-proxy-ca-bundle at it, with a create-time readability check. The standalone dev gateway task passes OPENSHELL_SANDBOX_PROXY_CA_BUNDLE through to the generated podman config, so a local gateway can be pointed at a TLS-intercepting proxy without hand-editing the regenerated TOML. Refs #1792 Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(sandbox): reject CA bundles with valid PEM framing but invalid X.509 DER The proxy CA bundle validation counted PEM blocks that base64-decoded successfully, but did not verify the decoded bytes were accepted as trust anchors by RootCertStore. A bundle with syntactically valid PEM framing but invalid DER would pass the startup check while contributing zero usable anchors, causing opaque TLS failures at runtime instead of a fail-closed startup error. Validate decoded certificates through RootCertStore::add_parsable_certificates and reject the bundle unless at least one is accepted. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(sandbox,podman): allow proxy auth without insecure acknowledgement for https:// proxies For an https:// proxy the Proxy-Authorization credential travels inside the verified TLS session, so the proxy_auth_allow_insecure acknowledgement is unnecessary. Previously both http:// and https:// proxies required it, producing a misleading cleartext-risk diagnostic for a path that is already encrypted. Skip the requirement when the proxy URL uses https://; the acknowledgement is still tolerated if set. Updated in both the supervisor and Podman driver validation paths, with docs and tests. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix: format Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(kubernetes): use truly unsupported scheme in proxy validation test https:// is now a supported proxy scheme after a13c4dce, so the unsupported-scheme test must use a genuinely unsupported scheme. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(e2e): sign the https proxy fixture listener cert with a CA The corporate-proxy E2E fixture served a single `openssl req -x509` certificate as its TLS listener identity. OpenSSL marks that certificate `basicConstraints: critical, CA:TRUE`, and rustls refuses a CA certificate presented as an end-entity certificate (CaUsedAsEndEntity). The supervisor's TLS handshake with the proxy therefore failed, the upstream dial errored, and the workload's CONNECT was dropped without a response, so podman_corporate_proxy_trusts_ca_bundle_for_https_proxy failed on the approved destination while policy denial still worked. Generate a corporate CA and a separate listener leaf signed by it, serve the leaf chain, and publish only the CA as the bundle the supervisor trusts. This is what an intercepting proxy actually presents, and it exercises the corporate-CA trust path rather than pinning the listener certificate itself. Refs #1792 Signed-off-by: Philippe Martin <phmartin@redhat.com> --------- Signed-off-by: Philippe Martin <phmartin@redhat.com> |