298 Commits
Author SHA1 Message Date
Evan Lezar 2263685cf3 test(tmachine): migrate Keycloak provider refresh coverage (#3404)
* test(tmachine): add Keycloak provider refresh suite

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* refactor(tmachine): share container runtime detection

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* ci(tmachine): run feature suites in GitHub Actions

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* ci(tmachine): run conformance with Podman tests

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* ci(tmachine): cover provider refresh with Podman

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* ci(integration): split input preparation from runners

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-09-18 13:56:07 +02:00
John T. Myers 1d010f4187 feat(sandbox): validate configuration before workload activation (#3259)
* feat(sandbox): validate configuration before workload activation

Closes #3145

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(sandbox): bound startup failures and preserve activation history

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(sandbox): enforce deadlines on startup RPC attempts

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(sandbox): isolate provider auto-create policy fixture

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(sandbox): remove unnecessary fixture string delimiters

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(sandbox): capture startup logs before ephemeral cleanup

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(sandbox): preserve credential revocation after rebase

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(server): retain provider revision helper for endpoint reports

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(server): import provider object trait in production

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(supervisor): preserve admission across boundary extraction

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(sandbox): use existing boundary discovery imports

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(sandbox): distinguish workload and supervisor containers

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(server): reconcile admission schema with timestamp migration

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(server): reconcile admission with typed deletion schema

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(server): reconcile admission with mutation request IDs

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(supervisor): adapt local startup fixture to admission state

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(supervisor): deduplicate startup quarantine diagnostics

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(tui): show configuration blockers in sandbox notes

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* feat(sandbox): expire provisioning repair attempts after five minutes

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(supervisor): reconcile admission with provider readiness

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(tui): separate configuration summaries from full diagnostics

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(tui): shorten invalid configuration note

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(server): refresh schema inventory after rebase

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-09-18 00:16:18 +00:00
Philippe MartinandJohn Myers 04146692d9 refactor(providers)!: make provider profiles import-only (#3383)
* chore(providers): remove dead provider plugin modules

Twelve modules under crates/openshell-providers/src/providers/ were never
declared in providers/mod.rs, so they have not been compiled since the plugin
registry was narrowed to the two adapters it still registers. Four of them
(generic, gitlab, opencode, outlook) key off provider type identifiers that
normalize_provider_type already retired.

Keep google_cloud and vertex, which are the only plugins ProviderRegistry::new
registers.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* test(providers): load example profiles from providers/ at test time

Add an example_profiles module that reads the YAML under providers/ from the
source checkout at run time and parses it with the existing profile loader. It
is gated behind a new non-default example-profiles feature so it is available to
this crate's own tests and, once wired into dev-dependencies, to gateway and CLI
tests, while never reaching a release binary.

Nothing consumes it yet; later changes move the test fixtures off the compiled
catalog and onto these files.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* test: source profile fixtures from providers/ instead of the compiled catalog

Unit and integration fixtures reached into builtin_profiles() to get a profile
to work with, which ties the tests to the compiled catalog rather than to the
files an operator would import. Point them at the example_profiles loader
instead.

The profiles.rs tests keep their coverage unchanged and become explicit golden
tests over providers/*.yaml: they are what keeps those files valid once nothing
compiles them. builtin_profiles_are_sorted_by_id widens into a test that the
whole example set parses, sorts, and lints clean as one catalog.

The CLI fake gateways serve the example profiles the way a real gateway serves
what an operator imported, through a shared helper.

No behavior change: the gateway still loads the built-in source by default and
serves the same profiles.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* docs(providers): document the example provider profiles

Each file under providers/ now opens with a header naming its expected client
binary identities, the image layout those paths assume, the credential scope,
the endpoint access it grants, and a smoke test. Add a README covering the
import commands and why a profile should be copied and edited rather than
imported unchanged.

Several of these profiles bind network access to paths that only exist in the
OpenShell Community image — /sandbox/.venv, /app/.venv, /sandbox/.cursor-server,
/usr/lib/node_modules. Imported unchanged into another image the profile matches
nothing: the catalog still advertises it, but the credential is never injected
and the traffic is denied. The headers say so where it applies.

Comments only; the profile schema has no documentation fields and all fifteen
files still lint clean.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* test(server): seed unit-test state with the example provider profiles

Server unit tests inherited the builtin + user source default from ServerState,
so around a hundred and forty assertions about github, openai and the rest
resolved against the compiled catalog. Point the test state at the user source
alone and import the example profiles from providers/ into its store first, the
way an operator would. Every one of those assertions keeps passing unchanged,
which is the point: it proves the gateway behaves identically with an imported
catalog before the default moves.

Four tests asserted builtin-source semantics specifically. Under import-only the
only profiles a gateway cannot edit are the ones a non-user source vends, so
they now exercise a source-managed profile composed with the user source; the
read-only guard they cover is the one that still applies to interceptor
catalogs. The list test asserts the imported profile's user/platform identity
instead of builtin with an empty scope.

test_server_state_with_user_only_github_profile is gone: the default test state
now is a user-only gateway with github imported.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* refactor(cli): resolve credential suggestions from the gateway catalog

The --env credential warning scanned a profile table compiled into the CLI, so
its suggestions described the binary rather than the gateway the user is talking
to: a profile the gateway does not serve was suggested anyway, and an imported
custom profile never was.

Fetch the catalog from the connected gateway instead. The warning moves out of
argument parsing and into sandbox create and sandbox template create, where a
client already exists. A catalog fetch failure is not fatal — the warning
degrades to its generic form rather than blocking sandbox creation.

Extract the ListProviderProfiles paging loop from provider list-profiles into a
shared fetch_provider_profile_catalog, which the profile-driven paths now share.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* refactor(cli)!: infer providers from the gateway catalog

Command-to-provider inference went through a hardcoded alias table in
openshell-providers, a second copy of the built-in catalog's identifiers. A
custom profile could never be inferred no matter what binaries it declared, and
the table drifted from the profiles it mirrored.

Infer from the connected gateway's catalog instead: match the command's basename
against each profile's ID and against the basenames of the binaries the profile
authorizes. A profile that names /usr/bin/claude is the profile for running
claude. The match must be unique — where several profiles claim a command, the
user names one with --provider — and an empty catalog infers nothing.

No command is special-cased. `binaries` is the operator's authorization
statement, so a profile that declares a binary claims the command that runs it,
whatever that binary is; narrowing that belongs in the profile rather than in a
list compiled into the CLI, which could never cover an unbounded catalog anyway.

Breaking: the retired aliases stop resolving, and commands the old table never
listed can now infer. git, pip and uv are declared by the github and pypi
example profiles, so they infer where they previously did not.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* refactor(providers,server)!: resolve provider profiles by exact ID

normalize_provider_type was a hardcoded alias map — gh to github, claude to
claude-code, vertex to google-vertex-ai — and a second copy of the built-in
catalog's identifiers, independent of the YAML it mirrored. It let a provider
type resolve to a profile the operator never named, and it made the built-in IDs
behave as a reserved namespace.

Remove it, along with detect_provider_from_command and the alias-normalizing
ProviderRegistry::inject_env. A provider type now names a profile exactly:

- the effective catalog resolves an ID or reports it absent, with no alias retry
- plugins activate only for the ID of a profile the gateway resolved; a provider
  with no resolvable profile gets no plugin projection, instead of falling back
  to an alias guess
- the CLI surfaces the gateway's not-found instead of retrying under an alias
- telemetry buckets by profile ID, and the gitlab, opencode and outlook buckets
  go with the aliases that were their only source

Breaking: `--type gh`, `--type claude` and the other aliases no longer resolve.
Use the profile's own ID.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* feat(server): fail closed when a provider's profile is absent

A sandbox composed from a provider whose profile the gateway cannot resolve
started anyway: the credential and policy builders warned and skipped, so the
sandbox came up carrying none of that provider's credentials or network policy.
The operator learned about it later, as a denied connection or a missing
environment variable, rather than as the configuration error it is.

Check at the two composition boundaries — CreateSandbox and
AttachSandboxProvider — right beside the catalog snapshot already taken there,
and reject with a bounded diagnostic naming the provider, the profile it refers
to, and the import command that supplies it. Scope-aware: a platform-scoped
provider is told to import with --global.

Read paths are untouched. ListProviders, GetProvider and profile export keep
working so an operator can see and recover an affected provider, and the shared
policy builders keep their warn-and-skip for the diagnostic paths that also
reach them.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* refactor(server): scope the vendor base-URL pin to declared endpoints

provider_profile_endpoints_are_active withheld a credential and the provider's
policy layer when an openai or anthropic provider pointed its client somewhere
other than the public vendor endpoint. The guard was keyed on
profile.source == "builtin" and on those two profile IDs, so it protected only
profiles OpenShell shipped. Once profiles are import-only no profile is ever
builtin, and the guard would silently stop applying — including to an operator
who imported providers/openai.yaml verbatim.

Key it on the profile instead of on where the profile came from. A profile's
endpoints are the boundary its credential is bound to, so if the provider
configures a *_BASE_URL pointing at a host the profile does not declare, the
profile no longer describes where that credential goes and is treated as
endpointless. Host matching reuses the DNS-label-aware matcher in
openshell-core, so wildcard endpoints such as Vertex's
*-aiplatform.googleapis.com resolve correctly.

A profile with no declared endpoints has no boundary to contradict, and a config
value that names no host is not a redirect this can reason about. Both keep the
profile active. The control now covers every endpoint-bearing profile, including
an operator's own, and the last "builtin" string leaves the gateway.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* test(e2e): import example provider profiles during gateway bring-up

The e2e suites create providers from github, openai, nvidia, claude-code and
google-cloud, and the GitHub lane runs a real git clone through the profile's
binary attribution. A gateway serves only the profiles an operator imported, so
the lanes have to import them.

Add e2e_import_example_provider_profiles to the shared bring-up helpers and call
it from the Docker, Podman, Kubernetes and VM wrappers once the gateway is
healthy and registered. It runs the same command the upgrade notes give
operators, against the repository's own providers/ directory, so the lanes
exercise the documented path rather than a test-only shortcut.

Lands before the default changes: a same-ID user profile already shadows a
built-in, so importing works today.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* refactor(providers)!: make provider profiles import-only

OpenShell compiled fifteen provider profile YAML files into every release binary
and selected them by default, so a fresh gateway published a catalog it was
never configured with. Those profiles are not image-neutral: their binary
selectors name paths that exist in the OpenShell Community image, so changing
the sandbox image could make a profile inert while the catalog still advertised
it. Every endpoint, credential name and binary path in providers/ was also
effectively part of the 0.1.0 public contract.

A gateway's catalog is now exactly what an operator imported:

- providers/*.yaml is no longer include_str!'d, and builtin_profiles() is gone
  from the openshell-providers API
- the default provider_profile_sources is [{ type = "user" }], and a gateway
  with nothing imported reaches ready and serves an empty catalog — an empty
  catalog is a valid state, not a startup failure
- the builtin source type is removed from the configuration schema, and a
  gateway.toml that still names it is rejected at parse time with the import
  command rather than an unknown-variant error
- nothing reserves the canonical identifiers any more, so github, pypi,
  anthropic and the rest import at their own IDs and the imported profile is the
  only definition for that ID; the static_fallback that kept a shipped
  definition resident behind an imported one is gone with the source that
  produced it

A collision between an interceptor-vended profile and an imported one still
fails closed, which is the behavior that shadowing quietly bypassed for
built-ins.

Operators upgrading should export the profiles their deployment relies on
before upgrading, or copy them from the providers/ directory of the matching
release tag, then import them at the scope their providers use.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* docs(providers): document import-only provider profiles

Rewrite the provider profile documentation around a catalog the operator builds
rather than one the gateway ships.

- gateway-config: the default is [{ type = "user" }], the builtin source type is
  gone and rejected at startup, and an empty catalog is a valid ready state
- profiles: replace the built-in profile table and its shadowing semantics with
  the import workflow, and say plainly that a profile whose binary paths do not
  match the image is inert
- manage-providers: replace the two fixed provider-type tables with
  `provider list-profiles`, and describe catalog-driven command inference
- inference-routing: drop the "still loads built-in profiles for compatibility"
  transition text; its migration walkthrough is now the normal path
- quickstart and the Docker Compose, GitHub, AWS, Google Cloud and Vertex
  tutorials: import the profile before creating the provider, since nothing
  resolves without it
- release notes: a 0.1.0 migration section covering export-before-upgrade,
  importing from the release tag's providers/ directory, what happens to a
  provider whose profile is missing, and the removal of the legacy type aliases
- README, architecture, the openshell-cli and debug-openshell-cluster skills,
  and the governance interceptor example follow the same change; the example's
  smoke assertion now imports a profile to prove an authoritative interceptor
  hides it

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(cli,e2e): distinguish catalog failures and authenticate OIDC seeding

Addresses two findings from review (GATOR-c1bd9867-02 and -03).

The provider profile catalog lookup collapsed its error into an empty
catalog, so an authorization, availability or transport failure on
ListProviderProfiles was indistinguishable from a gateway that genuinely has
no profiles. Command inference then resolved nothing and the sandbox was
created without the provider it needed, deferring the failure to the workload.

Keep the lookup's outcome instead of discarding it. The credential warning is
advisory and still degrades to its generic form, but inference now consults the
catalog only when there is a command to resolve and surfaces the lookup failure
when there is, naming the gateway and pointing at explicit --provider selection.
Two regression tests cover it through a fake gateway whose ListProviderProfiles
returns UNAVAILABLE while every other RPC succeeds: sandbox creation fails
without sending a provider-less create request, and a sandbox with no trailing
command still succeeds because it needs no catalog.

The e2e profile seeding also ran unauthenticated in the OIDC lanes. Those lanes
deliberately skip gateway registration and start the gateway without a TLS
client CA, so no mTLS identity exists and no token has been acquired when the
import runs; the wrapper exited during setup. Skipping the import is not
sufficient because the provider tests now require the claude-code profile.

Add e2e_register_oidc_admin_session, which mints an administrator token with
Keycloak's password grant — the same grant the OIDC test helpers use — and
writes the gateway metadata and token bundle that an interactive login would
have stored, so the import runs as an authenticated administrator. The mTLS
lanes keep the direct import unchanged. Both affected wrappers are covered:
with-podman-gateway.sh had the same defect as with-docker-gateway.sh.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* refactor(cli)!: remove command-derived provider attachment

Addresses GATOR-c1bd9867-01.

A profile's `binaries` list authorizes a binary to reach that profile's
endpoints. It is not a statement that running the binary asks for the provider,
and reading it as attachment intent let a command silently gain provider
authority: the aws-s3 example declares /bin/bash, so `sandbox create -- bash`
resolved an existing aws-s3 provider and attached its credential-backed
capability and network policy to the shell and its descendants. Attachment of
an already-created provider never prompted, so the confirmation flow did not
guard it.

The distinction that would make inference sound — whether a declared binary is
a profile's client or merely a permitted runtime — cannot be expressed:
NetworkBinary carries only a path, and the field that encoded it was removed in
0.1.0. Any substitute is a guess. Restricting the guess to a unique claimant
does not help, because uniqueness measures how sparse the catalog is rather
than what the user intended, and a compiled list of "generic" commands could
never cover an unbounded operator catalog.

Remove trailing-command inference rather than approximate it. A provider is
attached only when named with --provider, which still creates a missing
provider from local discovery when the name matches an imported profile ID.
Sandboxes with no providers remain a normal, fully supported state.

Removing inference also settles GATOR-c1bd9867-02: with no consumer deriving
authority from the catalog, its only remaining use is the advisory credential
warning, so a failed lookup degrades that warning instead of blocking creation.
The regression test now asserts that an unreachable catalog still creates the
sandbox and attaches nothing.

While repurposing the deduplication test, a pre-existing defect surfaced:
repeating a name in --provider auto-created it twice and the second attempt
failed with "provider already exists", because the explicit pass never
consulted the set of names it had already handled. Guard it.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(e2e): load the OIDC token and trust the gateway when seeding profiles

Addresses the carried finding GATOR-c1bd9867-03.

e2e_register_oidc_admin_session established a session the CLI never used. The
CLI decides whether to load a stored bearer token by matching on the auth_mode
field of the gateway metadata alone; the helper omitted that field, so the
metadata fell through to the default arm and oidc_token.json stayed on disk
unread. The profile import went out unauthenticated exactly as it had before
the helper existed. The helper also installed no trust anchor, leaving the CLI
unable to verify the gateway's self-signed serving certificate.

Write auth_mode = "oidc" so the stored token is loaded, and install the CA at
<gateway>/mtls/ca.crt. Only the CA is installed: with no client certificate or
key on disk the CLI falls back to CA-only server verification and authenticates
with the bearer token, which is what these lanes need because they start the
gateway without --tls-client-ca. Certificate verification stays on; no insecure
transport override is introduced.

Assert the session before anything depends on it. ListProviderProfiles is
annotated auth_mode: "bearer", so it cannot succeed unless the token was loaded
and accepted. The helper now fails at that point, naming the gateway config
directory and echoing the CLI output, rather than letting the defect surface
later as an opaque profile import error.

Both affected wrappers pass the PKI directory and CLI binary the helper needs;
the mTLS lanes are untouched.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(e2e): request the OpenShell scopes when minting the admin token

The OIDC lanes failed at the session assertion added in b25dd0298 with
PERMISSION_DENIED and "scope 'provider:read' required". Authentication was
working -- the gateway logged ListProviderProfiles at gRPC status 7, which it
can only reach once the bearer token has been loaded and accepted. The token
simply carried no OpenShell scope.

sandbox:*, provider:*, config:*, workspace:* and openshell:all are optional
client scopes on the openshell-cli client in scripts/keycloak-realm.json, so
Keycloak mints them only when the request asks for them. The password grant
here asked for nothing, leaving the realm defaults (openid, profile, email,
roles, web-origins, acr) and an access token that authorizes no RPC.

Request "openid openshell:all", as e2e/python/oidc/oidc_auth_test.py already
does for its administrator tokens. openshell:all is SCOPE_ALL in
crates/openshell-server/src/auth/authz.rs, so one scope covers the setup calls
without enumerating them. Verified against the realm: the token goes from
"email profile openid" to "email profile openshell:all openid".

Record the same scopes in metadata.json. The helper writes the bundle
openshell gateway login would have stored, and oidc_scopes is the field that
login path reads back, so leaving it out would misdescribe the stored token.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(e2e): seed provider profiles once per VM lane gateway

The VM lanes failed on the second test target with "custom provider profile
'anthropic' already exists" for all fifteen example profiles, and the import
exited non-zero.

e2e_import_example_provider_profiles was called from run_e2e_test, so it ran
once per target -- four times against one long-lived gateway. That was
harmless while import overwrote silently, but profiles are now import-only:
ImportProviderProfiles is create-only and reports an existing id as an
error-severity diagnostic, with no overwrite flag on the request. The first
import therefore succeeds and every later one fails.

Hoist the call to just after the conformance run, which is where the docker,
podman and kube lanes already seed their catalogs. The profiles persist for
the gateway's lifetime, so every target still finds them, and the
E2E_TEST_OVERRIDE path is covered by the same single call.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* test(e2e): name the provider explicitly in the OIDC workspace-user test

user_can_create_sandbox_with_inferred_provider_command reached the gateway
once the admin session was fixed, and then failed: it asserted a
missing-provider error, but the sandbox was created and died provisioning
with "failed to spawn sandbox entrypoint process 'claude-code'".

The test drove provider resolution by passing claude-code as the trailing
command and relying on the CLI to infer the provider type from it. That
inference is what this branch removed, so the trailing word is now nothing
but an entrypoint, and the image has no such binary.

The regression the test guards is not inference itself: it is that a
workspace user resolving a provider is not gated behind Platform Admin.
Name the provider with --provider, the only remaining way to attach one.
The CLI still has to fetch the claude-code profile before it can auto-create
the provider, so the lookup a workspace user must be allowed to make still
happens, and auto-creation still stops at the non-interactive branch with
"missing required provider". Both assertions therefore keep their meaning.

Rename the test and rework its comments to describe what it now exercises;
the old name would otherwise outlive the behavior it was named for.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(providers): reconcile import-only profiles with main

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: Philippe Martin <phmartin@redhat.com>
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-09-17 22:59:27 +00:00
John T. Myers b7a932a4d0 docs(inference): remove stale managed endpoint references (#3428)
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-09-17 22:54:20 +00:00
Shiju a72351d370 fix(policy)!: require explicit L7 append targets and scope (#3380)
Make allow and deny appends identify a rule and endpoint and declare every
affected binary and port. Reject incomplete, stale, ambiguous, and provider
targets atomically so a small append cannot silently change a broader scope.

Update CLI previews, wire requests, Go types and generated bindings, SDK
regressions, operator documentation, and live policy-update coverage.

BREAKING CHANGE: AddAllowRules and AddDenyRules require L7RuleTarget instead
of host and port. CLI appends require a rule name and explicit binary scope.

Signed-off-by: Shiju <shiju@nvidia.com>
2026-09-17 22:53:07 +00:00
Piotr MlocekandDrew Newberry cb93f62bfe fix(e2e): support distroless supervisor fixture (#3431)
* fix(e2e): support distroless supervisor fixture

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* test(e2e): preserve provider fixture trust bundle

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
2026-09-17 14:34:14 -07:00
Shiju 9708ba9999 feat(providers): report applied sandbox provider changes (#3391)
* feat(providers): report applied sandbox provider changes

Record exact provider mutation targets in shared configuration operations.
Require authenticated evidence that credentials, effective policy, and the
workload launch environment have been installed before reporting readiness.

Add bounded CLI and Rust SDK status and wait support, preserving ordinary
revision-scoped references for existing processes. Verify new-client rotation
and acknowledged detach revocation without external-stable resolver changes.

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(providers): align readiness times with protobuf contracts

Represent readiness receipts, status, and operation times with Timestamp
and report intervals with Duration. Reserve the scalar field tags, update
all consumers and generated bindings, and preserve timestamp presence
and nanosecond identity through storage and client validation.

Qualify both empty-map constructors in the Linux boundary test so its
module compiles while retaining the explicit default required by Clippy.

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(cli): preserve provider mutation storage uncertainty

Recognize the gateway's exact structured storage-uncertainty reason for
provider attach, detach, and update. Explain that the change may already
be saved and must be reconciled before retrying, without exposing server
messages or metadata. Preserve uncertainty ahead of generic retry hints.

Exercise saved mutations through the CLI and verify single submission,
redaction, missing receipt handling, and untrusted error-detail rejection.
Document the recovery guidance for users and the public CLI skill.

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(cli): explain denied provider profile lookups

Report exact and alias profile lookup denials with fixed permission and workspace guidance. Keep backend details redacted and stop before provider mutations.

Cover denied create and update calls through the CLI. Verify the complete provider list independently in the cross-workspace OIDC regression, extracting its JSON object from surrounding startup diagnostics.

Signed-off-by: Shiju <shiju@nvidia.com>

---------

Signed-off-by: Shiju <shiju@nvidia.com>
2026-09-17 19:10:43 +00:00
Drew NewberryandEvan Lezar 8de26878f9 fix(supervisor): serialize child registration with reaping (#3142)
* test(supervisor): reproduce fast child reaping race

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(supervisor): serialize child registration with reaping

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(supervisor): cover child registration reaper race

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Evan Lezar <elezar@nvidia.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Evan Lezar <elezar@nvidia.com>
2026-09-17 16:59:32 +00:00
krishicks 3dd5ce3314 test(parity): align fixture supervisor provenance (#3387)
The deterministic fixture still modeled Alpine after the supervisor image
switched to Debian, causing the wrapper to exit before producing parity
evidence. Update its image expectations and launch evidence to match the
current harness contract.

Force the intentional replacement of the read-only staged artifact so BSD mv
does not prompt when the deterministic parity test runs from a terminal.

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-09-17 15:55:58 +00:00
Mrunal Patel a316fd7832 feat(api): add durable workspace mutation admission and replay (#3321)
* feat(api): add durable workspace mutation admission and replay

Part of #3051 (phase 3a). Preserve unresolved admissions, reauthorize replay, and protect same-name replacements.

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* feat(api): extend mutation replay through gateway interceptors (#3323)

Add typed durable replay receipts for 24 ordinary unary mutations, protect sensitive payload fingerprints, and revalidate intercepted retries without repeating post-commit observation.

Part of #3051 (phase 3b).

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(api): scope workspace request IDs by target

Include the requested workspace name in create/delete admission keys while leaving workspace UUID guards unset. Cover cross-target UUID reuse, replay, and missing targets with server and live gateway regressions.

Merge the latest phase-two SDK fixes and preserve the approved interceptor replay changes.

Refs #3051.

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

---------

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
2026-09-17 14:51:56 +00:00
Carlos Villela 7e7a8d5610 fix(sandbox): pass declared environment to the initial process (#3392)
Apply the workload environment before stripping supervisor-only values and injecting provider placeholders in both canonical launch paths. The RFC 0012 bootstrap carries these values separately from inherited supervisor environment.

Add child-process coverage for declared values, explicit HOME, stripped identity material, and provider precedence. Add a Docker-backed lifecycle regression for initial and exec environments in TTY and non-TTY modes.

Validation: the child-process test fails without env application; the Docker regression fails with the pinned unpatched runtime and passes with the rebuilt runtime. Sandbox library tests: 201 passed, one ignored on Rust 1.95.0. mise run pre-commit passed. Full mise run ci is blocked by the host linker missing libz3. Fixes #3377.

Signed-off-by: Carlos Villela <cvillela@nvidia.com>
2026-09-17 05:46:06 +00:00
John T. Myers 0fd3385726 fix(container): use distroless Debian 13 for supervisor (#3393)
* fix(container): use distroless Debian 13 for supervisor

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(docker): preserve supervisor bootstrap ownership

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-09-17 05:29:47 +00:00
Mrunal Patel c502be9fd7 feat(api): return typed deletion outcomes with explicit missing-target semantics (#3317)
* feat(api): return typed deletion outcomes with explicit missing-target semantics

Implement phase 2 of #3051 across the public gateway API and first-party SDKs. Preserve asynchronous sandbox deletion and observed resource identities, reserve legacy wire fields, and document the coordinated migration.

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(sdk): bind deletion waits to sandbox identity

Track accepted deletions by original sandbox identity in Rust and TypeScript. Preserve name-only waits and cover replacement races and lookup failures.

Merge the phase-one cleanup fix and adapt its interceptor regressions to typed deletion outcomes.

Refs #3051.

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* test(e2e): accept asynchronous stopped sandbox deletion

Allow either completed or accepted deletion output, then continue polling for actual sandbox absence. This preserves the lifecycle assertion under the phase-two deletion contract.

Refs #3051.

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

---------

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
2026-09-16 22:11:19 +00:00
Derek Carr d68b7069c3 refactor(proto)!: use well-known time types (#3113)
* refactor(proto)!: use well-known time types

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(proto): preserve time migration behavior

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(proto): preserve timestamp boundary semantics

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(proto): convert sandbox token expiry to timestamp

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(proto): preserve time compatibility semantics

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(proto): preserve exact endpoint and profile times

Signed-off-by: Derek Carr <decarr@redhat.com>

* test(e2e): use duration for interactive exec timeout

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(sdk-go)!: remove legacy profile duration fields

BREAKING CHANGE: Go provider profile callers must use RefreshBefore, MaxLifetime, and CacheTTL with ProfileDuration instead of the whole-second fields.

Signed-off-by: Derek Carr <decarr@redhat.com>

---------

Signed-off-by: Derek Carr <decarr@redhat.com>
2026-09-16 19:51:45 +00:00
Johnny Greco 2ccef97769 feat(policy): establish one canonical authored policy representation (#3334)
* chore(policy): restart schema implementation

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* feat(policy-schema): add canonical authored policy model

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* refactor(policy): use canonical authored schema

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* refactor(prover): project canonical policy documents

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(policy): describe shared schema boundary

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(policy): close reviewed parser gaps

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(policy-schema): fail closed on unsupported fields

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(policy): preserve partial process identities

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(policy-schema): harden authored policy inspection

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* refactor(policy): rename policy schema crate

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* revert(policy): restore policy schema crate

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* test(e2e): serialize OIDC PKCE scenarios

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

---------

Signed-off-by: Johnny Greco <jogreco@nvidia.com>
2026-09-16 17:36:31 +00:00
Drew Newberry b3e4ad4579 fix(ci): repair RFC 0012 post-merge checks (#3360)
* fix(ci): lock nextest for Windows ARM64

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): preserve GPU access in sandbox runtime

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(windows): restore host proxy identity binding

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(windows): satisfy cross-platform network lint

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(windows): satisfy MXC lint

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(windows): exclude Unix supervisor runtime

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(mxc): satisfy Windows test lint

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(windows): escape proxy test policy paths

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): accept admitted GPU runtime groups

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(docker): serialize proxy pipeline scenarios

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-16 04:08:00 +00:00
Drew Newberry c1f2e7189f feat(isolation): implement the RFC 0012 sandbox architecture (#2942)
* feat(isolation): add RFC 0012 backend contract

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>

* refactor(isolation): name the interface crate explicitly

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): expose trusted host gateway

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(agents): inventory the MXC driver

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): add mediated DNS transport

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): tighten interface error and digest contracts

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(isolation): remove unrelated driver inventory

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): define capability-free launch contract

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): seal confirmed boundary state

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): validate confirmation for external backend implementations

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): clarify mediated DNS identity

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(isolation): generalize loopback connector

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(isolation): unify typed network mediation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): bind launches to sandbox sessions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(mxc): initialize extended sandbox status

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): add boundary protocol and Linux primitives

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): harden signals and separate process status from transport

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): validate remote confirmation through public contract

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): validate wire state and propagate snapshot failures

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(isolation): import owned agent specification explicitly

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(isolation): describe mediated DNS channel

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): bound mediation attach without nested retries

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(isolation): generalize loopback protocol

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): add transport-neutral session authentication

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(isolation): separate sandbox backend protocol

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): harden runtime boundary controls

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): add terminal boundary operation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): split supervisor and sandbox runtimes

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): harden boundary isolation and lifecycle ownership

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): reject private root redirects and adopt typed errors

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): preserve accept thread ownership on musl

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(sandbox): isolate credential probes from filtered threads

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): return retained exec exit status to independent waiters

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): bound network mediation and preserve socket authorization

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): bound control admission and retire stale mediation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci(e2e): select migrated drivers per stack layer

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(sandbox): implement loopback connector

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): authenticate the Sandbox Protocol

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(supervisor): rotate launch-scoped authentication

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(sandbox): consume dedicated backend crate

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(sandbox): align topology session fixture

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): align projected bootstrap bundle

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(auth): validate refreshed credentials before rotation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): fail closed across supervisor disconnects

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): repair rebased sandbox CI

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* build(runtime): publish separate sandbox and supervisor images

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(config): configure the sandbox runtime image

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(ci): validate sandbox binary linkage

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(isolation): use backend and runtime terminology

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(sandbox): use a scratch runtime image

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(ci): refresh schema and dependency policy

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): bind reconnects to supervisor process

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs: align runtime split operational guidance

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* chore(security): document Kubernetes runtime RBAC

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): enforce runtime lifecycle invariants

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(compute): identify sandbox start generations

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(server): restore sandbox launch sessions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): support authenticated runtime replacement

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(auth): bind sandbox session successors

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(auth): retry pending sandbox successors

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(vm): run the supervisor outside the guest workload

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(vm): use sandbox backend protocol

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): repair rebase integration

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): use unified build toolchain

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(vm): use sandbox runtime terminology

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(vm): own guest network bootstrap

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): expose guest init version

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): select native supervisor artifacts

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): guard guest init Linux symbols

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): scope Linux test imports

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): avoid guest interface casts

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): reconcile admitted sandbox identity

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): share resolved sandbox identity

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): surface host supervisor failures

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): include guest logs on supervisor exit

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): rotate and clean runtime generations

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): make sandbox starts generation-aware

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): rotate restored sandbox sessions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(vm): keep shared paths in the base layer

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): bind sandbox session lineage

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(docker): isolate workloads behind the host supervisor

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(docker): rotate launch-scoped authentication

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(docker): use sandbox backend protocol

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(docker): use host networking for supervisor

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): preserve host gateway alias resolution

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(docker): use separate sandbox and supervisor images

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): restore startup validation after rebase

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(docker): name the sandbox runtime directly

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): narrow supervisor CA runtime storage

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): close companion isolation gaps

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(docker): align mediated network expectations

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(docker): exercise mediated network paths

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): attach supervisor to managed network

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): defer supervisor recovery until gateway is ready

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): make sandbox starts generation-aware

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): rotate restored sandbox sessions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): preserve workloads during session rotation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(docker): remove unrelated configuration RFC changes

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): bind sandbox session lineage

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(kubernetes): add proxy-pod isolation topology

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(kubernetes): use sandbox backend protocol

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): use stable sandbox service authority

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(kubernetes): split sandbox and supervisor images

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): adapt proxy pods to current runtime APIs

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(kubernetes): describe the single runtime placement

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(kubernetes): simplify sandbox orchestration

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): validate deployment prerequisites

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): update Trivy Helm profile inventory

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(kubernetes): update Trivy scan inventory count

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): reuse preloaded runtime images in e2e

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): type and clean runtime resources

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): make sandbox restarts recoverable

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): rotate restored sandbox sessions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): preserve supervisor egress

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): bind sandbox session lineage

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(podman): adopt isolated sandbox and supervisor containers

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): stage bootstrap archives at named volume destinations

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(podman): rotate launch-scoped authentication

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(podman): use sandbox backend protocol

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(podman): use host networking for supervisor

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(podman): split sandbox and supervisor images

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): repair rebase integration

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(podman): name the sandbox runtime directly

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): provision supervisor CA runtime storage

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): address isolation review findings

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): inspect Debian supervisor provenance

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): use libpod-compatible tmpfs options

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): bind verified sandbox runtime binary

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): provide external driver data directory

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): start sandbox before joining user namespace

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): separate supervisor user namespace

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): make sandbox starts generation-aware

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): rotate restored sandbox sessions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): bind sandbox session lineage

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* perf(isolation): add TCP and DNS benchmark harnesses

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(perf): align benchmark timing and supported protocols

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(perf): report TCP benchmark metrics accurately

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(perf): cancel failed worker startup

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): build matching local supervisor image

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): make local sandbox smoke test runnable

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): wire local sandbox runtime image

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): narrow sandbox service RBAC

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(ci): validate split runtime artifacts

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): harden runtime session handling

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(supervisor): add standalone network proxy role

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(rfc): remove implementation companion notes

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(vm): standardize runtime release name

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): pin renamed runtime artifacts

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(auth): persist sandbox runtime identity

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(runtime): restore branch validation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): reconcile main after rebase

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(network): close unframed HTTP 1.0 responses

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* chore(isolation): preserve upstream OCSF updates

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(security): close credential and TLS replay paths

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(auth): make sandbox refresh retries idempotent

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-16 00:49:14 +00:00
Seth Jennings dbe36eaf85 fix(security): harden Vault credential transport (#3329)
Reject non-loopback plaintext Vault endpoints, disable redirects, and support private CA bundles without weakening hostname verification. Update Helm configuration, documentation, operator skills, and regression coverage for OSSR-002.

Signed-off-by: Seth Jennings <sjenning@redhat.com>
2026-09-15 19:56:09 +00:00
Mrunal Patel b799fccb8b fix(auth): harden OIDC trust root retrieval (#3332)
* fix(auth): harden OIDC trust root retrieval

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(e2e): pass OIDC HTTP acknowledgement value

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

---------

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
2026-09-15 17:00:45 +00:00
Jorge 42e9bcf2b0 feat(e2e): support the Vault credential-driver lane on OpenShift (#3312)
Implements https://github.com/NVIDIA/OpenShell/issues/3212

Running e2e:kubernetes with the Vault credential driver failed on
OpenShift in two ways: the OpenBao fixture pod was rejected by the
restricted-v2 SCC, and provider-creating tests hit HTTP 403 from
auth/kubernetes/login because OpenBao's Kubernetes auth was provisioned
only inside a single feature-gated test.

- Deploy OpenBao with the chart's OpenShift mode (global.openshift=true)
  when OpenShift is detected, so the pod inherits a namespace-assigned,
  SCC-compliant security context with no manual SCC grant. Hoist
  OpenShift detection ahead of the credential-driver fixtures so the
  flag is set before the fixture is deployed.
- Provision the OpenBao KV store, Kubernetes auth method, storage
  policy, and gateway login role in the harness (deploy_vault_fixture),
  making a Vault-backed gateway usable by the whole suite instead of
  only the credential_drivers test. Remove the now-redundant
  configure_vault_storage helper from the test.
- Harden openbao_exec so it tolerates only the idempotent "path is
  already in use" error on reruns and fails fast with output on any
  other error, instead of a blanket `|| true` that masked genuine
  failures (e.g. an unresponsive pod) until a later cryptic write.
- Document the OpenShift Vault credential-store SCC and Kubernetes-auth
  403 troubleshooting in the debug-openshell-cluster skill.

Signed-off-by: Jorge Garcia Oncins <jgarciao@redhat.com>
2026-09-14 21:55:43 +00:00
krishicks 5b9daab935 fix(ci): restore mise run ci on macOS (#3294)
- Replace BSD-incompatible in-place sed calls with portable temp-file rewrites.
- Remove test-only shell interception and capture generated gateway config
  directly.
- Allow parity tests to use supplied supervisor binaries without resolving a
  Linux target.
- Normalize temporary-directory paths and use portable RPM config installation.
- Set a valid setuptools-scm version for Python protobuf generation in Jujutsu
  checkouts.

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-09-12 00:17:15 +00:00
Jesse JaggarsandDrew Newberry 02b664bb0d refactor(config): normalize and enforce gateway schema v2 (#2814)
* refactor(config): normalize compute driver field names

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* refactor(config): introduce canonical gateway fields

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* refactor(config): enforce gateway schema version 2

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): preserve compute driver runtime guarantees

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): address schema v2 review regressions

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): complete schema v2 migration safeguards

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(config): expand schema v2 regression coverage

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(config): add schema v2 parity manifest

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): correct parity manifest inventory

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* docs(config): record schema v2 intentional changes

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* docs(config): disposition schema v2 parity gaps

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): add dual schema parity harness

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): establish compute lifecycle parity baseline

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): preserve gateway option compatibility

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): record gateway option parity

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* docs(config): close gateway-wide parity gaps

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(podman): apply configured pids limit

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): validate Podman option parity

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): add Kubernetes option parity harness

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): record Kubernetes option parity

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): disposition VM parity lanes

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): add external driver parity lane

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(e2e): preserve external driver pull policy

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): attest parity artifacts and launches

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): require clean parity build sources

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): bind parity runtime artifacts

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(e2e): use isolated supervisor tags

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(e2e): qualify parity image tags

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(e2e): serve parity supervisor locally

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): isolate parity podman services

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): harden parity evidence provenance

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): pin parity sandbox artifacts

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): attest parity runtime inputs

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): bind parity runtime evidence

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): record compute boundary parity

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): disposition cross-cutting parity lanes

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(packaging): preflight gateway config upgrades

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): preserve rebase integration guarantees

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(ci): isolate temporary git signing config

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): update remaining schema v2 consumers

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(ci): provide e2fs tools to VM tests

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): align preflight with gateway startup

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(vm): preserve rootfs tar configuration

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* chore(config): adopt duration unit constructors

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(packaging): preflight RPM gateway config

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): address driver review findings

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(e2e): require fresh semantic parity evidence

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(docker): update tests for renamed sandbox label

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(gateway): preserve selective driver coverage after rebase

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
2026-09-11 05:00:24 +00:00
Drew Newberry 1860010850 feat(sdk): add lazy pagination pagers (#3256)
* feat(sdk): add lazy pagination pagers

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(sdk): harden pager edge cases

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(sdk): cover initial resume token

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(go): fix all-workspaces pager examples

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-11 00:02:08 +00:00
Drew Newberry 33bbda3d33 refactor(persistence): adopt continuation-token pagination (#3249)
* refactor(persistence): adopt continuation-token pagination

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(pagination): address continuation review findings

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(tui): recover completed list refreshes

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(pagination): address review scalability findings

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(pagination): repair branch validation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(go): use page size in template example

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-11 00:02:07 +00:00
John T. MyersandJohn Myers 226a83b323 fix(supervisor): classify credential placeholders in request bodies (#3246)
* fix(supervisor): classify credential placeholders in request bodies

Closes #2904

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* fix(supervisor): preserve same-provider placeholders in request bodies

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <johntmyers@users.noreply.github.com>
2026-09-10 23:10:01 +00:00
Mrunal PatelandJohn Myers 90dbe5454b feat(api): add typed workspace selectors (#3245)
* feat(api)!: add typed workspace selectors

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(cli): preserve template workspace metadata

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* test(e2e): migrate workspace request selectors

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(api): update public schema inventory

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-09-10 20:38:50 +00:00
ShijuandDrew Newberry 8211274bfb fix(e2e): follow credential storage identity (#3202)
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Signed-off-by: Shiju <shiju@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
2026-09-09 22:44:40 +00:00
Shiju ea8eda6d5b feat(supervisor): enforce MCP request protocol versions (#3241)
* feat(supervisor): enforce MCP request protocol versions

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(supervisor): enforce MCP versions across HTTP forwarding

Apply shared request-version guards before authorization and after forward-request rewriting. Require version metadata to survive HTTP header cleanup, and cover valid initialization and selected-revision forwarding through middleware.

Signed-off-by: Shiju <shiju@nvidia.com>

---------

Signed-off-by: Shiju <shiju@nvidia.com>
2026-09-09 20:28:55 +00:00
John T. Myers f4dc6be4b2 refactor(inference): remove managed inference routes (#3195)
* refactor(inference): remove managed inference routes

Closes #3172

Remove the inference route control plane, inference.local data path, built-in router crate, and SDK surface. Move inference workloads to explicitly imported provider profiles and native endpoints, with migration cleanup and updated tests and documentation.

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(policy): preserve alternate upstream isolation

Restore the provider policy activation guard so legacy OpenAI and Anthropic providers configured for alternate base URLs do not grant egress to the built-in public vendor endpoints.

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-09-09 18:47:22 +00:00
Simon Scatton 48c449d8c8 chore(deps): replace ring with AWS-LC (#3243)
* chore(deps): replace ring with AWS-LC

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* fix(lint): address warnings after dependency upgrades

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* fix(tls): limit provider initialization to reqwest clients

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

---------

Signed-off-by: Simon Scatton <sscatton@nvidia.com>
2026-09-09 17:57:42 +00:00
Gaizka Menendez 8af79a7f4b fix(podman): resolve macOS Podman socket dynamically (#3135)
* docs(podman): document macOS socket path mismatch and dynamic lookup

On macOS, Homebrew-installed Podman does not create the default socket
path that the Podman driver probes. Document the OPENSHELL_PODMAN_SOCKET
override and the podman machine inspect lookup in both the compute
drivers reference and the debug-openshell-cluster skill.

Fixes #1690

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* fix(podman): resolve macOS Podman socket dynamically

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* chore: restore debug skill file

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* chore: drop legacy debug skill path

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* fix(podman): trim unrelated e2e changes

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* fix(e2e): harden shell array expansion

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* chore: remove unrelated skill note

* ci: retrigger checks

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

---------

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>
2026-09-09 14:10:59 +00:00
Evie Howard 6e6b3c8905 refactor(cli): remove local Dockerfile image builds (#3214)
Signed-off-by: Evie Howard <evhoward@redhat.com>
2026-09-09 13:28:17 +00:00
Philippe Martin 118b250f01 feat(sandbox): support rootfs tar as --from source for VM driver (#2863)
* feat(sandbox): support rootfs tar as --from source for VM driver

Accept flat rootfs tar archives (.tar, .tar.gz, .tgz) via the --from
flag for VM-backed gateways. The CLI detects the archive extension,
validates that the gateway uses the VM compute driver, and passes the
tar path through driver_config. The VM driver copies the tar into its
staging area and feeds it into the existing rootfs extraction and ext4
disk creation pipeline, skipping the container image pull/export steps.

Closes #2175

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(sandbox): validate rootfs tar path at the VM driver boundary

The rootfs_tar_path field in driver_config was passed from the API
caller directly to tokio::fs::copy without validation. An authenticated
user bypassing the CLI could supply arbitrary host paths (e.g.
/dev/zero for disk exhaustion, or readable host files for data
exfiltration).

Introduce a trusted staging directory that the VM driver creates on
startup and advertises via GetCapabilities. The CLI now copies the tar
into the staging directory before creating the sandbox, and the driver
validates that the received path is a regular file inside the staging
root and within a configurable size limit (default 10 GiB) before any
I/O.

New VmDriverConfig options:
- rootfs_tar_staging_dir: override the staging directory
  (default: <state_dir>/rootfs-tar-staging)
- rootfs_tar_max_bytes: override the size limit (default: 10 GiB)

Addresses GATOR-28b5152e-01.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(sandbox): request-scoped staging, size pre-check, and cleanup for rootfs tar

Tighten the rootfs tar staging flow to address the remaining GATOR-01
obligations:

- Request-scoped staging: the CLI creates a unique per-request
  subdirectory (req-<pid>) under the staging root instead of placing
  files directly in the shared directory. The driver enforces that the
  tar path is at depth 2 (staging_root/<subdir>/<file>), preventing
  cross-request path selection.

- Size pre-check: the driver advertises rootfs_tar_max_bytes via
  GetCapabilities. The CLI reads this limit and rejects oversized files
  before copying, avoiding disk exhaustion in the staging directory.

- Cleanup: the driver removes the request staging subdirectory after
  consuming the tar (on cache hit, copy success, or copy failure),
  ensuring staged data does not persist beyond the request.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(vm): restore rootfs-tar sandboxes from persisted image identity

On restore or restart, the one-shot staged tar archive has already been
cleaned up. Reading the persisted image identity from the sandbox state
directory and resolving the cached disk path directly avoids re-accessing
the deleted staging path.

Addresses GATOR-168b9210-01.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(cli): use random staging dirs and enforce byte limit during rootfs tar copy

Replace PID-based request staging directories with tempfile-generated
random names to prevent collisions and make paths unpredictable.
Replace bare tokio::fs::copy with a streaming copy loop that enforces
the advertised max_bytes limit during transfer, closing the TOCTOU gap
between the pre-copy size check and the actual copy.

Signed-off-by: Philippe Martin <phmartin@nvidia.com>
Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(sandbox): issue rootfs tar staging slots from the gateway

A caller could name any host path in `driver_config.vm.rootfs_tar_path`,
which the privileged VM driver then read. The CLI-side locality check did
not apply to direct API requests.

The gateway now owns staging. `BeginRootfsTarStaging` allocates a
request-scoped directory under the driver-advertised staging root and
returns an opaque single-use token; `CreateSandbox` carries the token, and
the gateway substitutes the path it allocated before dispatching to the
driver. `template.driver_config.<driver>.rootfs_tar_path` is rejected
outright in request validation, so a caller-supplied path never reaches
privileged I/O.

Tokens are bound to the issuing workspace and subject, consumed once, and
expire after 30 minutes. Outstanding slots are capped per caller and
overall, so one caller can neither exhaust the staging filesystem nor
starve others. An RAII guard reclaims the directory on every failure path
after consumption, and an age-gated sweep runs at startup and on each
reconcile pass for directories whose driver died before its own cleanup.

The token is stripped from the public sandbox before persistence: the
stored copy is returned verbatim by GetSandbox, ListSandboxes and
WatchSandbox to every member of the workspace.

Also fixes two defects this exposed:

- The CLI wrote `rootfs_tar_path` at the top level of `driver_config`, but
  the gateway forwards only `driver_config.<driver_name>`, silently
  dropping unmatched keys. The archive never reached the VM driver, so the
  documented `--from ./rootfs.tar` flow did not work at all. Config is now
  nested under `vm` and deep-merged, so a caller's existing VM settings
  survive instead of being clobbered by a shallow extend.

- Staging previously required `GetGatewayInfo`, which is restricted to
  `platform_admin`, making the feature unusable for ordinary users on any
  RBAC-enabled gateway. The new RPC matches CreateSandbox at
  `sandbox:write` / `workspace_role: user`.

`compute_driver.proto` is unchanged; the gateway reads the staging root
from the capabilities it already stores.

Refs #2175

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(vm): derive rootfs tar cache identity from archive contents

The prepared-disk cache key combined the archive's full path with an mtime
truncated to seconds, then mapped punctuation to `-`. Distinct paths such as
`/tmp/a/b.tar` and `/tmp/a-b.tar` collapsed onto the same key and reused each
other's disk, a rewrite within the same second kept stale contents, and a long
path could exceed filesystem component limits.

Identity is now a SHA-256 of the archive contents. This is also what makes the
cache work at all now that the gateway allocates a fresh staging directory per
request: a path-derived key would miss on every create.

The archive is hashed, the cache checked, and only on a miss copied — so a hit
skips writing a multi-gigabyte file. The copy is hashed as it is written and
rejected if the digest differs from the first pass, which closes the window
where the source changes during staging rather than approximating it with a
re-stat.

Refs #2175

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(vm): decompress gzip rootfs tar archives during staging

`--from` accepts `.tar.gz` and `.tgz`, but the driver staged whatever bytes
it was given and the guest image-prep VM extracts the staged file with a
plain `tar -xpf`. Compressed sources therefore depended on the guest tar
auto-detecting gzip, and the prepared disk was sized from the compressed
length, which is far too small for the expanded rootfs.

Staging now detects gzip from the archive's magic bytes -- the driver only
ever sees a gateway-issued staging path, never the caller's file name -- and
writes an uncompressed tar. The digest still covers the source bytes, so the
"archive changed while staging" check is unaffected, and expansion is bounded
by `rootfs_tar_max_bytes` so a compression bomb cannot fill the host disk.

`extract_rootfs_archive_to` sniffs gzip as well, so the host-side extraction
path matches.

Adds unit coverage for gzip staging, bounded expansion, and gzip extraction,
plus an e2e sandbox created from a gzip-compressed export.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

---------

Signed-off-by: Philippe Martin <phmartin@redhat.com>
Signed-off-by: Philippe Martin <phmartin@nvidia.com>
2026-09-08 22:36:40 +00:00
Jorge 519e5eb35f feat(e2e): make e2e:kubernetes work transparently on OpenShift (#3183)
* feat(e2e): make e2e:kubernetes work transparently on OpenShift

Running `mise run e2e:kubernetes` on OpenShift required manual namespace
creation, SCC grants, Helm value overrides, and cleanup. A separate
`e2e:openshift` task existed but only checked pod readiness without
running the Rust e2e test suite, and even with the suite wired up the
SSH-relay `sandbox connect` path stalled to the ready timeout because
`kubectl port-forward` cannot carry round-trip-heavy SSH over the
internet.

The harness now auto-detects OpenShift via the `route.openshift.io` API
group and, on OpenShift, both configures the cluster and switches the
gateway transport automatically:

- Drives the gateway through a passthrough OpenShift Route secured with
  mandatory mTLS instead of port-forward, so the connect suites
  (live_policy_update, port_forward, sync, connect-based
  sandbox_lifecycle, settings_management) actually pass. Computes the
  Route host from the cluster ingress domain, extracts client mTLS
  material from the openshell-client-tls secret, waits for the Route to
  serve mTLS, asserts a certless caller is rejected at the TLS
  handshake, and registers an mTLS CLI gateway pointing at the Route.
- Applies an SCC-compatible Helm values overlay that removes hardcoded
  runAsUser/fsGroup, letting OpenShift assign UIDs from the namespace
  range.
- Grants the privileged SCC to openshell-sandbox before Helm install
  and removes it during cleanup.
- Grants the anyuid SCC to the PostgreSQL fixture service account in
  DB scenarios and removes it during cleanup.
- All oc commands use --context to target the correct cluster.

The OpenShift e2e overlay (ci/values-openshift-e2e.yaml) turns TLS back
on, enables the Route, promotes the cert-verified caller to a dev
principal, and forces `image.pullPolicy`/`supervisor.image.pullPolicy`
to Always so runs against the `latest` upstream image use it instead of
a stale copy cached on the cluster nodes. Every OpenShift branch is
gated on OPENSHIFT_DETECTED, so the vanilla-Kubernetes port-forward path
is unchanged.

The Helm template for podSecurityContext is wrapped with {{- with }} so
null values omit the block instead of rendering invalid YAML.

The separate e2e:openshift task and e2e-openshift.sh script are removed
since e2e:kubernetes now covers OpenShift.

TESTING.md is updated with Kubernetes e2e documentation including
OpenShift auto-detection, dropping the e2e-host-gateway feature on
remote clusters, pinning IMAGE_TAG when the CLI and image versions
differ, task variants, and environment variables.

The debug-openshell-cluster skill gains an OpenShift platform row and
two SCC failure patterns (gateway rejected over hardcoded runAsUser,
sandbox missing the privileged SCC) covering the SCC handling and
podSecurityContext behavior this change introduces.

Signed-off-by: Jorge Garcia Oncins <jgarciao@redhat.com>

* fix(e2e): harden OpenShift SCC cleanup, mTLS gate, and Route timeout

Track the anyuid SCC grant for the PostgreSQL fixture with a dedicated
OPENSHIFT_POSTGRES_SCC_GRANTED flag set before the fixture apply, so a
failed apply no longer leaks the binding; cleanup now revokes it whenever
the grant succeeded, independent of deploy state.

Validate the Route server cert in the certless security gate (curl
--cacert instead of -k) and classify curl's exit code so only a TLS
client-auth rejection (35/56) counts as the expected certless rejection;
an unrelated DNS/timeout/TLS failure now fails loudly instead of masking
a potential mTLS hole.

Raise the OpenShift Route timeout in the e2e overlay. The default HAProxy
Route timeout is 30s, which severed long-lived transfers (large sandbox
upload/download, SSH-relay `sandbox connect`) mid-stream and failed the
sync e2e tests. Set both haproxy.router.openshift.io/timeout and
timeout-tunnel to 300s: a passthrough Route proxies in TCP mode, so
timeout-tunnel governs the established tunnel while timeout covers the
pre-tunnel phase.

Document the OpenShift transport exception, oc prerequisites and SCC
grants, and make the skopeo tag-check example copy-safe in TESTING.md.

Signed-off-by: Jorge Garcia Oncins <jgarciao@redhat.com>

---------

Signed-off-by: Jorge Garcia Oncins <jgarciao@redhat.com>
2026-09-08 20:37:19 +00:00
Evan Lezar e4369adcd0 chore(deps): replace serde_yml with noyalib (#3031)
* chore(deps): replace serde_yml with noyalib

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(providers): annotate generic YAML test values

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-09-07 13:22:27 +00:00
Shiju 592df3e014 feat(policy): preserve exact MCP revision allowlists (#3027)
* feat(mcp): add version-aware wire profile metadata

Signed-off-by: Shiju <shiju@nvidia.com>

* feat(policy): canonicalize MCP version allowlists

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(policy): align MCP policy tests with current main

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(policy): canonicalize supervisor protobuf ingress

Materialize defaultable MCP revisions before ambiguity checks, OPA construction, and sidecar delivery. Reject invalid sidecar policies with bounded errors.

Signed-off-by: Shiju <shiju@nvidia.com>

---------

Signed-off-by: Shiju <shiju@nvidia.com>
2026-09-05 04:24:49 +00:00
Evan LezarandDrew Newberry 1e1a8b5810 test(e2e): keep lifecycle sandboxes running (#3128)
* test(e2e): keep lifecycle sandboxes running

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(e2e): cover sandbox terminal lifecycle semantics

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(e2e): synchronize detached main exit checks

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Evan Lezar <elezar@nvidia.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
2026-09-02 18:58:02 +00:00
Philippe Martin 857af42a16 feat(vm): support corporate HTTP forward proxy egress for microVM sandboxes (#3090)
* feat(vm): support corporate HTTP forward proxy egress for microVM sandboxes

The corporate forward proxy machinery from #1792 is driver-agnostic and
already merged: openshell-supervisor-network implements CONNECT chaining,
NO_PROXY matching, credentials, https:// proxies and corporate CA trust, and
openshell-sandbox exposes it as six argv-only flags. Podman gained the driver
half in #2245/#2512 and Kubernetes in #2633; the VM driver had none of it, so
VM sandboxes on proxy-only networks could not reach any destination requiring
the proxy even when policy allowed it.

The blocking piece was not proxy logic but delivery: the VM guest init script
runs as PID 1 and execs a fixed supervisor command line, and libkrun's
krun_set_exec receives an empty argv, so there was no channel for driver-owned
supervisor arguments. The supervisor's proxy flags deliberately have no
environment fallback, and build_guest_environment merges user-supplied
environment, so the guest env is not a safe transport either.

Add a driver-authored argument file, mirroring the existing init.d manifest:
the driver writes /opt/openshell/supervisor-args into the overlay upperdir on
every launch and the guest reads it verbatim, one argument per line, appending
it to every supervisor exec. It is written even when empty, which is what makes
the channel unforgeable -- the upperdir always shadows the read-only image
layer, so an image can neither supply its own arguments nor disable the
operator's by omitting the file. Because both launch backends exec the same
init script, this covers libkrun and QEMU without touching either.

A microVM has no bind mounts or container secrets, so the credential and CA
bundle are staged into the per-sandbox overlay the way the gateway JWT already
is: credential root-only at 0600, CA at 0644, both rewritten every launch so a
removed setting clears prior material, and both deleted with the sandbox state
directory. This places the credential at rest in the overlay image on the
gateway host, which differs from the Podman secret model and is documented as
an explicit security consideration.

Validation is fail-closed and shared: a new
openshell_core::driver_utils::validate_upstream_proxy_settings holds the
pairing rules the Podman driver established, and both the gateway and the
driver call it so an invalid table names the offending key instead of
surfacing as an opaque driver-readiness timeout.

Guest egress leaves through gvproxy, so a proxy on the gateway host's loopback
is reachable only through host.openshell.internal; the guest to gateway
callback is unaffected.

Closes #3088

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(vm): bound the proxy CA read and scope the host-loopback recipe

Two review findings on the corporate forward proxy support for microVM
sandboxes.

The driver read the operator's proxy_ca_bundle with an unbounded fs::read and
accepted it on a substring match for the PEM BEGIN CERTIFICATE marker. A
special file such as /dev/zero therefore grew driver memory without bound on
every authorized sandbox create, and a PEM block holding invalid DER passed
the host check but contributes no trust anchor in the guest, so every
supervisor would fail after boot with an error attributed to the sandbox
rather than to the setting.

Move the read into openshell-core as read_upstream_proxy_ca_bundle_file: it
reuses the credential reader's bounded-read path (non-regular files rejected
on fstat, size capped, read bounded even if the file grows), then requires at
least one anchor that RootCertStore::add_parsable_certificates accepts. The
supervisor's own reader now delegates to it, so host acceptance and guest
acceptance are the same function and cannot drift.

The published host-loopback recipe was written for libkrun only. gvproxy NATs
host.openshell.internal to the gateway host's 127.0.0.1, but GPU sandboxes run
on the QEMU/TAP backend where that name resolves to the TAP host address and
the driver's own nftables input chain accepts only the gateway port from the
guest — no proxy on the gateway host is reachable there at any bind address,
so an operator following the generic recipe lost all proxy-required egress
while configuration validation succeeded.

Scope the recipe to libkrun in every reference and reject a gateway-host proxy
URL when a launch plan resolves to QEMU, naming the reason, instead of booting
a sandbox whose policy-approved CONNECTs all time out.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(vm): match the QEMU proxy preflight to the selected TAP host

The gateway-host proxy guard added for the QEMU/TAP backend classified the
wrong set of addresses in both directions.

It ran at the top of configure_qemu_launch_plan, before the subnet allocation
that settles plan.host_ip, so it could not compare against the address the
guest actually reaches the host on. An operator pointing https_proxy at the
sandbox's own TAP host address, such as 10.0.128.1, passed the check, and the
driver's nftables input chain — which accepts only the gateway port from the
guest — then dropped every policy-approved CONNECT, which is exactly the
silent timeout the guard exists to prevent.

In the other direction it rejected 192.168.127.254 unconditionally. That
address is special only to libkrun/gvproxy; on QEMU/TAP it is an ordinary
address that may be routable through the guest's masqueraded egress, so the
guard refused a working configuration.

Run the check after the launch plan's network allocation, on both the
freshly-allocated and already-complete paths, and compare IP literals with
that sandbox's selected TAP host. Loopback literals, localhost, and the
documented host aliases that write_host_gateway_aliases seeds to the TAP host
still classify as the gateway host, and the failure names the address. The
gvproxy host-loopback constant returns to being a documentation anchor.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

---------

Signed-off-by: Philippe Martin <phmartin@redhat.com>
2026-09-02 14:56:08 +00:00
grs 74960ebfae feat(server): add sandbox templates (#2833)
* feat(server): add sandbox workload templates

Signed-off-by: Gordon Sim <gsim@redhat.com>

* feat(go-sdk): add sandbox workload template support

Signed-off-by: Gordon Sim <gsim@redhat.com>

* feat(rust-sdk): add sandbox workload template support

Signed-off-by: Gordon Sim <gsim@redhat.com>

* feat(python-sdk): add sandbox workload template support

Signed-off-by: Gordon Sim <gsim@redhat.com>

* feat(typescript-sdk): add sandbox workload template support

Signed-off-by: Gordon Sim <gsim@redhat.com>

* docs(agents): document sandbox workload templates

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(cli): support default GPU requests in sandbox templates

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(server): cap sandbox templates per workspace

Signed-off-by: Gordon Sim <gsim@redhat.com>

* docs(architecture): document sandbox workload template boundaries

Signed-off-by: Gordon Sim <gsim@redhat.com>

* feat(cli+sdk): expose sandbox workload template provenance

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(cli): include sandbox template annotations in output

Signed-off-by: Gordon Sim <gsim@redhat.com>

* feat(sandbox): add label selectors to template listing

Signed-off-by: Gordon Sim <gsim@redhat.com>

* test(e2e): cover sandbox template failure paths

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(server): preserve command and ttl when creating sandbox from template

Signed-off-by: Gordon Sim <gsim@redhat.com>

* test(server): add field coverage test for template merge

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(go-sdk): add pagination support to fake client

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(cli): warn on env vars that looks like secrets

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(sdk-ts): propagate sandbox workspace through lifecycle calls

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(docs): update workspace management docs

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(ts-sdk): support command and tty when creating from template

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(go-sdk): support command and tty when creating from template

Signed-off-by: Gordon Sim <gsim@redhat.com>

* test(python-sdk): verify command and tty handling when creating from template

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(go-sdk): update docs and ClientInterface

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(server): validate sandbox create specs before I/O

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(go-sdk): guard empty DNS-1123 label validation

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(python-sdk): allow empty template builder mappings

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(cli): align template GPU JSON default output

Signed-off-by: Gordon Sim <gsim@redhat.com>

---------

Signed-off-by: Gordon Sim <gsim@redhat.com>
2026-09-02 00:55:29 +00:00
Drew Newberry 9ca19e6c80 refactor(compute): decouple gateway driver composition (#2823)
* refactor(compute): decouple gateway driver composition

Move first-party composition and VM process ownership into openshell-gateway, leaving openshell-server backend-independent. Update packaging and build references with the new crate, simplify the compiled-driver boundary, and keep the driver-free gateway path buildable with bundled Z3 tooling.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(telemetry): bound compute driver categories

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(core): keep runtime transport generic

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(compute): complete server driver decoupling

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(compute): preserve driver integrations after rebase

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>

* fix(compute): preserve docker tracing after decoupling

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>

* fix(compute): preserve driver behavior after extraction

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(compute): remove MXC policy side channel

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(compute): separate policy delivery from readiness

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>
2026-09-01 21:13:45 +00:00
John T. MyersandJohn Myers 07df822090 feat(providers): make profiles authoritative (#2962)
* feat(providers): make profiles authoritative

Closes #1988

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* docs(providers): move profiles into provider navigation

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* docs(providers): clarify provider attachment lifecycle

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* fix(tui): scroll provider profile picker

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* fix(providers): honor profile credential semantics

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* fix(providers): prefer exact profile IDs

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* fix(providers): harden authoritative profile adoption

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* test(oidc): align provider fixtures with profiles

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* fix(providers): preserve authoritative profile lifecycle

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <johntmyers@users.noreply.github.com>
2026-09-01 19:22:07 +00:00
John T. MyersandJohn Myers a547dc9f4d fix(sandbox): reconcile early container exits (#3101)
* fix(sandbox): reconcile early container exits

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* test(oidc): isolate RBAC checks from lifecycle exits

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <johntmyers@users.noreply.github.com>
2026-09-01 18:42:48 +00:00
Evan Lezar bb70461878 test(e2e): run conformance in gateway lanes (#2925)
* test(e2e): isolate VM-specific smoke assertions

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(e2e): add portable CLI conformance baseline

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* feat(conformance): add standalone CLI runner

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(e2e): run conformance in gateway lanes

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-09-01 12:36:51 +00:00
Drew Newberry 69a05ebb3b fix(sandbox): complete successful main processes (#2884) 2026-08-28 18:21:20 -07:00
Simon Scatton 981606d2f8 ci: build release binaries with Nix (#2977)
* ci: build release binaries with Nix

Refs #1683

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* ci: build VM artifacts with Nix

Refs #1683

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* ci: build images from Nix artifacts

Refs #1683

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* fix(nix): prevent host header leakage

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* ci: parallelize artifact builds

Refs #1683

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* ci: build external driver test artifacts

Refs #1683

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* fix(nix): disable mold in musl shells

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* ci: key Rust cache by Nix shell derivation

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* ci: refactor end-to-end workflows

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* ci: split platform binary workflows

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* ci: remove obsolete native build workflows

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* ci: replace disallowed mise action

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* ci: fix refactored e2e lanes

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* ci: check out local result action

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* ci: cache mise installations

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* ci: run docker builds on host runners

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* ci: disable unstable kubernetes e2e lanes

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* fix(ci): scope binary builds to cargo packages

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* fix(ci): address zizmor template injection findings

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* fix(ci): resolve remaining zizmor annotations

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

---------

Signed-off-by: Simon Scatton <sscatton@nvidia.com>
2026-08-27 17:25:06 +00:00
Evan Lezar 9f88f8ff9b ci: remove rootless podman e2e lane (#2981)
Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-08-27 08:48:44 +00:00
Mrunal Patel 18ce13b9b1 feat(providers): expose actionable OAuth refresh failures (#2887)
* fix(providers): classify OAuth refresh failures

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* test(providers): add Keycloak refresh e2e lane

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(providers): harden OAuth refresh recovery

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(providers): classify post-mint refresh failures

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(providers): tolerate malformed OAuth subtypes

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

---------

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
2026-08-25 18:00:00 +00:00
Seth Jennings 5206bc51b2 feat(sdk): add OAuth Client Credentials support to SDKs (#2907)
* feat(sdk): add renewable client credentials auth

Implement lazy OAuth client-credentials acquisition and renewal for the Python, TypeScript, and Go SDK clients, with shared security conformance coverage and service-account documentation.\n\nCloses #2803

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(sdk): address client credentials review

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(sdk): address client credentials review

Signed-off-by: Seth Jennings <sjenning@redhat.com>

---------

Signed-off-by: Seth Jennings <sjenning@redhat.com>
2026-08-25 17:53:47 +00:00
Russell Bryant 72b9c4acef test(e2e): pin direct podman calls to harness socket on macOS (#2909)
The podman_gateway_start e2e test shells out to `podman` directly (unlike
every other Podman e2e test, which routes through the gateway). The harness
`e2e/with-podman-gateway.sh` clobbers `XDG_CONFIG_HOME` with an empty dir to
isolate CLI/SDK gateway metadata. On macOS the podman client resolves its VM
connection config from `XDG_CONFIG_HOME`, so a bare `podman ps` falls back to a
nonexistent native rootless socket and fails with "unable to connect to Podman
socket". On Linux the socket resolves via `XDG_RUNTIME_DIR`, so the test passes
there and this is a macOS-only false failure.

Target the same API socket the gateway uses by passing `--url unix://$SOCKET`
when the harness-exported `OPENSHELL_PODMAN_SOCKET` is set, mirroring the
shell's `podman_cmd` helper. When the var is unset (running the test outside the
harness), fall back to plain `podman`, so Linux behavior is unchanged.

Signed-off-by: Russell Bryant <rbryant@redhat.com>
2026-08-24 22:04:25 +00:00
Philippe Martin 0a1f246587 feat(sandbox,podman): trust corporate CA for https:// proxies and intercepted TLS (#2512)
* feat(sandbox,podman): trust corporate CA for https:// proxies and intercepted TLS

The corporate proxy chaining only accepted plain http:// proxy URLs, so
operators whose forward proxy terminates TLS with a private corporate CA
had no way to reach it, and TLS-intercepting proxies (mitmproxy, squid
ssl-bump) that re-sign tunneled server certificates broke every upstream
handshake after CONNECT.

The supervisor now accepts https:// proxy URLs: it wraps the connection to
the proxy in TLS before the CONNECT handshake, verifying the proxy
certificate against the built-in Mozilla roots, the system CA bundle, and an
optional operator corporate CA bundle. The upstream dial returns a
Plain/Tls stream enum consumed generically by the relay paths.

The corporate CA is delivered as a driver-supplied command-line argument
(--upstream-proxy-ca-bundle), never an environment variable, matching the
hardened proxy-config model where a sandbox image cannot influence the
operator's egress boundary. It is folded into the sandbox combined trust
bundle (write_ca_files) and the L7 upstream verification store
(build_upstream_client_config) at startup, so intercepted upstream
handshakes succeed and sandbox workloads trust the re-signed certificates.
Configuration is fail-closed: a CA bundle set without a proxy, or an
unreadable or certificate-free file, is fatal rather than silently
weakening the trust boundary.

The shared parse_upstream_proxy_url validator accepts https:// (recording
the scheme so the driver and supervisor agree), keeping the explicit-port
requirement. The Podman driver gains a proxy_ca_bundle operator setting
(TOML, --sandbox-proxy-ca-bundle, OPENSHELL_SANDBOX_PROXY_CA_BUNDLE) that
bind-mounts the host PEM read-only into the sandbox (a CA certificate is not
secret) and points --upstream-proxy-ca-bundle at it, with a create-time
readability check.

The standalone dev gateway task passes OPENSHELL_SANDBOX_PROXY_CA_BUNDLE
through to the generated podman config, so a local gateway can be pointed at
a TLS-intercepting proxy without hand-editing the regenerated TOML.

Refs #1792

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(sandbox): reject CA bundles with valid PEM framing but invalid X.509 DER

The proxy CA bundle validation counted PEM blocks that base64-decoded
successfully, but did not verify the decoded bytes were accepted as
trust anchors by RootCertStore. A bundle with syntactically valid PEM
framing but invalid DER would pass the startup check while contributing
zero usable anchors, causing opaque TLS failures at runtime instead of
a fail-closed startup error.

Validate decoded certificates through RootCertStore::add_parsable_certificates
and reject the bundle unless at least one is accepted.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(sandbox,podman): allow proxy auth without insecure acknowledgement for https:// proxies

For an https:// proxy the Proxy-Authorization credential travels inside
the verified TLS session, so the proxy_auth_allow_insecure
acknowledgement is unnecessary. Previously both http:// and https://
proxies required it, producing a misleading cleartext-risk diagnostic
for a path that is already encrypted.

Skip the requirement when the proxy URL uses https://; the
acknowledgement is still tolerated if set. Updated in both the
supervisor and Podman driver validation paths, with docs and tests.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix: format

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(kubernetes): use truly unsupported scheme in proxy validation test

https:// is now a supported proxy scheme after a13c4dce, so the
unsupported-scheme test must use a genuinely unsupported scheme.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(e2e): sign the https proxy fixture listener cert with a CA

The corporate-proxy E2E fixture served a single `openssl req -x509`
certificate as its TLS listener identity. OpenSSL marks that certificate
`basicConstraints: critical, CA:TRUE`, and rustls refuses a CA
certificate presented as an end-entity certificate (CaUsedAsEndEntity).
The supervisor's TLS handshake with the proxy therefore failed, the
upstream dial errored, and the workload's CONNECT was dropped without a
response, so podman_corporate_proxy_trusts_ca_bundle_for_https_proxy
failed on the approved destination while policy denial still worked.

Generate a corporate CA and a separate listener leaf signed by it, serve
the leaf chain, and publish only the CA as the bundle the supervisor
trusts. This is what an intercepting proxy actually presents, and it
exercises the corporate-CA trust path rather than pinning the listener
certificate itself.

Refs #1792

Signed-off-by: Philippe Martin <phmartin@redhat.com>

---------

Signed-off-by: Philippe Martin <phmartin@redhat.com>
2026-08-24 21:10:25 +00:00