Commit Graph
221 Commits
Author SHA1 Message Date
0a770d9173 feat(kubernetes): support HA gateway rebalancing (#1868)
* feat(kubernetes): support HA gateway rebalancing

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* perf(server): cache peer connections, tokens, and owner lookups

Every forwarded relay rebuilt its setup from scratch: an owner lookup, a
blocking read of the peer token, a TLS connect to the owning replica, and
a TokenReview plus Pod GET on the receiving side. Sandbox service routing
does this per HTTP request, so the apiserver calls scaled with traffic.

Cache all of it on ServerState:

- peer channels pooled per endpoint, so relays multiplex over one
  connection instead of redialing
- peer tokens keyed by SHA-256, expiring at min(ttl, token exp) so a hit
  cannot accept an expired token
- owner records for 3s against a 45s ownership TTL, still freshness
  checked before use

Entries are evicted when a relay fails. Also raise HTTP/2
max_concurrent_streams to 1024, since pooling funnels every relay between
two replicas onto one connection and hyper's default of 200 sits below
the 256 pending-relay budget.

Signed-off-by: divesh <dgude@nvidia.com>

* perf(server): pool upstream connections for sandbox services

Each HTTP request to a sandbox service opened its own supervisor relay,
paying a new TCP connection and HTTP/1 handshake every time. Worse, it
counted against the 32 in-flight relay cap, so a service handling more
than 32 concurrent requests failed outright.

Pool idle upstreams per endpoint and port, up to 8 each for 15s. Reuse is
safe because the pool only returns a connection hyper reports as ready,
and HTTP/1 cannot start a request until the previous body has drained.
Upgrades are never pooled since they take the connection over, and a
failed send evicts that endpoint. Pruning is bounded per key, with the
full sweep limited to once per 30s.

Signed-off-by: divesh <dgude@nvidia.com>

* fix(server): address HA gateway review findings (#3449)

- Let a gateway own supervisor sessions without a peer endpoint. Requiring
  one whenever the store is PostgreSQL broke every single-instance
  PostgreSQL deployment, because no sandbox supervisor could connect.
  A cross-replica request to an owner that advertises no endpoint now fails
  immediately naming the cause, instead of retrying until the wait timeout.
- Close a supervisor session on heartbeat only when another replica owns it,
  or after renewals fail for the ownership TTL. A database error no longer
  drops every session heartbeating during an outage.
- Clamp owner record ages at zero so a skewed or corrupt stored timestamp
  cannot produce a negative age.
- Bound the cross-object advisory lock with a lock timeout, so a stuck holder
  fails instead of blocking every mutation in the fleet.
- Refuse to start when a peer endpoint is configured on a multi-replica
  backend but peer authentication is unavailable, and warn when a
  multi-replica backend has no peer endpoint at all.
- Reject a plaintext peer endpoint when the gateway serves TLS.
- Skip the sandbox watch poller on single-replica backends, where the local
  update bus already sees every write.
- Rate-limit the peer owner cache sweep so an insert no longer scans the
  whole map under the lock.
- Retry GET and HEAD on a pooled upstream the sandbox closed, instead of
  returning 502, and drop an emptied endpoint from the pool right away.
- Document the gateway peer environment variables and the post-rollout
  ownership skew operators should expect.

Signed-off-by: divesh <dgude@nvidia.com>

* fix(server): harden HA supervisor ownership

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Signed-off-by: divesh <dgude@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: divesh <dgude@nvidia.com>
Co-authored-by: Divesh Chowdary <47188680+FrostGod@users.noreply.github.com>
2026-09-21 21:22:12 +00:00
Drew Newberry dee4f98dda chore(vm): refresh runtime defaults and hardening (#3446)
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-21 10:19:30 +00:00
Drew Newberry d6f3e1f5f2 fix(packaging): keep SPDX comments out of Debian control (#3483)
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-19 18:55:33 -07:00
Drew Newberry 17ce738bfb fix(ci)!: remove gateway callback listener dependency (#3365)
* fix(ci): repair post-merge release canary

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(packaging): bootstrap canary runtime prerequisites

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci(canary): collect macOS VM diagnostics

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci(canary): pin libkrun-compatible macOS runner

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci(canary): limit macOS smoke test to package startup

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(compute)!: remove gateway callback listeners

Run Docker supervisors on host networking so they use the operator-configured primary gateway endpoint. Remove the unused compute-driver callback listener negotiation and listener-scoped routing machinery.

BREAKING CHANGE: The ComputeDriver API no longer exposes GetGatewayListenerRequirements or GatewayListenerRequirement. External drivers must regenerate bindings and connect supervisors to the configured primary gateway endpoint.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(docker): use sandbox runtime image in launcher

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(podman): exercise production endpoint selection

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(e2e): route supervisors to reachable gateways

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(e2e): align Podman endpoint fixtures

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): preserve host aliases for supervisors

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): align sandbox host gateway pin

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(e2e): address Docker fixtures by bridge IP

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(e2e): serialize sandbox lifecycle cases

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(e2e): host Docker TCP fixture with gateway

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(e2e): use loopback for host-network supervisor

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-18 20:55:55 +00:00
Drew Newberry d91b1999a0 feat(api)!: use sandbox names as canonical RPC references (#3272)
* feat(api)!: use sandbox names as canonical references

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(cli): update forward color fixture for workspace scope

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(supervisor): use sandbox names for settings lookup

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(e2e): use canonical sandbox request fields

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(e2e): use canonical sandbox receipt field

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(server): harden sandbox mutation handling

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(ci): update rebased sandbox references

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(api)!: standardize canonical entity references

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(api): codify protobuf API conventions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(api): preserve workspace selector semantics

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(api): restore workspace selector parity

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(api): preserve descriptive name fields

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(api): update e2e request fixtures

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(core): omit workspace selector during bootstrap

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* chore(api): remove proto convention checker

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* chore(api): refresh schema fingerprints after rebase

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(cli): use canonical provider receipt field

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-18 11:39:55 -07:00
Florent BENOIT c5a8c4d220 chore(vm): bump libkrun to v1.19.4 and libkrunfw to v5.6.1 (#3451)
The libkrun and libkrunfw projects moved from the containers/ GitHub
org to libkrun/. Update all clone URLs and version pins accordingly.

libkrun v1.19.4 introduces krun_add_virtiofs4 with a permissions
semantics parameter, enabling correct file ownership mapping on
host-to-guest bind mounts — required for issue #2585.

- libkrun: v1.17.4 → v1.19.4 (libkrun/libkrun)
- libkrunfw: 463f717b → v5.6.1 (libkrun/libkrunfw)
- Centralise LIBKRUN_REF in pins.env instead of hardcoding in each
  build script
- Source pins.env in the macOS build script for consistency with the
  Linux build script
- Remove the init/init make target step — since libkrun/libkrun@05c4eb7
  the init blob moved into the init_blob crate and is built by cargo

Ref: https://github.com/NVIDIA/OpenShell/issues/2585

Signed-off-by: Florent Benoit <fbenoit@redhat.com>
2026-09-18 13:32:04 +00:00
Evan Lezar 2263685cf3 test(tmachine): migrate Keycloak provider refresh coverage (#3404)
* test(tmachine): add Keycloak provider refresh suite

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* refactor(tmachine): share container runtime detection

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* ci(tmachine): run feature suites in GitHub Actions

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* ci(tmachine): run conformance with Podman tests

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* ci(tmachine): cover provider refresh with Podman

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* ci(integration): split input preparation from runners

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-09-18 13:56:07 +02:00
Piotr Mlocek c108c31696 docs(fern): sync announcement configuration (#3436)
* docs(fern): sync announcement configuration

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(docs): scope announcements to synced channel

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(fern): use dev as source version

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(fern): use channel-neutral logo link

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

---------

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-09-18 03:43:01 +00:00
Drew NewberryandPiotr Mlocek 8b77925eb7 feat(installer): support prerelease installations (#3364)
* feat(installer): support prerelease installations

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(readme): clarify 0.1.0 production readiness

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(readme): refine production readiness message

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(release): publish rolling prerelease channels

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(release): keep prereleases out of GitHub releases

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(readme): refine 0.1.0 notice

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(readme): expand 0.1.0 notice

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(readme): simplify 0.1.0 notice

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(installer): simplify prerelease alias

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(installer): require successful prerelease runs

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(release): reuse exact prerelease artifacts

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(installer): include prover in prerelease bundle

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci: remove prerelease post-publish validation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(fern): sync announcement configuration

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs: remove stale prerelease canary claim

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(fern): defer announcement synchronization

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(fern): simplify release announcement

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Co-authored-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-09-18 00:53:36 +00:00
Philippe MartinandJohn Myers 04146692d9 refactor(providers)!: make provider profiles import-only (#3383)
* chore(providers): remove dead provider plugin modules

Twelve modules under crates/openshell-providers/src/providers/ were never
declared in providers/mod.rs, so they have not been compiled since the plugin
registry was narrowed to the two adapters it still registers. Four of them
(generic, gitlab, opencode, outlook) key off provider type identifiers that
normalize_provider_type already retired.

Keep google_cloud and vertex, which are the only plugins ProviderRegistry::new
registers.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* test(providers): load example profiles from providers/ at test time

Add an example_profiles module that reads the YAML under providers/ from the
source checkout at run time and parses it with the existing profile loader. It
is gated behind a new non-default example-profiles feature so it is available to
this crate's own tests and, once wired into dev-dependencies, to gateway and CLI
tests, while never reaching a release binary.

Nothing consumes it yet; later changes move the test fixtures off the compiled
catalog and onto these files.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* test: source profile fixtures from providers/ instead of the compiled catalog

Unit and integration fixtures reached into builtin_profiles() to get a profile
to work with, which ties the tests to the compiled catalog rather than to the
files an operator would import. Point them at the example_profiles loader
instead.

The profiles.rs tests keep their coverage unchanged and become explicit golden
tests over providers/*.yaml: they are what keeps those files valid once nothing
compiles them. builtin_profiles_are_sorted_by_id widens into a test that the
whole example set parses, sorts, and lints clean as one catalog.

The CLI fake gateways serve the example profiles the way a real gateway serves
what an operator imported, through a shared helper.

No behavior change: the gateway still loads the built-in source by default and
serves the same profiles.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* docs(providers): document the example provider profiles

Each file under providers/ now opens with a header naming its expected client
binary identities, the image layout those paths assume, the credential scope,
the endpoint access it grants, and a smoke test. Add a README covering the
import commands and why a profile should be copied and edited rather than
imported unchanged.

Several of these profiles bind network access to paths that only exist in the
OpenShell Community image — /sandbox/.venv, /app/.venv, /sandbox/.cursor-server,
/usr/lib/node_modules. Imported unchanged into another image the profile matches
nothing: the catalog still advertises it, but the credential is never injected
and the traffic is denied. The headers say so where it applies.

Comments only; the profile schema has no documentation fields and all fifteen
files still lint clean.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* test(server): seed unit-test state with the example provider profiles

Server unit tests inherited the builtin + user source default from ServerState,
so around a hundred and forty assertions about github, openai and the rest
resolved against the compiled catalog. Point the test state at the user source
alone and import the example profiles from providers/ into its store first, the
way an operator would. Every one of those assertions keeps passing unchanged,
which is the point: it proves the gateway behaves identically with an imported
catalog before the default moves.

Four tests asserted builtin-source semantics specifically. Under import-only the
only profiles a gateway cannot edit are the ones a non-user source vends, so
they now exercise a source-managed profile composed with the user source; the
read-only guard they cover is the one that still applies to interceptor
catalogs. The list test asserts the imported profile's user/platform identity
instead of builtin with an empty scope.

test_server_state_with_user_only_github_profile is gone: the default test state
now is a user-only gateway with github imported.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* refactor(cli): resolve credential suggestions from the gateway catalog

The --env credential warning scanned a profile table compiled into the CLI, so
its suggestions described the binary rather than the gateway the user is talking
to: a profile the gateway does not serve was suggested anyway, and an imported
custom profile never was.

Fetch the catalog from the connected gateway instead. The warning moves out of
argument parsing and into sandbox create and sandbox template create, where a
client already exists. A catalog fetch failure is not fatal — the warning
degrades to its generic form rather than blocking sandbox creation.

Extract the ListProviderProfiles paging loop from provider list-profiles into a
shared fetch_provider_profile_catalog, which the profile-driven paths now share.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* refactor(cli)!: infer providers from the gateway catalog

Command-to-provider inference went through a hardcoded alias table in
openshell-providers, a second copy of the built-in catalog's identifiers. A
custom profile could never be inferred no matter what binaries it declared, and
the table drifted from the profiles it mirrored.

Infer from the connected gateway's catalog instead: match the command's basename
against each profile's ID and against the basenames of the binaries the profile
authorizes. A profile that names /usr/bin/claude is the profile for running
claude. The match must be unique — where several profiles claim a command, the
user names one with --provider — and an empty catalog infers nothing.

No command is special-cased. `binaries` is the operator's authorization
statement, so a profile that declares a binary claims the command that runs it,
whatever that binary is; narrowing that belongs in the profile rather than in a
list compiled into the CLI, which could never cover an unbounded catalog anyway.

Breaking: the retired aliases stop resolving, and commands the old table never
listed can now infer. git, pip and uv are declared by the github and pypi
example profiles, so they infer where they previously did not.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* refactor(providers,server)!: resolve provider profiles by exact ID

normalize_provider_type was a hardcoded alias map — gh to github, claude to
claude-code, vertex to google-vertex-ai — and a second copy of the built-in
catalog's identifiers, independent of the YAML it mirrored. It let a provider
type resolve to a profile the operator never named, and it made the built-in IDs
behave as a reserved namespace.

Remove it, along with detect_provider_from_command and the alias-normalizing
ProviderRegistry::inject_env. A provider type now names a profile exactly:

- the effective catalog resolves an ID or reports it absent, with no alias retry
- plugins activate only for the ID of a profile the gateway resolved; a provider
  with no resolvable profile gets no plugin projection, instead of falling back
  to an alias guess
- the CLI surfaces the gateway's not-found instead of retrying under an alias
- telemetry buckets by profile ID, and the gitlab, opencode and outlook buckets
  go with the aliases that were their only source

Breaking: `--type gh`, `--type claude` and the other aliases no longer resolve.
Use the profile's own ID.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* feat(server): fail closed when a provider's profile is absent

A sandbox composed from a provider whose profile the gateway cannot resolve
started anyway: the credential and policy builders warned and skipped, so the
sandbox came up carrying none of that provider's credentials or network policy.
The operator learned about it later, as a denied connection or a missing
environment variable, rather than as the configuration error it is.

Check at the two composition boundaries — CreateSandbox and
AttachSandboxProvider — right beside the catalog snapshot already taken there,
and reject with a bounded diagnostic naming the provider, the profile it refers
to, and the import command that supplies it. Scope-aware: a platform-scoped
provider is told to import with --global.

Read paths are untouched. ListProviders, GetProvider and profile export keep
working so an operator can see and recover an affected provider, and the shared
policy builders keep their warn-and-skip for the diagnostic paths that also
reach them.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* refactor(server): scope the vendor base-URL pin to declared endpoints

provider_profile_endpoints_are_active withheld a credential and the provider's
policy layer when an openai or anthropic provider pointed its client somewhere
other than the public vendor endpoint. The guard was keyed on
profile.source == "builtin" and on those two profile IDs, so it protected only
profiles OpenShell shipped. Once profiles are import-only no profile is ever
builtin, and the guard would silently stop applying — including to an operator
who imported providers/openai.yaml verbatim.

Key it on the profile instead of on where the profile came from. A profile's
endpoints are the boundary its credential is bound to, so if the provider
configures a *_BASE_URL pointing at a host the profile does not declare, the
profile no longer describes where that credential goes and is treated as
endpointless. Host matching reuses the DNS-label-aware matcher in
openshell-core, so wildcard endpoints such as Vertex's
*-aiplatform.googleapis.com resolve correctly.

A profile with no declared endpoints has no boundary to contradict, and a config
value that names no host is not a redirect this can reason about. Both keep the
profile active. The control now covers every endpoint-bearing profile, including
an operator's own, and the last "builtin" string leaves the gateway.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* test(e2e): import example provider profiles during gateway bring-up

The e2e suites create providers from github, openai, nvidia, claude-code and
google-cloud, and the GitHub lane runs a real git clone through the profile's
binary attribution. A gateway serves only the profiles an operator imported, so
the lanes have to import them.

Add e2e_import_example_provider_profiles to the shared bring-up helpers and call
it from the Docker, Podman, Kubernetes and VM wrappers once the gateway is
healthy and registered. It runs the same command the upgrade notes give
operators, against the repository's own providers/ directory, so the lanes
exercise the documented path rather than a test-only shortcut.

Lands before the default changes: a same-ID user profile already shadows a
built-in, so importing works today.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* refactor(providers)!: make provider profiles import-only

OpenShell compiled fifteen provider profile YAML files into every release binary
and selected them by default, so a fresh gateway published a catalog it was
never configured with. Those profiles are not image-neutral: their binary
selectors name paths that exist in the OpenShell Community image, so changing
the sandbox image could make a profile inert while the catalog still advertised
it. Every endpoint, credential name and binary path in providers/ was also
effectively part of the 0.1.0 public contract.

A gateway's catalog is now exactly what an operator imported:

- providers/*.yaml is no longer include_str!'d, and builtin_profiles() is gone
  from the openshell-providers API
- the default provider_profile_sources is [{ type = "user" }], and a gateway
  with nothing imported reaches ready and serves an empty catalog — an empty
  catalog is a valid state, not a startup failure
- the builtin source type is removed from the configuration schema, and a
  gateway.toml that still names it is rejected at parse time with the import
  command rather than an unknown-variant error
- nothing reserves the canonical identifiers any more, so github, pypi,
  anthropic and the rest import at their own IDs and the imported profile is the
  only definition for that ID; the static_fallback that kept a shipped
  definition resident behind an imported one is gone with the source that
  produced it

A collision between an interceptor-vended profile and an imported one still
fails closed, which is the behavior that shadowing quietly bypassed for
built-ins.

Operators upgrading should export the profiles their deployment relies on
before upgrading, or copy them from the providers/ directory of the matching
release tag, then import them at the scope their providers use.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* docs(providers): document import-only provider profiles

Rewrite the provider profile documentation around a catalog the operator builds
rather than one the gateway ships.

- gateway-config: the default is [{ type = "user" }], the builtin source type is
  gone and rejected at startup, and an empty catalog is a valid ready state
- profiles: replace the built-in profile table and its shadowing semantics with
  the import workflow, and say plainly that a profile whose binary paths do not
  match the image is inert
- manage-providers: replace the two fixed provider-type tables with
  `provider list-profiles`, and describe catalog-driven command inference
- inference-routing: drop the "still loads built-in profiles for compatibility"
  transition text; its migration walkthrough is now the normal path
- quickstart and the Docker Compose, GitHub, AWS, Google Cloud and Vertex
  tutorials: import the profile before creating the provider, since nothing
  resolves without it
- release notes: a 0.1.0 migration section covering export-before-upgrade,
  importing from the release tag's providers/ directory, what happens to a
  provider whose profile is missing, and the removal of the legacy type aliases
- README, architecture, the openshell-cli and debug-openshell-cluster skills,
  and the governance interceptor example follow the same change; the example's
  smoke assertion now imports a profile to prove an authoritative interceptor
  hides it

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(cli,e2e): distinguish catalog failures and authenticate OIDC seeding

Addresses two findings from review (GATOR-c1bd9867-02 and -03).

The provider profile catalog lookup collapsed its error into an empty
catalog, so an authorization, availability or transport failure on
ListProviderProfiles was indistinguishable from a gateway that genuinely has
no profiles. Command inference then resolved nothing and the sandbox was
created without the provider it needed, deferring the failure to the workload.

Keep the lookup's outcome instead of discarding it. The credential warning is
advisory and still degrades to its generic form, but inference now consults the
catalog only when there is a command to resolve and surfaces the lookup failure
when there is, naming the gateway and pointing at explicit --provider selection.
Two regression tests cover it through a fake gateway whose ListProviderProfiles
returns UNAVAILABLE while every other RPC succeeds: sandbox creation fails
without sending a provider-less create request, and a sandbox with no trailing
command still succeeds because it needs no catalog.

The e2e profile seeding also ran unauthenticated in the OIDC lanes. Those lanes
deliberately skip gateway registration and start the gateway without a TLS
client CA, so no mTLS identity exists and no token has been acquired when the
import runs; the wrapper exited during setup. Skipping the import is not
sufficient because the provider tests now require the claude-code profile.

Add e2e_register_oidc_admin_session, which mints an administrator token with
Keycloak's password grant — the same grant the OIDC test helpers use — and
writes the gateway metadata and token bundle that an interactive login would
have stored, so the import runs as an authenticated administrator. The mTLS
lanes keep the direct import unchanged. Both affected wrappers are covered:
with-podman-gateway.sh had the same defect as with-docker-gateway.sh.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* refactor(cli)!: remove command-derived provider attachment

Addresses GATOR-c1bd9867-01.

A profile's `binaries` list authorizes a binary to reach that profile's
endpoints. It is not a statement that running the binary asks for the provider,
and reading it as attachment intent let a command silently gain provider
authority: the aws-s3 example declares /bin/bash, so `sandbox create -- bash`
resolved an existing aws-s3 provider and attached its credential-backed
capability and network policy to the shell and its descendants. Attachment of
an already-created provider never prompted, so the confirmation flow did not
guard it.

The distinction that would make inference sound — whether a declared binary is
a profile's client or merely a permitted runtime — cannot be expressed:
NetworkBinary carries only a path, and the field that encoded it was removed in
0.1.0. Any substitute is a guess. Restricting the guess to a unique claimant
does not help, because uniqueness measures how sparse the catalog is rather
than what the user intended, and a compiled list of "generic" commands could
never cover an unbounded operator catalog.

Remove trailing-command inference rather than approximate it. A provider is
attached only when named with --provider, which still creates a missing
provider from local discovery when the name matches an imported profile ID.
Sandboxes with no providers remain a normal, fully supported state.

Removing inference also settles GATOR-c1bd9867-02: with no consumer deriving
authority from the catalog, its only remaining use is the advisory credential
warning, so a failed lookup degrades that warning instead of blocking creation.
The regression test now asserts that an unreachable catalog still creates the
sandbox and attaches nothing.

While repurposing the deduplication test, a pre-existing defect surfaced:
repeating a name in --provider auto-created it twice and the second attempt
failed with "provider already exists", because the explicit pass never
consulted the set of names it had already handled. Guard it.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(e2e): load the OIDC token and trust the gateway when seeding profiles

Addresses the carried finding GATOR-c1bd9867-03.

e2e_register_oidc_admin_session established a session the CLI never used. The
CLI decides whether to load a stored bearer token by matching on the auth_mode
field of the gateway metadata alone; the helper omitted that field, so the
metadata fell through to the default arm and oidc_token.json stayed on disk
unread. The profile import went out unauthenticated exactly as it had before
the helper existed. The helper also installed no trust anchor, leaving the CLI
unable to verify the gateway's self-signed serving certificate.

Write auth_mode = "oidc" so the stored token is loaded, and install the CA at
<gateway>/mtls/ca.crt. Only the CA is installed: with no client certificate or
key on disk the CLI falls back to CA-only server verification and authenticates
with the bearer token, which is what these lanes need because they start the
gateway without --tls-client-ca. Certificate verification stays on; no insecure
transport override is introduced.

Assert the session before anything depends on it. ListProviderProfiles is
annotated auth_mode: "bearer", so it cannot succeed unless the token was loaded
and accepted. The helper now fails at that point, naming the gateway config
directory and echoing the CLI output, rather than letting the defect surface
later as an opaque profile import error.

Both affected wrappers pass the PKI directory and CLI binary the helper needs;
the mTLS lanes are untouched.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(e2e): request the OpenShell scopes when minting the admin token

The OIDC lanes failed at the session assertion added in b25dd0298 with
PERMISSION_DENIED and "scope 'provider:read' required". Authentication was
working -- the gateway logged ListProviderProfiles at gRPC status 7, which it
can only reach once the bearer token has been loaded and accepted. The token
simply carried no OpenShell scope.

sandbox:*, provider:*, config:*, workspace:* and openshell:all are optional
client scopes on the openshell-cli client in scripts/keycloak-realm.json, so
Keycloak mints them only when the request asks for them. The password grant
here asked for nothing, leaving the realm defaults (openid, profile, email,
roles, web-origins, acr) and an access token that authorizes no RPC.

Request "openid openshell:all", as e2e/python/oidc/oidc_auth_test.py already
does for its administrator tokens. openshell:all is SCOPE_ALL in
crates/openshell-server/src/auth/authz.rs, so one scope covers the setup calls
without enumerating them. Verified against the realm: the token goes from
"email profile openid" to "email profile openshell:all openid".

Record the same scopes in metadata.json. The helper writes the bundle
openshell gateway login would have stored, and oidc_scopes is the field that
login path reads back, so leaving it out would misdescribe the stored token.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(e2e): seed provider profiles once per VM lane gateway

The VM lanes failed on the second test target with "custom provider profile
'anthropic' already exists" for all fifteen example profiles, and the import
exited non-zero.

e2e_import_example_provider_profiles was called from run_e2e_test, so it ran
once per target -- four times against one long-lived gateway. That was
harmless while import overwrote silently, but profiles are now import-only:
ImportProviderProfiles is create-only and reports an existing id as an
error-severity diagnostic, with no overwrite flag on the request. The first
import therefore succeeds and every later one fails.

Hoist the call to just after the conformance run, which is where the docker,
podman and kube lanes already seed their catalogs. The profiles persist for
the gateway's lifetime, so every target still finds them, and the
E2E_TEST_OVERRIDE path is covered by the same single call.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* test(e2e): name the provider explicitly in the OIDC workspace-user test

user_can_create_sandbox_with_inferred_provider_command reached the gateway
once the admin session was fixed, and then failed: it asserted a
missing-provider error, but the sandbox was created and died provisioning
with "failed to spawn sandbox entrypoint process 'claude-code'".

The test drove provider resolution by passing claude-code as the trailing
command and relying on the CLI to infer the provider type from it. That
inference is what this branch removed, so the trailing word is now nothing
but an entrypoint, and the image has no such binary.

The regression the test guards is not inference itself: it is that a
workspace user resolving a provider is not gated behind Platform Admin.
Name the provider with --provider, the only remaining way to attach one.
The CLI still has to fetch the claude-code profile before it can auto-create
the provider, so the lookup a workspace user must be allowed to make still
happens, and auto-creation still stops at the non-interactive branch with
"missing required provider". Both assertions therefore keep their meaning.

Rename the test and rework its comments to describe what it now exercises;
the old name would otherwise outlive the behavior it was named for.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(providers): reconcile import-only profiles with main

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: Philippe Martin <phmartin@redhat.com>
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-09-17 22:59:27 +00:00
Piotr Mlocek 769273f096 feat(middleware): add a hook to inspect HTTP responses (#3074)
* feat(middleware): implement HTTP response processing

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): allow one-byte response stream units

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* feat(examples): separate content guard from middleware protocol demos

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): address HTTP response review findings

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* refactor(examples): defer protocol demo to a separate PR

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* refactor(middleware): keep response body timeout local

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): honor fail-open for unrepresentable responses

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(middleware): trim runtime docs and extract troubleshooting reference

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* chore(middleware): split build fix and simplify test and skill guidance

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* refactor(middleware): isolate HTTP response processing

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): hide response credential headers

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): end invalid preflight streams

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(middleware): remove hard-wrapped prose

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): use protobuf request timeouts

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

---------

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-09-17 18:54:38 +00:00
John T. Myers 9b9f7905e1 fix(dev): extract Docker sandbox runtime from sandbox image (#3422)
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-09-17 18:21:11 +00:00
Johnny Greco 58b5f8f976 feat(prover): add standalone policy boundary checker (#3289)
* feat(prover): add standalone policy maximum checker

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* refactor(prover): simplify check scope schema

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(prover): align containment and cancellation with runtime

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(prover): stabilize containment checks in CI

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* test(prover): avoid solver in fast-path guard test

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(prover): align string containment with runtime

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(prover): reject ambiguous z3 string escapes

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(prover): align containment with runtime boundaries

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(prover-cli): harden cancellation and invalid input

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* feat(packaging): install policy prover with OpenShell

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(prover): clarify installation and check results

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(build): describe prover distribution directly

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* refactor(prover): rename maximum policy to boundary

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* refactor(prover): localize fail-closed validation

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* test(prover): cover fail-closed CLI surfaces

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* test(prover): allow CI load for REST solver proof

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(prover): use canonical policy schema for containment

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(prover): preserve uncertainty for runtime binary globs

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(prover): bound policy validation work

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(prover): document validation resource limits

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(prover): make containment API extensible

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(prover): define containment API contract

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(ci): integrate prover with consolidated builds

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* fix(ci): declare release packaging dependency

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

---------

Signed-off-by: Johnny Greco <jogreco@nvidia.com>
2026-09-17 17:30:03 +00:00
Florent BENOIT f419b9c1df fix(python): use portable empty-array expansion for macOS Bash 3.x (#3413)
macOS ships Bash 3.2 which errors on `"${arr[@]}"` when the array is
empty under `set -u`. Use `${arr[@]+"${arr[@]}"}` to safely expand
VERIFY_ARGS only when it has elements.

Signed-off-by: Florent Benoit <fbenoit@redhat.com>
2026-09-17 16:30:58 +00:00
krishicks 9c41f057c3 fix(python): stabilize development tasks under jj (#3354)
setuptools-scm can select jj's Git-bridge dev tag when uv resolves the local
package. Pin a development-only version for every Python task that performs
that resolution.

The 0.0.0 value is valid local editable-package metadata for development tasks;
it is neither a release version nor a Git tag. Wheel builds remain unpinned and
derive their published version from the actual release tag.

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-09-16 16:24:38 +00:00
Drew Newberry b3e4ad4579 fix(ci): repair RFC 0012 post-merge checks (#3360)
* fix(ci): lock nextest for Windows ARM64

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): preserve GPU access in sandbox runtime

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(windows): restore host proxy identity binding

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(windows): satisfy cross-platform network lint

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(windows): satisfy MXC lint

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(windows): exclude Unix supervisor runtime

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(mxc): satisfy Windows test lint

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(windows): escape proxy test policy paths

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): accept admitted GPU runtime groups

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(docker): serialize proxy pipeline scenarios

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-16 04:08:00 +00:00
Prekshi Vyas 314c73343c fix(ci): restore prebuilt Z3 on Windows (#3353)
Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>
2026-09-16 01:41:53 +00:00
Drew Newberry c1f2e7189f feat(isolation): implement the RFC 0012 sandbox architecture (#2942)
* feat(isolation): add RFC 0012 backend contract

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>

* refactor(isolation): name the interface crate explicitly

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): expose trusted host gateway

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(agents): inventory the MXC driver

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): add mediated DNS transport

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): tighten interface error and digest contracts

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(isolation): remove unrelated driver inventory

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): define capability-free launch contract

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): seal confirmed boundary state

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): validate confirmation for external backend implementations

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): clarify mediated DNS identity

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(isolation): generalize loopback connector

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(isolation): unify typed network mediation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): bind launches to sandbox sessions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(mxc): initialize extended sandbox status

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): add boundary protocol and Linux primitives

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): harden signals and separate process status from transport

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): validate remote confirmation through public contract

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): validate wire state and propagate snapshot failures

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(isolation): import owned agent specification explicitly

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(isolation): describe mediated DNS channel

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): bound mediation attach without nested retries

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(isolation): generalize loopback protocol

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): add transport-neutral session authentication

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(isolation): separate sandbox backend protocol

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): harden runtime boundary controls

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): add terminal boundary operation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): split supervisor and sandbox runtimes

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): harden boundary isolation and lifecycle ownership

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): reject private root redirects and adopt typed errors

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): preserve accept thread ownership on musl

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(sandbox): isolate credential probes from filtered threads

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): return retained exec exit status to independent waiters

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): bound network mediation and preserve socket authorization

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): bound control admission and retire stale mediation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci(e2e): select migrated drivers per stack layer

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(sandbox): implement loopback connector

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): authenticate the Sandbox Protocol

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(supervisor): rotate launch-scoped authentication

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(sandbox): consume dedicated backend crate

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(sandbox): align topology session fixture

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): align projected bootstrap bundle

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(auth): validate refreshed credentials before rotation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): fail closed across supervisor disconnects

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): repair rebased sandbox CI

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* build(runtime): publish separate sandbox and supervisor images

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(config): configure the sandbox runtime image

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(ci): validate sandbox binary linkage

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(isolation): use backend and runtime terminology

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(sandbox): use a scratch runtime image

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(ci): refresh schema and dependency policy

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): bind reconnects to supervisor process

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs: align runtime split operational guidance

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* chore(security): document Kubernetes runtime RBAC

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): enforce runtime lifecycle invariants

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(compute): identify sandbox start generations

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(server): restore sandbox launch sessions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): support authenticated runtime replacement

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(auth): bind sandbox session successors

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(auth): retry pending sandbox successors

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(vm): run the supervisor outside the guest workload

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(vm): use sandbox backend protocol

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): repair rebase integration

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): use unified build toolchain

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(vm): use sandbox runtime terminology

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(vm): own guest network bootstrap

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): expose guest init version

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): select native supervisor artifacts

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): guard guest init Linux symbols

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): scope Linux test imports

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): avoid guest interface casts

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): reconcile admitted sandbox identity

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): share resolved sandbox identity

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): surface host supervisor failures

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): include guest logs on supervisor exit

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): rotate and clean runtime generations

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): make sandbox starts generation-aware

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): rotate restored sandbox sessions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(vm): keep shared paths in the base layer

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): bind sandbox session lineage

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(docker): isolate workloads behind the host supervisor

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(docker): rotate launch-scoped authentication

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(docker): use sandbox backend protocol

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(docker): use host networking for supervisor

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): preserve host gateway alias resolution

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(docker): use separate sandbox and supervisor images

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): restore startup validation after rebase

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(docker): name the sandbox runtime directly

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): narrow supervisor CA runtime storage

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): close companion isolation gaps

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(docker): align mediated network expectations

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(docker): exercise mediated network paths

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): attach supervisor to managed network

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): defer supervisor recovery until gateway is ready

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): make sandbox starts generation-aware

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): rotate restored sandbox sessions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): preserve workloads during session rotation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(docker): remove unrelated configuration RFC changes

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): bind sandbox session lineage

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(kubernetes): add proxy-pod isolation topology

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(kubernetes): use sandbox backend protocol

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): use stable sandbox service authority

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(kubernetes): split sandbox and supervisor images

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): adapt proxy pods to current runtime APIs

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(kubernetes): describe the single runtime placement

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(kubernetes): simplify sandbox orchestration

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): validate deployment prerequisites

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): update Trivy Helm profile inventory

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(kubernetes): update Trivy scan inventory count

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): reuse preloaded runtime images in e2e

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): type and clean runtime resources

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): make sandbox restarts recoverable

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): rotate restored sandbox sessions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): preserve supervisor egress

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): bind sandbox session lineage

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(podman): adopt isolated sandbox and supervisor containers

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): stage bootstrap archives at named volume destinations

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(podman): rotate launch-scoped authentication

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(podman): use sandbox backend protocol

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(podman): use host networking for supervisor

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(podman): split sandbox and supervisor images

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): repair rebase integration

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(podman): name the sandbox runtime directly

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): provision supervisor CA runtime storage

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): address isolation review findings

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): inspect Debian supervisor provenance

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): use libpod-compatible tmpfs options

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): bind verified sandbox runtime binary

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): provide external driver data directory

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): start sandbox before joining user namespace

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): separate supervisor user namespace

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): make sandbox starts generation-aware

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): rotate restored sandbox sessions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): bind sandbox session lineage

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* perf(isolation): add TCP and DNS benchmark harnesses

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(perf): align benchmark timing and supported protocols

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(perf): report TCP benchmark metrics accurately

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(perf): cancel failed worker startup

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): build matching local supervisor image

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): make local sandbox smoke test runnable

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): wire local sandbox runtime image

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): narrow sandbox service RBAC

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(ci): validate split runtime artifacts

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): harden runtime session handling

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(supervisor): add standalone network proxy role

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(rfc): remove implementation companion notes

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(vm): standardize runtime release name

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): pin renamed runtime artifacts

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(auth): persist sandbox runtime identity

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(runtime): restore branch validation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): reconcile main after rebase

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(network): close unframed HTTP 1.0 responses

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* chore(isolation): preserve upstream OCSF updates

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(security): close credential and TLS replay paths

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(auth): make sandbox refresh retries idempotent

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-16 00:49:14 +00:00
Prekshi VyasandShailendra Singh d95bab5436 fix(ci): upstream Windows SDK validation support (#3327)
* fix(ci): upstream Windows SDK validation support

Port remaining Windows validation tooling from GitLab independently of the combined MXC runtime port. Preserve locked SDK dependencies and guard temporary launcher cleanup.

Co-authored-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* fix(ci): install native TypeScript dependencies before validation

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

---------

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>
Co-authored-by: Shailendra Singh <shailendras@nvidia.com>
2026-09-15 22:34:36 +00:00
Mrunal Patel b799fccb8b fix(auth): harden OIDC trust root retrieval (#3332)
* fix(auth): harden OIDC trust root retrieval

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(e2e): pass OIDC HTTP acknowledgement value

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

---------

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
2026-09-15 17:00:45 +00:00
26f2f96393 feat(mxc): add Windows host proxy for MXC sandbox network egress (#3163)
* Implement Windows host proxy integration and update dependencies for OpenShell

* Update README and gateway config to clarify egress proxy address handling and allocation

* Refactor ProxyIdentityMode to return Result for static_binary and add tests for binary path and SHA256 hash

* Enhance platform_hosts_path for Windows to use SystemRoot and improve error handling for hosts file reading

* Refactor FileFingerprint to use Option for mtime and ctime, simplifying metadata handling

* Add conditional compilation for Windows host module

* add unit tests for OPA policy evaluation and identity handling

* remove openshell-supervisor-network from unsupported driver package test exclusion list

* feat(mxc): enable host proxy TLS state generation

Generate per-sandbox TLS state for the MXC host proxy so HTTPS L7 enforcement can use the same MITM path as Linux. Grant generated CA material to the MXC process and inject standard trust env vars, while matching Linux behavior by disabling TLS termination on CA setup failure and relying on proxy fail-closed handling.

* fix(docs): remove outdated notes on governed egress from docs

* fix(tests): update TLS environment variable paths to use temporary directory

* fix(examples): make run-mxc-e2e harness correct and orphan-free

The MXC e2e harness never actually exercised the fs scenarios: it started
the gateway once and patched agent_command per scenario AFTERWARDS, so the
running gateway kept launching the default demo agent (not shipped in the
kit) and every fs scenario failed with CreateProcessW error:2. It also
scored on the `sandbox create` exit code (non-zero due to the harmless
interactive attach), wrote sandbox records to the persistent gateway DB
(leaving orphans that collided on later runs), and its deny scenarios never
proved denial.

Changes:
- Start a FRESH gateway per scenario so each scenario's agent_command is
  actually loaded (root cause of CreateProcessW error:2).
- Score by on-disk artifact / expected outcome, not `sandbox create` exit.
- Real deny assertions: a control write to a granted path must succeed
  (proves the agent ran) while the denied write must be absent. fs-empty
  probes an ungranted out-of-share path (share_dir is mapped rw by design).
- Run the gateway on an ephemeral in-memory DB (sqlite::memory:) so the
  harness never writes to the persistent store and cannot leave orphan
  sandbox records; also use unique per-run sandbox names + pre-delete.
- Fix the process_container probe: use a real cwd + absolute cmd.exe
  (canonical wxc-exec does not expand %TEMP% -> 0x8007010B).
- Fix summary counts (@() so a single FAIL is counted and exit is non-zero).

Verified PASS=4 FAIL=0 on 7F203-MXC-003 (no BaseContainer velocity keys)
using a canonical wxc-exec build (AppContainer fallback).

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(e2e): probe timeout is milliseconds (10ms->30000ms)

MXC process.timeout is wall-clock ms (wire.rs). The 10 value meant 10ms,
which the base-container tier (7F203-MXC-001/.181) enforced strictly and
timed the probe out. AppContainer path (.18/-003) happened to slip under
it. Bump to 30000ms so the process_container preflight probe is reliable
across both tiers.

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc): use native paths in real runtime probes

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc): make processcontainer work with mxc-latest-released wxc-exec

Three fixes to support the release wxc-exec binary (BaseContainer dispatcher)
in addition to mxc-fixes-env-vars:

1. Seed process env from host (driver.rs)
   ProcessContainer starts with a completely blank environment -- no PATH,
   SystemRoot, or anything.  Seed the process env from the gateway host
   environment so the agent binary can locate DLLs and run.  Skip internal
   Windows drive-letter variables (keys starting with '=') which cause
   CreateProcessW to return ERROR_ENVVAR_NOT_FOUND.  User agent_env entries
   and TLS CA vars are applied as overrides on top of the host env.

2. Remove TLS readonly_paths grant (driver.rs)
   The release wxc-exec (BaseContainer dispatcher) requires write-DAC
   permission on every path in readonly_paths to set up AppContainer ACLs.
   Adding the proxy's temp TLS directory caused a DACL error and exit -1.
   The CA cert paths remain available to the agent via TLS env vars.

3. Remove allowedHosts from network JSON (mxc.rs)
   The release wxc-exec rejects network.allowedHosts / network.blockedHosts
   on Windows with "not yet supported".  Removed the loopback exemption
   attempt (127.0.0.1, ::1, localhost) from the network section.
   Intra-container loopback works natively in the release binary without
   it -- the spawner can connect to the server at 127.0.0.1:22000 directly.

Additional changes:
- mxc-ws-agent.rs: add relay-debug.txt error capture and relay-ready.txt
  marker for reliable timing of host client connections.
- mxc-ws-gateway.toml: debug = true for JSON config dump during diagnosis.
- run-ws-agent-test.ps1: default port changed to 17670 (gateway default);
  relay-ready.txt polling before ws-echo to avoid connecting before the
  spawner has established the proxy bridge.

Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(e2e): address CodeRabbit review on run-mxc-e2e.ps1 (MR !46)

Four robustness/correctness fixes from CodeRabbit:

1. Start-Gw: kill the spawned gateway before the "did not start within 30s"
   throw. If the process is alive but never binds the port, $gw is not yet
   assigned in the caller, so the finally block cannot reap it -> orphan
   gateway holding the port for the next run.

2. create-fail scoring: a non-zero `sandbox create` exit alone is not proof
   of a policy rejection (gateway-registration/transport/fixture errors also
   exit non-zero and would false-pass). PASS now requires a genuine
   rejection signal (network / invalid_argument / network_policies) AND that
   it is not an infrastructure failure; other non-zero exits go to FAIL with
   output captured.

3. deny scenarios (ControlTarget path): snapshot the deny target AFTER
   Wait-File lands the control artifact, so a late denied write (enforcement
   regression racing the control write) can no longer be recorded as PASS.

4. -KeepRunning: break out of the scenario loop after the first scenario so
   a later scenario does not start a second gateway on the same port
   (previously a reliable port collision instead of a usable debug mode).

Re-verified PASS=4 FAIL=0 on both boxes (7F203-MXC-001 base-container and
7F203-MXC-003 AppContainer fallback); network-policy-rejected correctly
scores as "policy rejection".

Signed-off-by: Akber Raza <akberr@nvidia.com>

* feat(mxc-e2e): collect run-mxc-e2e output into a results bundle

Mirror the sibling run-*.ps1 scripts by collecting every run's logs into a
timestamped results-e2e-<stamp>\ folder and zipping it. The bundle contains the
console transcript, per-scenario gateway stdout/stderr, the exact TOML rendered
for each scenario, the policy fixture used, and a summary.txt with the verdict
table.

Per-scenario gateway logs now land in gateway.<scenario>.log/.err.log inside the
bundle instead of a single fixed gateway.e2e.log in the script directory.

Wrap pre-flight, mode setup, scenario definitions, and the scenario loop in a
single try/catch/finally so the finally always writes the summary, stops the
transcript, and zips the bundle -- even on a pre-flight failure. The existing
per-scenario gateway-cleanup try/finally stays nested inside. All scenario
logic, scoring rules, and comments are preserved.

Signed-off-by: Prashant S Khodade <pkhodade@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc-e2e): address CodeRabbit review on run-mxc-e2e.ps1

- Require -Scenario when -KeepRunning: the loop breaks after the first
  scenario, so a full-suite run would execute only one scenario yet still
  report the suite as PASS. Fail fast so a partial run can't be mislabeled
  complete.
- Start-Transcript now runs inside the guarded try block with a
  $transcriptStarted flag; Stop-Transcript is only called when it actually
  started, so a Start-Transcript failure still yields the results bundle.
- Wrap the -Scenario filter in @() so a single exact match stays an array
  (reliable .Count and a proper array for the scenario loop on PS 5.1).

Signed-off-by: Prashant S Khodade <pkhodade@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(examples): pass gateway config via OPENSHELL_GATEWAY_CONFIG for spaced paths

Start-Process -ArgumentList does not quote array elements, so launching the
gateway with a bare --config <path> token split on any space in the install
path (e.g. C:\Users\First Last\...), and clap rejected the fragment with
'unrecognized subcommand'. Every MXC example launcher that started the gateway
hit this when the kit was unzipped under a path containing a space.

Pass the config path through the OPENSHELL_GATEWAY_CONFIG env var (which the
gateway already reads via clap) and drop the --config token. Env vars carry
spaces safely.

Affected: run-ocsf-audit, run-mxc-e2e, run-demo, run-inference-test,
run-ollama-test. run-mtls-test was not affected (its launch passes no config
path). Root-caused and fix-verified on 7F203-MXC-003 from a spaced path.

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(run-mxc-e2e): improve scoring logic and enhance command execution handling

* fix(mxc): reconcile proxy support after rebase

Restore the proxy-enabled OCSF audit example removed by 13185f6e now that the host CONNECT proxy is present. Adapt the proxy lifecycle test to the target branch's DriverSandboxSpec policy delivery contract.

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc): grant sandbox access to proxy CA

- Share the per-sandbox public CA bundle with the AppContainer
- Add real wxc-exec HTTPS proxy coverage and document trust isolation

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc): reject unsupported network middleware

- Reject middleware-bearing MXC policies before sandbox lifecycle begins
- Guard host proxy startup and document the unsupported registry path
- Add mapper, lifecycle, and host proxy regression coverage

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc): reconcile host proxy with main

- remove obsolete inference routing from the host proxy adapter
- use the workspace AWS-LC provider in host-proxy tests
- adapt the forward-proxy test to ProxyIdentityMode

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(build): switch to bundled Z3 for Windows MSVC builds

* fix(mxc): reconcile host proxy after rebase

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Akber Raza <akberr@nvidia.com>
Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
Signed-off-by: Prashant S Khodade <pkhodade@nvidia.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Jamie King <jamiek@nvidia.com>
Co-authored-by: Shailendra Singh <shailendras@nvidia.com>
Co-authored-by: Prashant Khodade <pkhodade@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
2026-09-14 22:58:49 +00:00
krishicks 5b9daab935 fix(ci): restore mise run ci on macOS (#3294)
- Replace BSD-incompatible in-place sed calls with portable temp-file rewrites.
- Remove test-only shell interception and capture generated gateway config
  directly.
- Allow parity tests to use supplied supervisor binaries without resolving a
  Linux target.
- Normalize temporary-directory paths and use portable RPM config installation.
- Set a valid setuptools-scm version for Python protobuf generation in Jujutsu
  checkouts.

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-09-12 00:17:15 +00:00
Piotr Mlocek 5b57f0d154 fix(ci): restore Windows test portability (#3288)
* fix(ci): restore Windows test portability

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(tasks): skip Unix lockfile check on Windows

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(tasks): use buf shim on Windows

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

---------

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-09-11 21:40:08 +00:00
Piotr Mlocek b92620e838 ci(rust): reject stale Cargo lockfiles (#3227)
* ci(rust): reject stale Cargo lockfiles

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(ci): clarify lockfile validation policy

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* refactor(ci): structure and test Cargo lockfile validation

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* chore(ci): remove lockfile validator regression tests

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(rust): complete locked validation and lint examples

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(rust): check lockfile diffs after validation

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(build): shorten lockfile validation notes

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(rust): skip lockfile check after failures

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

---------

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-09-11 08:52:41 +00:00
Jesse JaggarsandDrew Newberry 02b664bb0d refactor(config): normalize and enforce gateway schema v2 (#2814)
* refactor(config): normalize compute driver field names

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* refactor(config): introduce canonical gateway fields

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* refactor(config): enforce gateway schema version 2

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): preserve compute driver runtime guarantees

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): address schema v2 review regressions

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): complete schema v2 migration safeguards

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(config): expand schema v2 regression coverage

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(config): add schema v2 parity manifest

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): correct parity manifest inventory

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* docs(config): record schema v2 intentional changes

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* docs(config): disposition schema v2 parity gaps

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): add dual schema parity harness

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): establish compute lifecycle parity baseline

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): preserve gateway option compatibility

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): record gateway option parity

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* docs(config): close gateway-wide parity gaps

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(podman): apply configured pids limit

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): validate Podman option parity

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): add Kubernetes option parity harness

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): record Kubernetes option parity

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): disposition VM parity lanes

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): add external driver parity lane

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(e2e): preserve external driver pull policy

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): attest parity artifacts and launches

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): require clean parity build sources

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): bind parity runtime artifacts

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(e2e): use isolated supervisor tags

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(e2e): qualify parity image tags

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(e2e): serve parity supervisor locally

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): isolate parity podman services

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): harden parity evidence provenance

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): pin parity sandbox artifacts

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): attest parity runtime inputs

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): bind parity runtime evidence

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): record compute boundary parity

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): disposition cross-cutting parity lanes

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(packaging): preflight gateway config upgrades

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): preserve rebase integration guarantees

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(ci): isolate temporary git signing config

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): update remaining schema v2 consumers

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(ci): provide e2fs tools to VM tests

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): align preflight with gateway startup

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(vm): preserve rootfs tar configuration

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* chore(config): adopt duration unit constructors

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(packaging): preflight RPM gateway config

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): address driver review findings

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(e2e): require fresh semantic parity evidence

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(docker): update tests for renamed sandbox label

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(gateway): preserve selective driver coverage after rebase

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
2026-09-11 05:00:24 +00:00
Drew Newberry 38f2aef930 feat(gateway): support selective compute driver builds (#3118)
* feat(gateway): support selective compute driver builds

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(gateway): support selective Windows MXC builds

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-11 00:36:09 +00:00
Piotr Mlocek ddc8bba967 ci(windows): add Windows MSVC CI jobs (#2738)
* fix(ci): preserve Windows Rust build cache

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(ci): invalidate empty Windows caches

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* perf(ci): cache Windows builds with sccache

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(windows): restore target directory caching

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* perf(ci): use prebuilt Z3 on Windows

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* perf(ci): layer sccache on Windows target cache

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(windows): split PR checks from main validation

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(windows): separate checks builds and cache seeding

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(ci): simplify Windows build dependency

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(ci): rely on Windows job dependency status

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(ci): use valid opt-in Windows ARM runner

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(windows): keep ARM64 validation local

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(ci): install Clippy for Windows validation

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(windows): focus platform lint coverage

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(licenses): explain bzip2 allowance

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(windows): simplify workflow name

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(windows): allow async platform stub

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(network): make file fingerprints portable

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(windows): lint supported deliverables

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(windows): allow platform-gated lint

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* chore(ci): align Windows cache action with main

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(ci): align Windows validation with prerequisites

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(windows): pin Rust toolchain action

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(ci): use enterprise-approved Windows actions

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(windows): restore strict MSVC validation

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(windows): run Rust tests with nextest

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(ci): normalize nextest lock provenance

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(windows): add native arm64 validation

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(ci): lock nextest for Windows ARM64

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(windows): resolve duplicate MXC authentication method

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* test(conformance): use native absolute paths on Windows

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(windows): address MSVC review feedback

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(windows): simplify cache key names

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(ci): isolate Windows Rust toolchains for stable caches

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(ci): configure Rustup home in runner setup

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(windows): surface sccache server write diagnostics

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(windows): remove temporary cache diagnostics

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* ci(windows): address review feedback

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(windows): reconcile merged driver capabilities

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(deps): preserve AWS-LC-only lockfile

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* chore(deps): allow z3 prebuilt TLS wrapper

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

---------

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-09-10 23:53:01 +00:00
Piotr Mlocek ce25acca5a fix(build): honor Cargo target directory when staging binaries (#3262)
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-09-10 23:19:13 +00:00
Piotr Mlocek a0814443f1 feat(docs): publish versioned release snapshots (#3149)
* feat(docs): add version availability labels

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(docs): use supported Python for sync

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* feat(docs): publish versioned docs from releases

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(docs): format dev version label

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* chore(docs): upgrade Fern CLI to 5.112.0

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(docs): make release publishing monotonic

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(docs): preserve snapshot release identity

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(fern): document versioned publishing

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* test(docs): cover explicit snapshot rollback

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* refactor(docs): use Fern refs for versions

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(docs): bundle components for ref versions

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* revert(docs): keep complete version copies

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(docs): publish latest and dev channels

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

---------

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-09-09 23:34:34 +00:00
John T. Myers f4dc6be4b2 refactor(inference): remove managed inference routes (#3195)
* refactor(inference): remove managed inference routes

Closes #3172

Remove the inference route control plane, inference.local data path, built-in router crate, and SDK surface. Move inference workloads to explicitly imported provider profiles and native endpoints, with migration cleanup and updated tests and documentation.

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(policy): preserve alternate upstream isolation

Restore the provider policy activation guard so legacy OpenAI and Anthropic providers configured for alternate base URLs do not grant egress to the built-in public vendor endpoints.

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-09-09 18:47:22 +00:00
Gaizka Menendez 8af79a7f4b fix(podman): resolve macOS Podman socket dynamically (#3135)
* docs(podman): document macOS socket path mismatch and dynamic lookup

On macOS, Homebrew-installed Podman does not create the default socket
path that the Podman driver probes. Document the OPENSHELL_PODMAN_SOCKET
override and the podman machine inspect lookup in both the compute
drivers reference and the debug-openshell-cluster skill.

Fixes #1690

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* fix(podman): resolve macOS Podman socket dynamically

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* chore: restore debug skill file

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* chore: drop legacy debug skill path

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* fix(podman): trim unrelated e2e changes

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* fix(e2e): harden shell array expansion

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

* chore: remove unrelated skill note

* ci: retrigger checks

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>

---------

Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com>
2026-09-09 14:10:59 +00:00
alangou 3693b32841 ci(trivy): add artifact and PR configuration scans (#3185)
* ci(trivy): add artifact and PR configuration scans

Signed-off-by: Adrien Langou <alangou@nvidia.com>

* fix(ci): harden Trivy gate detection and finding diff

Signed-off-by: Adrien Langou <alangou@nvidia.com>

* feat(ci): scan released artifacts in release pipelines

Signed-off-by: Adrien Langou <alangou@nvidia.com>

* fix(ci): harden and simplify Trivy scans

Signed-off-by: Adrien Langou <alangou@nvidia.com>

* fix(ci): consolidate Trivy reports and prevent collisions

Signed-off-by: Adrien Langou <alangou@nvidia.com>

---------

Signed-off-by: Adrien Langou <alangou@nvidia.com>
2026-09-09 13:58:06 +00:00
Jorge 519e5eb35f feat(e2e): make e2e:kubernetes work transparently on OpenShift (#3183)
* feat(e2e): make e2e:kubernetes work transparently on OpenShift

Running `mise run e2e:kubernetes` on OpenShift required manual namespace
creation, SCC grants, Helm value overrides, and cleanup. A separate
`e2e:openshift` task existed but only checked pod readiness without
running the Rust e2e test suite, and even with the suite wired up the
SSH-relay `sandbox connect` path stalled to the ready timeout because
`kubectl port-forward` cannot carry round-trip-heavy SSH over the
internet.

The harness now auto-detects OpenShift via the `route.openshift.io` API
group and, on OpenShift, both configures the cluster and switches the
gateway transport automatically:

- Drives the gateway through a passthrough OpenShift Route secured with
  mandatory mTLS instead of port-forward, so the connect suites
  (live_policy_update, port_forward, sync, connect-based
  sandbox_lifecycle, settings_management) actually pass. Computes the
  Route host from the cluster ingress domain, extracts client mTLS
  material from the openshell-client-tls secret, waits for the Route to
  serve mTLS, asserts a certless caller is rejected at the TLS
  handshake, and registers an mTLS CLI gateway pointing at the Route.
- Applies an SCC-compatible Helm values overlay that removes hardcoded
  runAsUser/fsGroup, letting OpenShift assign UIDs from the namespace
  range.
- Grants the privileged SCC to openshell-sandbox before Helm install
  and removes it during cleanup.
- Grants the anyuid SCC to the PostgreSQL fixture service account in
  DB scenarios and removes it during cleanup.
- All oc commands use --context to target the correct cluster.

The OpenShift e2e overlay (ci/values-openshift-e2e.yaml) turns TLS back
on, enables the Route, promotes the cert-verified caller to a dev
principal, and forces `image.pullPolicy`/`supervisor.image.pullPolicy`
to Always so runs against the `latest` upstream image use it instead of
a stale copy cached on the cluster nodes. Every OpenShift branch is
gated on OPENSHIFT_DETECTED, so the vanilla-Kubernetes port-forward path
is unchanged.

The Helm template for podSecurityContext is wrapped with {{- with }} so
null values omit the block instead of rendering invalid YAML.

The separate e2e:openshift task and e2e-openshift.sh script are removed
since e2e:kubernetes now covers OpenShift.

TESTING.md is updated with Kubernetes e2e documentation including
OpenShift auto-detection, dropping the e2e-host-gateway feature on
remote clusters, pinning IMAGE_TAG when the CLI and image versions
differ, task variants, and environment variables.

The debug-openshell-cluster skill gains an OpenShift platform row and
two SCC failure patterns (gateway rejected over hardcoded runAsUser,
sandbox missing the privileged SCC) covering the SCC handling and
podSecurityContext behavior this change introduces.

Signed-off-by: Jorge Garcia Oncins <jgarciao@redhat.com>

* fix(e2e): harden OpenShift SCC cleanup, mTLS gate, and Route timeout

Track the anyuid SCC grant for the PostgreSQL fixture with a dedicated
OPENSHIFT_POSTGRES_SCC_GRANTED flag set before the fixture apply, so a
failed apply no longer leaks the binding; cleanup now revokes it whenever
the grant succeeded, independent of deploy state.

Validate the Route server cert in the certless security gate (curl
--cacert instead of -k) and classify curl's exit code so only a TLS
client-auth rejection (35/56) counts as the expected certless rejection;
an unrelated DNS/timeout/TLS failure now fails loudly instead of masking
a potential mTLS hole.

Raise the OpenShift Route timeout in the e2e overlay. The default HAProxy
Route timeout is 30s, which severed long-lived transfers (large sandbox
upload/download, SSH-relay `sandbox connect`) mid-stream and failed the
sync e2e tests. Set both haproxy.router.openshift.io/timeout and
timeout-tunnel to 300s: a passthrough Route proxies in TCP mode, so
timeout-tunnel governs the established tunnel while timeout covers the
pre-tunnel phase.

Document the OpenShift transport exception, oc prerequisites and SCC
grants, and make the skopeo tag-check example copy-safe in TESTING.md.

Signed-off-by: Jorge Garcia Oncins <jgarciao@redhat.com>

---------

Signed-off-by: Jorge Garcia Oncins <jgarciao@redhat.com>
2026-09-08 20:37:19 +00:00
Piotr Mlocek 039b265096 feat(middleware): define HTTP response pre-return interface (#3073)
* feat(middleware): define HTTP response pre-return interface

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(middleware): clarify HTTP response interface

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* refactor(middleware): align response result actions

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* refactor(middleware): expose response reason codes

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* refactor(middleware): share session end reasons

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* feat(middleware)!: finalize HTTP response pre-return contract

Replace the separate body_end event with HttpResponseBodyUnit.end_of_stream.
Every body-inspecting stage receives exactly one flagged unit, which may be
empty; a zero-byte body is one empty flagged unit and OpenShell never reads
ahead to set the flag.

Defer response trailers from V1 and reserve their field numbers. HTTP/1.0
clients and Content-Length bodies cannot carry trailers and that behavior was
undefined.

Add HttpResponsePreflight.permitted_body_modes, computed once from the
original upstream head so every stage sees the same list, and make an
unlisted selection a failure rather than a downgrade. Add the block_delivery
preflight action as a successful decision enforced regardless of on_error.

Expose Content-Length, Content-Encoding, and Content-Range read-only in
preflight. Cap STREAM_BYTES input units at half of max_payload_bytes and
permit deferring bytes across replacements only for fail_closed bindings,
surfaced as deferral_permitted.

Split PEER_DISCONNECT into DOWNSTREAM_DISCONNECT and UPSTREAM_DISCONNECT and
attribute WebSocket relay failures by direction instead of a generic peer
error. Compile the content-guard example in lint and branch checks so proto
renames cannot break it silently.

BREAKING CHANGE: WebSocketSessionEndReason and WebSocketSessionEnd are
replaced by the shared MiddlewareSessionEndReason and MiddlewareSessionEnd.
NORMAL_CLOSE is now NORMAL, UPSTREAM_REJECTED is now UPSTREAM_FAILURE, and
PEER_DISCONNECT is split into DOWNSTREAM_DISCONNECT and UPSTREAM_DISCONNECT.
Enum numbers are unchanged so binary wire compatibility is preserved;
generated symbols and JSON names change.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(middleware): describe skip as opting out of inspection

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* feat(middleware): add body-phase block_delivery and skip_remaining actions

Body results may now stop delivery or opt out of inspecting the rest of the
response after a prefix. One HttpResponseBlockDelivery message is shared by
preflight and body results and documents the difference between blocking
before and after head commitment. Drop the field reservations, since nothing
in this contract has shipped, and renumber session_end to close the gap.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* refactor(middleware): share HTTP body leaf messages across directions

HttpBodyUnit, HttpBodyPassThrough, HttpBodyTransform, HttpBodySkipRemaining,
and HttpBodyMode carry no response-specific semantics, so name them for reuse
by the streaming request hook. Envelopes, results, preflight, and
block_delivery stay response-specific because commitment semantics differ.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* refactor(middleware): keep HTTP body leaf messages response-specific

Reverts the shared HttpBody* naming. A direction-specific payload such as a
response-only semantic mode would otherwise add unreachable variants to the
other direction or force a source-breaking fork after 0.1.0. The streaming
request hook defines its own HttpRequestBody* messages and copies the shape;
SDKs present a direction-neutral body handler over both.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(middleware): simplify response proto comments

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): reject undispatched response bindings

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(middleware): simplify phase field comment

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* refactor(middleware): rename HTTP response preflight result

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* feat(middleware): add HTTP response trailer results

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(middleware): define response block delivery

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(middleware): define stage-local response body modes

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(middleware): define final response body units

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(middleware): define streaming response deferral

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(middleware): defer whole-body accumulation timeout

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(middleware): align response result diagnostics

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* test(middleware): cover upstream WebSocket disconnect

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(middleware): keep response streams unit-local

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(middleware): trim disconnect compatibility note

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

---------

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-09-04 21:18:02 +00:00
Piotr Mlocek d7cb6e456d fix(dev): inherit non-expiring sandbox JWT in local gateway scripts (#2636)
* fix(dev): inherit non-expiring sandbox JWT in local gateway scripts

The local gateway launcher scripts hardcode gateway_jwt.ttl_secs = 3600,
which overrides the non-expiring default introduced in #1721. Local
Docker, Podman, and VM sandboxes are still unrecoverable when the gateway
is down longer than that TTL: the on-disk token expires and only the
Kubernetes ServiceAccount path can rebootstrap, so the supervisor
crash-loops on policy fetch and the sandbox never leaves Provisioning.

Drop the override so local drivers inherit the default. gateway.sh also
serves the kubernetes driver, which is a shared deployment and must keep
a positive TTL, so it now emits ttl_secs only for that driver.

The e2e regression test added in #1721 does not catch this because the
e2e harness uses its own configs, which already set ttl_secs = 0.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(dev): expand sandbox JWT TTL in gateway config

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

---------

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-09-04 17:16:44 +00:00
alangou 64a858dade fix(ci): restore Codex Security scan execution (#3124)
* refactor(ci): resolve Codex Security range in Python

Signed-off-by: Adrien Langou <alangou@nvidia.com>

* fix(ci): allow unprivileged userns for Codex sandbox

Signed-off-by: Adrien Langou <alangou@nvidia.com>

---------

Signed-off-by: Adrien Langou <alangou@nvidia.com>
2026-09-03 07:38:51 +00:00
Dhiraj Bokde cc4ded2088 feat(helm): split gateway and workspace charts (#2643)
* feat(helm): split gateway and workspace charts

Signed-off-by: Dhiraj Bokde <dbokde@nvidia.com>

* fix(helm): preserve split chart upgrade compatibility

Keep workspace manifests valid after value validation and default legacy reused values to the combined resource topology.

* fix(ci): preserve VM runtime for E2E

The Rust cache restores target/ after VM runtime artifacts are staged,
overwriting target/vm-runtime-compressed before openshell-driver-vm is built.
Stage the compressed runtime outside target and pass that location through
OPENSHELL_VM_RUNTIME_COMPRESSED_DIR so build.rs can embed the supervisor.

Also locate the Helm split-ownership test repository root from the script
path rather than git rev-parse. The test runs in a container where the
GitHub checkout can be owned by a different UID and rejected as dubious
ownership.

Signed-off-by: Dhiraj Bokde <dbokde@nvidia.com>

* fix(ci): install yq for Helm ownership test

The split-chart ownership regression uses yq to inspect rendered YAML,
but the Helm CI container installs only tools declared in mise.
Declare and lock yq so mise install --locked provides the test dependency.

Signed-off-by: Dhiraj Bokde <dbokde@nvidia.com>

---------

Signed-off-by: Dhiraj Bokde <dbokde@nvidia.com>
2026-09-02 00:23:39 +00:00
Drew Newberry 9ca19e6c80 refactor(compute): decouple gateway driver composition (#2823)
* refactor(compute): decouple gateway driver composition

Move first-party composition and VM process ownership into openshell-gateway, leaving openshell-server backend-independent. Update packaging and build references with the new crate, simplify the compiled-driver boundary, and keep the driver-free gateway path buildable with bundled Z3 tooling.

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(telemetry): bound compute driver categories

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(core): keep runtime transport generic

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(compute): complete server driver decoupling

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(compute): preserve driver integrations after rebase

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>

* fix(compute): preserve docker tracing after decoupling

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>

* fix(compute): preserve driver behavior after extraction

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(compute): remove MXC policy side channel

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(compute): separate policy delivery from readiness

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>
2026-09-01 21:13:45 +00:00
Russell BryantandJohn Myers 5b925dd8af feat(build): add defaults-without-telemetry feature alias (#2843)
* feat(build): add defaults-without-telemetry feature alias

Cargo cannot subtract a single default feature, so compiling telemetry out
meant `--no-default-features` plus a hand-maintained keep-list of the crate's
other defaults. That keep-list was already wrong for operators: telemetry is
the only default on openshell-server and openshell-driver-vm, but
openshell-sandbox also defaults to `bundled-ca-roots`, so a bare
`--no-default-features` silently swapped the supervisor onto the platform
trust store.

Add a `defaults-without-telemetry` alias to each of the three telemetry-
carrying binary crates, enumerating every default except `telemetry`.
Telemetry-free builds become `--no-default-features --features
defaults-without-telemetry` and stay correct as the default set grows.

The alias is a keep-list, not a switch. Enabling it on top of the defaults
would otherwise produce a telemetry-on binary that reads as telemetry-free, so
each crate root carries a `compile_error!` for the `telemetry` +
`defaults-without-telemetry` combination.

Add `rust:verify:defaults-without-telemetry` to guard both properties: each
alias still equals its crate's defaults minus `telemetry`, and the
mutual-exclusion error is wired up. The additive-misuse check matches on the
`compile_error!` text rather than a nonzero exit code so it cannot pass
vacuously on hosts where openshell-driver-vm fails to build for unrelated
reasons. `rust:verify:telemetry-off` now builds through the alias.

Signed-off-by: Russell Bryant <rbryant@redhat.com>

* fix feature alias for openshell-server

Signed-off-by: Russell Bryant <rbryant@redhat.com>

* fix(ci): run Rust verification in Nix shell

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: Russell Bryant <rbryant@redhat.com>
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-09-01 19:55:42 +00:00
Simon Scatton a4f9c762ce fix(release): handle prerelease tag builds (#3094)
Signed-off-by: Simon Scatton <sscatton@nvidia.com>
2026-09-01 16:12:01 +00:00
Simon Scatton c8f13205e3 ci(release): publish prerelease artifacts (#3093)
Signed-off-by: Simon Scatton <sscatton@nvidia.com>
2026-09-01 15:31:18 +00:00
alangou f7180c0fd6 feat(ci): add Codex Security release qualification (#3087)
* feat(ci): add Codex Security release qualification

Scan cumulative release-train diffs through NVIDIA inference and publish findings to Code Scanning.

Signed-off-by: alangou <alangou@nvidia.com>

* fix(ci): disable package cache for security scan

Prevent cache poisoning in the tag-triggered Codex Security workflow.

Signed-off-by: alangou <alangou@nvidia.com>

* refactor(ci): simplify Codex Security reporting

Remove custom inference cost accounting so the workflow remains focused on scanning and SARIF publication.

Signed-off-by: alangou <alangou@nvidia.com>

---------

Signed-off-by: alangou <alangou@nvidia.com>
2026-09-01 13:57:14 +00:00
Evan Lezar bb70461878 test(e2e): run conformance in gateway lanes (#2925)
* test(e2e): isolate VM-specific smoke assertions

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(e2e): add portable CLI conformance baseline

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* feat(conformance): add standalone CLI runner

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(e2e): run conformance in gateway lanes

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-09-01 12:36:51 +00:00
krishicks 197b41371d fix(dev): harden local cluster and gateway startup (#2993)
* fix(helm): refresh kubeconfig for existing k3d clusters

Docker can recreate the k3d load balancer on a new API port. Start existing
clusters and prefer fresh k3d entries so create does not retain a stale
endpoint.

Signed-off-by: Kris Hicks <khicks@nvidia.com>

* fix(dev): conditionally enable local OTLP export

Probe port 4317 before adding OTLP configuration for the VM, Docker,
and Podman gateway tasks. Document the startup behavior and troubleshooting
for local collector availability.

Signed-off-by: Kris Hicks <khicks@nvidia.com>

---------

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-08-27 17:55:01 +00:00
krishicks f68867b869 feat(gateway): identify gateways in exported traces (#2647)
* feat(gateway): add installation name configuration

Add a first-class operator-assigned gateway name with TOML, CLI, environment,
and Helm configuration surfaces. Local gateways default to openshell, while
Helm defaults to the chart fullname; operators sharing a collector across
namespaces or clusters can set a globally distinct name.

Signed-off-by: Kris Hicks <khicks@nvidia.com>

* feat(gateway): identify gateways in exported traces

Attach the configured gateway installation name and compute driver to the
gateway OpenTelemetry resource so operators can filter traces from multiple
installations that share a collector. Forward the gateway name and OTLP
endpoint to managed external drivers so their distinct service resources carry
the same installation identity.

Keep service.name stable per process type, omit blank resource values, and
leave per-span operation names and request attributes unchanged.

Refs #2507

Signed-off-by: Kris Hicks <khicks@nvidia.com>

---------

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-08-27 17:27:34 +00:00
alangou 37072ee81c feat(build): publish OCI SBOM and provenance attestations (#2836)
* feat(build): embed auditable Rust dependency metadata

Signed-off-by: Adrien Langou <alangou@nvidia.com>

* feat(build): publish OCI SBOM and provenance attestations

Signed-off-by: Adrien Langou <alangou@nvidia.com>

---------

Signed-off-by: Adrien Langou <alangou@nvidia.com>
2026-08-27 14:53:33 +00:00
Evan Lezar 9f88f8ff9b ci: remove rootless podman e2e lane (#2981)
Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-08-27 08:48:44 +00:00
bcd517bbe0 feat(driver-mxc): native Windows MXC compute driver + server wiring (#2721)
* feat(driver): add MXC compute driver for Windows isolation sessions

Introduces the openshell-driver-mxc crate implementing ComputeDriver
backed by Microsoft MXC isolation sessions (Windows only). Wires the
new driver into the server's build_compute_runtime dispatch and adds
the Mxc variant to ComputeDriverKind.

Also adds a local protobuf-src stub (tools/protobuf-src-local) to
unblock Windows builds that lack MSYS2/MinGW, and pins the zig
Windows x64 toolchain in mise.lock.

(cherry picked from commit 4f7012224efb18fbfeb47aa87e0cfd3f036f32f0)
Signed-off-by: Jamie King <jamiek@nvidia.com>

* wip(mxc): checkpoint hung-agent work (recon, policy_map embed, A1 wiring, demo artifacts)

Safety checkpoint of uncommitted work from the background agent run that stalled mid-Step-7. Includes: mxc-driver-recon.md (Step 0.5), policy_map.rs (~876L embedded mapper), A1 policy-threading edits across driver.rs/policy.rs/mxc.rs/compute/mod.rs, and examples/ (demo.yaml + mxc-gateway.toml). Not yet verified to compile end-to-end; to be reorganized into the skill's Step 11 commit sequence.

(cherry picked from commit 38e42c03870be3d10e984a54f17a3b61122ff510)
Signed-off-by: Jamie King <jamiek@nvidia.com>

* test(mxc): fix lifecycle and policy unit-test compile drift

- Bring futures::StreamExt into scope for the watch-stream `.next()` call in
  driver::lifecycle_tests so the negative policy proof test compiles.
- Bind a local `mapper` and drop the unused/deprecated NetworkBinary in the
  embedded-mapper network-policy rejection test.

Signed-off-by: Jamie King <jamiek@nvidia.com>
(cherry picked from commit 039b0baf98735ca672dae52be8d3af2417dc0c1a)
Signed-off-by: Jamie King <jamiek@nvidia.com>

* fix(mxc): downgrade missing sandbox_token to debug log

The gateway mints `sandbox_token` only when a sandbox-JWT issuer is
configured. There is no in-sandbox supervisor on MXC (supervisor-removal
design — D1/D4), so no component ever consumes the token; requiring it
on the driver side blocks the demo's `--disable-tls` smoke gateway with a
spurious `invalid_argument`. Log the absence and proceed instead.

Signed-off-by: Jamie King <jamiek@nvidia.com>
(cherry picked from commit cea209797d0edcb1d152251e748900b0a63cca62)
Signed-off-by: Jamie King <jamiek@nvidia.com>

* fix(mxc): keep sandbox Ready after a successful one-shot agent exec

monitor_exec demoted Ready->Error on exit 0 (reason ExecCompleted), so the positive demo (write hello.txt + exit) landed in Error phase. Keep Ready=True (reason AgentCompleted) on success; only non-zero exits go to ExecFailed. Tighten the positive lifecycle test to assert the terminal condition stays Ready=True/AgentCompleted. Verified live via gateway mock round-trip: phase now Provisioning->Ready with no demotion.

(cherry picked from commit 54ab030f03ca0f83d0050d8dd843b633651684ad)
Signed-off-by: Jamie King <jamiek@nvidia.com>

* feat(mxc): add processContainer backend for default-deny enforcement

Add a backend selector to the MXC driver (isolation_session default | process_container). process_container drives a one-shot AppContainer that is genuinely default-deny: a write to any ungranted path is denied by the OS, unlike isolation_session which is grant-only and cannot deny. The lifecycle forks on the flag - isolation_session keeps provision/start/exec, process_container runs a single ephemeral container via run_oneshot.

Also: run-demo.ps1 gains -Backend and hardens the CLI register/create calls; docs corrected to state isolation_session does NOT deny out-of-policy writes and that the negative proof requires process_container.

Verified end-to-end on a real demo box (gateway -> CLI -> driver -> MXC): in-policy write succeeds, out-of-policy write denied (PermissionDenied), OVERALL: PASS.

(cherry picked from commit c6cde3860bbe1b8edb3147d3e840f6bf0ece32d8)
Signed-off-by: Jamie King <jamiek@nvidia.com>

* refactor(driver-mxc): embed policy mapper as a module; remove standalone crate

Adopt the proto-based mapper (map_to_mxc) as the single source of truth,
embedded in openshell-driver-mxc as a Windows-gated `policy_map` module.
Rewire EmbeddedPolicyMapper to call it directly on the typed SandboxPolicy,
deleting the serde_yaml proto->YAML bridge. Move the CLI to a windows-gated
example and the parity tests into the crate; delete openshell-policy-mapper.

- gate policy_map + seam Windows-only (MXC is Windows-only)
- drop serde_yaml; add dev-deps openshell-policy, clap, anyhow
- normalize mapped paths to Windows form in the seam, in one place
- docs: add driver-mxc to AGENTS.md table; correct design doc section 17 test lane

Signed-off-by: Giedrius Burachas <gburachas@nvidia.com>
(cherry picked from commit f22f9c7a25b9a651c5c5cc73f62fb01c4d6c1a8d)
Signed-off-by: Jamie King <jamiek@nvidia.com>

* feat(driver-mxc): implement lossless split_policy for proxy-delegated egress

Signed-off-by: Giedrius Burachas <gburachas@nvidia.com>
(cherry picked from commit 96d6afa0e2dc6a1d54edd12c34a0ceb0a30dadd0)
Signed-off-by: Jamie King <jamiek@nvidia.com>

* feat(driver-mxc): implement Pattern-C governed-egress split through the policy seam

- split_policy: SocketAddr proxy_redirect (replaces bare port), processcontainer
  containment guard naming MXC M1, version preserved in the trimmed proxy_policy,
  delegation reported as an info loss item
- seam: MappedConfig carries trimmed_policy + proxy_addr; MapCtx.egress selects
  the split path; coarse path unchanged when egress is disabled
- driver: [openshell.drivers.mxc] egress_proxy / egress_proxy_addr config,
  validated at create (isolation_session rejected until M1); lifecycle threads
  the redirect into provision and stores the trimmed policy per sandbox,
  emitting an EgressRedirect platform event
- mxc: optional MxcNetwork block (defaultPolicy=block + proxy) in provision and
  one-shot configs; mock records configs for test assertions
- tests: lossless-invariant suite over all example policies (validate +
  serialize round-trip), split lifecycle proof, M1 rejection; example gains
  --split --proxy-addr writing mxc-config.json / trimmed-policy.yaml /
  loss-report.json

Signed-off-by: Giedrius Burachas <gburachas@nvidia.com>
(cherry picked from commit 34d54ad9f25dc6034c3ba15668555ff0d22cddd8)
Signed-off-by: Jamie King <jamiek@nvidia.com>

* fix(driver-mxc): emit MXC network.proxy as {localhost: port}

Verified against the real wxc-exec 0.6.0-alpha via --dry-run: MXC accepts
only the {localhost: N} proxy shape (the form the design doc specifies)
and rejects {host, port} with a parse error. Schema 0.6.0-alpha can
express only a loopback port, so non-127.0.0.1 redirect addresses are now
rejected: split_policy emits an error loss (no proxy block) and the driver
refuses egress_proxy_addr values off 127.0.0.1. Per-sandbox attribution
must use per-sandbox ports until the schema widens.

Signed-off-by: Giedrius Burachas <gburachas@nvidia.com>
(cherry picked from commit edde8d5434571fd5398409204fcf6862672c0793)
Signed-off-by: Jamie King <jamiek@nvidia.com>

* fix(driver-mxc): serialize isolation_session stop/deprovision as unit variants

Empirical contract finding from the real test lane (build 26300.8553,
wxc-exec 2026-06-10): the stop and deprovision experimental blocks are
unit variants in the wxc-exec schema and must serialize as null; sending
{} is rejected with malformed_request (invalid type: map, expected unit),
while provision/start accept maps. The production invoker, the real-lane
test, the probe script, and the e2e runner all sent {} - the driver could
provision and run an agent but never stop or delete an isolation-session
sandbox against this build. Pinned by a unit test.

Signed-off-by: Giedrius Burachas <gburachas@nvidia.com>
(cherry picked from commit 0df39ca0b22ebc21eb965b2b567a5b2cac26af32)
Signed-off-by: Jamie King <jamiek@nvidia.com>

* test(driver-mxc): add Tier-0 mapper coverage matrix with schema drift guard

Three-quadrant, table-driven matrix (38 tests): mappable fields assert
exact MXC output; every OpenShell field MXC cannot express asserts a loss
item with the expected severity (and seam rejection on error); an empty
policy asserts the restrictive default-deny posture for every MXC knob
OpenShell does not control. The handled_fields_inventory drift guard
serializes a fully-populated policy and compares its YAML keys against
the mapper-handled field lists, so a new openshell-policy field fails the
suite until consciously mapped, delegated, or reported as loss.

Re-exports the policy seam types for integration tests; adds serde_yml,
base64, serde_json as dev-dependencies.

Signed-off-by: Giedrius Burachas <gburachas@nvidia.com>
(cherry picked from commit 91807f984a3b16846e35d6ca0d5ec41057aafa3a)
Signed-off-by: Jamie King <jamiek@nvidia.com>

* feat(driver-mxc): inject agent_env into sandbox process.env

Add MxcComputeConfig.agent_env: each entry is either KEY=VALUE (verbatim) or a bare KEY resolved from the gateway host environment at launch, keeping secrets (e.g. inference API keys) out of the config file. Wire it into the agent process so gateway-launched agents can authenticate to cloud endpoints (process.env was previously hardcoded empty). Unit-tested via resolve_agent_env_passthrough_and_host_lookup.

Also add a gateway-driven cloud-inference (T1) test harness: mxc-inference.toml (agent_env + curl agent), inference.yaml policy, and run-inference-test.ps1 which starts the gateway, creates an isolation_session sandbox, runs an authenticated Nemotron call, and bundles redacted results. Documented agent_env in mxc-gateway.toml. Validated end-to-end on the test box (chat HTTP 200 + completion via the gateway).

(cherry picked from commit 94d9e827b8b77e0af9dc654943ea8e7d8400cc9f)
Signed-off-by: Jamie King <jamiek@nvidia.com>

* test(driver-mxc): avoid unsafe env mutation in resolve_agent_env test

Replace std::env::{set,remove}_var (unsafe + racy under parallel test
execution in edition 2024) with a read-only PATH lookup. Preserves all
three behaviors under test and drops the #[allow(unsafe_code)].

(cherry picked from commit ac5766eeab2db4e8cc6fcd8d8a97809edaf3df30)
Signed-off-by: Jamie King <jamiek@nvidia.com>

* fix(mxc): adapt MXC driver to current GitHub OpenShell API

The MXC driver crate was authored on GitLab against an earlier proto/core
API. Adapt it to the API on GitHub main:

- build_capabilities_response no longer takes supports_interactive_session
- DriverSandboxSpec.gpu (bool) is now resource_requirements; detect GPU via
  effective_driver_gpu_count(driver_gpu_requirements(..))
- DriverSandbox gained a `workspace` field
- SandboxPolicy gained `network_middlewares`: pass it through the proxy split,
  emit a loss item on the coarse MXC path, and account for it in the mapper
  drift-guard test

Verified: cargo check + 75 mock-based tests pass (lib 27, examples 10,
policy_mapper_matrix 38).

Signed-off-by: Jamie King <jamiek@nvidia.com>

* feat(server): wire the MXC compute driver into the gateway on Windows

Register openshell-driver-mxc as the Windows-only in-process compute
backend so compute_driver = "mxc" resolves to a working runtime:

- ComputeRuntime::new_mxc, adapted to the current 11-arg from_driver
- mxc_policy_sink A1 side channel, staged in create_sandbox before dispatch
- mxc_config_from_context loader and the Mxc dispatch arm (Windows
  constructs; other targets return an explicit "Windows-only" error)
- Windows-gated openshell-driver-mxc dependency
- Mxc arms for the telemetry, config-file required-fields, and CLI
  reserved-builtin matches to keep them exhaustive/correct

Verified with cargo check --workspace --features openshell-prover/bundled-z3
on x86_64-pc-windows-msvc, stacked on PR #2496.

Signed-off-by: Jamie King <jamiek@nvidia.com>

* fix(driver-mxc): implement GetGatewayListenerRequirements for #2496 base

Signed-off-by: Jamie King <jamiek@nvidia.com>

* test(driver-mxc): add probe-gated real wxc-exec test lane (no mocks)

- tests/wxc_exec_real.rs: ignored-by-default integration tests against a
  real wxc-exec. Six --dry-run contract tests run wherever the binary
  exists (they caught the network.proxy shape mismatch); enforcement
  tests (processcontainer default-deny positive/negative, isolation
  session lifecycle round trip with a deprovision drop-guard) probe the
  backend and SKIP with a recorded reason where it is not live.
- examples/probe-mxc-host.ps1: classifies a host (OS build, --probe,
  per-backend trial) and emits a JSON capability verdict.
- examples/run-mxc-e2e.ps1 + e2e-policies/: scenario runner generalizing
  run-demo.ps1 (fs-rw, fs-readonly, fs-default-deny-empty,
  network-policy-rejected) with PASS/FAIL/SKIP gating and a stale
  OPENSHELL_MXC_MOCK_WXC guard in real mode.
- tasks/windows.toml: windows:test:mxc-real:x64, windows:e2e:mxc,
  windows:e2e:mxc:mock.

Signed-off-by: Giedrius Burachas <gburachas@nvidia.com>
(cherry picked from commit 49afafe892caded59c4df50a9b652011ad97f41c)
Signed-off-by: Jamie King <jamiek@nvidia.com>

* fix(driver-mxc): address PR review feedback

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>

* docs: defer public MXC documentation

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>

* fix(driver-mxc): build Windows capabilities response

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>

* fix(server): gate in-tree tracing on Windows

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>

* test(windows): fix cross-platform test assumptions

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>

* test(windows): verify process and PEM portably

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>

* test(windows): keep lifecycle command in policy

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>

---------

Signed-off-by: Jamie King <jamiek@nvidia.com>
Signed-off-by: Giedrius Burachas <gburachas@nvidia.com>
Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>
Co-authored-by: Prashant K <pkhodade@nvidia.com>
Co-authored-by: Giedrius Burachas <gburachas@nvidia.com>
Co-authored-by: Shailendra Singh <shailendras@nvidia.com>
Co-authored-by: Drew Newberry <385+drew@users.noreply.github.com>
2026-08-27 07:24:04 +00:00
krishicks d0dfb22baf feat(kubernetes): export driver traces over OTLP (#2958)
Mirror the VM, Podman, and Docker driver tracing setup for Kubernetes.
Export standalone driver spans through OTLP/gRPC as the distinct
openshell-driver-kubernetes service, preserve gateway trace context, record
lifecycle operations and gRPC failures, and flush spans on shutdown.

Kubernetes currently runs in-process when selected as a built-in gateway
driver. Use the temporary server-boundary shim shared with Podman and Docker
so traces retain the shape they will have when Kubernetes moves to a
separate process. Move the common ComputeDriver RPC tracing layer into
openshell-otel to keep all drivers aligned.

Propagate the active W3C context through the controller-reserved Sandbox
annotation and enable Agent Sandbox OTLP export in the local k3s workflow.
This connects asynchronous controller reconciliation spans to the originating
OpenShell create trace.

Expose gateway OTLP configuration through Helm and add an Aspire collector
to the local k3s workflow. Extend helm:k3s:forward with OTLP ingest and trace
UI forwarding for Kubernetes and local container gateway development.

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-08-26 21:22:51 +00:00
krishicks c399342649 feat(dev): unify local Kubernetes gateway workflow (#2914)
Make the local k3s gateway workflow match the Docker and Podman flows by
registering and selecting successful plaintext Skaffold deployments with the
OpenShell CLI. Derive the registration name from the worktree-specific k3d
cluster name so parallel worktrees retain independent gateway metadata.

Add helm:k3s:forward as the standard way to expose the Kubernetes gateway on
localhost:8090, and update the development and debugging guidance to use the
active registered gateway instead of one-off endpoint flags.

Signed-off-by: Kris Hicks <khicks@nvidia.com>
2026-08-26 14:44:47 +00:00