mirror of
https://github.com/NVIDIA/OpenShell.git
synced 2026-10-03 16:11:17 +08:00
pull-request/3773
243
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
87929ad130 |
docs: keep page URLs aligned with file paths (#3713)
* docs: align page file names and nav labels with published URLs Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs: redirect moved dev pages and fix agent guide redirects Signed-off-by: Johnny Greco <jogreco@nvidia.com> * ci(docs): check that page URLs match their file paths Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(docs): align navigation checks with repo conventions Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs: omit historical redirect aliases Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(docs): preserve published overview and TypeScript URLs Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(docs): redirect unversioned overview and TypeScript URLs Signed-off-by: Johnny Greco <jogreco@nvidia.com> --------- Signed-off-by: Johnny Greco <jogreco@nvidia.com> |
||
|
|
8934b74a85 |
fix(docs): sync redirects with versioned snapshots (#3754)
* fix(docs): sync redirects with versioned snapshots Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(docs): make redirect validation types explicit Signed-off-by: Johnny Greco <jogreco@nvidia.com> --------- Signed-off-by: Johnny Greco <jogreco@nvidia.com> |
||
|
|
d1a19c70ee |
fix(homebrew): overwrite generated gateway config during migration (#3752)
Closes #3746 Use Homebrew atomic_write for exact legacy config migrations and keep the first-install write distinct in the formula test. Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
a67567e583 |
fix(snap): require mTLS for the snap gateway (#3726)
* fix(snap): require mTLS for the snap gateway Replace the installer opt-in with an authenticated snap gateway. The wrapper no longer forces plaintext, so the gateway serves TLS from the bundle it already generates in $SNAP_COMMON/tls. The install hook writes a config that enables mTLS user auth instead of unauthenticated access, and a new post-refresh hook migrates the exact legacy default on existing installs. install.sh waits for the gateway, detects whether it serves TLS, copies the client bundle into the target user's snap state directory, and registers the gateway over HTTPS. Older plaintext snap revisions still register over HTTP with a warning. The release canary asserts mTLS auth and HTTPS registration. Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(snap): pass config preflight and detect the mTLS gateway reliably An explicit [openshell.gateway.mtls_auth] table fails config preflight, which validates mTLS auth before the local TLS bundle supplies the client CA. Write a default that pins the Docker driver instead; with the wrapper's TLS bundle the gateway requires client certificates and enables mTLS user auth automatically, as the native packages do. The mTLS gateway rejects TLS handshakes without a client certificate, and it still answers plaintext loopback HTTP for sandbox service routing, so the installer could misdetect it as a legacy plaintext gateway. Probe HTTPS with the root-owned client bundle, and treat a gateway as legacy only when a plaintext gRPC Health call succeeds. Signed-off-by: Drew Newberry <anewberry@nvidia.com> * chore(snap): simplify install hook comment Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(snap): migrate insecure gateway configs on refresh Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(snap): simplify mTLS detection and config migration Detect the mTLS snap from the installed revision's post-refresh hook instead of probing plaintext gRPC, and drop the scheme global. Remove the installer's pre-hook config fallback, which is dead now that every channel ships the install hook and which wrote the insecure default. Give the install hook a single write path with a simple backup name, and shorten the manual client certificate steps. Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(snap): stop keeping a copy of replaced insecure configs Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(install): make the snap an opt-in install method Stop selecting the OpenShell snap just because the snap command exists. Linux installs default to the Debian or RPM package; OPENSHELL_INSTALL_METHOD=snap (or deb, rpm) selects the package explicitly. Hosts that already have the OpenShell snap keep refreshing it rather than gaining a second gateway on the same port. The release canary and snap repro script opt in explicitly. Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(snap): let the gateway auto-detect its compute driver Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(snap): restart the gateway after refresh Published revisions use refresh-mode: endure, and snapd honors the old revision's setting during a refresh, so the plaintext gateway kept running with the migrated config unused until a manual restart. Restart the gateway from the post-refresh hook so the mTLS config takes effect immediately. Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(snap): drop refresh notes from the snap description Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(snap): trim snap refresh notes from installation docs Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
6f596828ad |
docs(fern): publish ordered version snapshots (#3721)
* docs(fern): publish and order versioned release docs Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(fern): remove version availability badges Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(fern): keep version badges optional Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> --------- Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> |
||
|
|
c9257c8447 |
fix(policy): restore policy.local and proposal conformance (#3689)
* fix(policy): restore policy.local and proposal conformance Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * test(conformance): select policy scenarios by name Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * test(conformance): reduce policy scenario timing flakes Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * test(conformance): assert proposals target Bash Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> --------- Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> |
||
|
|
1374672967 |
fix(helm): restore Kubernetes e2e chart rendering (#3692)
Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
08548713c5 |
fix(install): honor pinned releases and speed up prerelease discovery (#3681)
* fix(install): use native packages for prereleases and speed up discovery Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(install): honor pinned versions and guard Snap migration Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(install): omit Snap migration guard Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
52cb8ecee7 |
fix(kubernetes): remove NetworkPolicy acknowledgement (#3677)
Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
4cd1351e18 |
fix(test): use portable file mode checks in snap installer tests (#3670)
The snap installer tests added in #3656 checked the gateway config mode with GNU stat -c, which BSD stat on macOS rejects, so mise run ci failed locally on macOS. Check the mode with find -perm instead, which matches the exact mode on both GNU and BSD systems. Signed-off-by: Kris Hicks <khicks@nvidia.com> |
||
|
|
6c864ec9ab |
fix(install): avoid installing incompatible docker snap (#3666)
* fix(install): avoid installing incompatible docker snap The work to land RFC-0012 added new restrictions when interacting with Docker by setting `NoNewPrivs`. This prevents the `docker` snap from transitioning its AppArmor profile from `snap.docker.dockerd` to `docker-default` when it tries to launch a container. Thus, the `docker` snap is currently incompatible with OpenShell. This commit prevents `install.sh` from installing the `docker` snap before installing the `openshell` snap, and instead requires the user to install a non-snap Docker daemon before proceeding with installing the snap. Systems without the `snap` command are unaffected, since they install native packages without checking for the presence of Docker or other compute providers. Signed-off-by: Oliver Calder <oliver.calder@canonical.com> * fix(snap): only snapd 2.76 for openshell snap since store installs work Signed-off-by: Oliver Calder <oliver.calder@canonical.com> * fixup! fix(install): avoid installing incompatible docker snap Signed-off-by: Oliver Calder <oliver.calder@canonical.com> * fixup! fix(snap): only snapd 2.76 for openshell snap since store installs work Signed-off-by: Oliver Calder <oliver.calder@canonical.com> --------- Signed-off-by: Oliver Calder <oliver.calder@canonical.com> |
||
|
|
8369bc11a5 |
fix(podman): restore host gateway alias mediation (#3606)
* fix(podman): restore host gateway alias mediation Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(podman-e2e-tests): enable broader test podman e2e coverage Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(tests): make test more reliable Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(podman): fix macos linting error Signed-off-by: Gordon Sim <gsim@redhat.com> --------- Signed-off-by: Gordon Sim <gsim@redhat.com> |
||
|
|
0518bd4c83 |
fix(sandbox): preserve local sessions across host sleep (#3573)
* fix(cli): recover sandbox connect transport Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(auth): propagate non-expiring local sandbox sessions Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(cli): distinguish main exit from transport loss Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(cli): bound sandbox connect recovery Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
e60098d748 |
fix(snap): install openshell snap via install.sh when snap available (#3656)
* fix(snap): update stale snap docs and tests Previously, the `openshell` snap required the `docker` snap. Now, it works with any Docker daemon running on the system. Furthermore, the `snap-declaration` assertion on the `openshell` snap when installed from the Snap Store causes the `openshell` snap to always connect to the system `:docker` slot, rather than a slot provided by the `docker` snap. This commit updates the documentation, including the `description` field in `snapcraft.yaml`, to ensure that all information is correct and up-to-date. Additionally, some tests connected the `openshell:docker` plug to the `docker` snap's `docker:docker-daemon` slot, which is inconsistent with how the `openshell` snap operates when installed from the store. Update those tests to connect to the system `:docker` slot as well. Signed-off-by: Oliver Calder <oliver.calder@canonical.com> * fix(snap): require snapd 2.76 for openshell snap Signed-off-by: Oliver Calder <oliver.calder@canonical.com> * fix(nix): align snap gateway reproducer timeout with release-canary Signed-off-by: Oliver Calder <oliver.calder@canonical.com> * fix(snap): require snapd 2.77 for openshell snap Signed-off-by: Oliver Calder <oliver.calder@canonical.com> * fixup! fix(snap): require snapd 2.77 for openshell snap Signed-off-by: Oliver Calder <oliver.calder@canonical.com> * feat(snap): install openshell snap via install.sh when snap available Change `install.sh` to install the `openshell` snap by default when snapd is installed on the host. This installs the snap from the `latest/stable` channel, which should match the most up-to-date release tag on github. If `OPENSHELL_VERSION=dev` is set for `install.sh`, then it will install the `openshell` snap from the `latest/edge` channel, which matches the latest dev release available on github. The `openshell` snap currently requires Docker in order to function. If a Docker daemon is already installed on the system, it will be used by the `openshell` snap. Otherwise, `install.sh` will install the `docker` snap first, wait for the Docker daemon to be ready, and then install the `openshell` snap. Also, update the `release-canary.yml` to split the `ubuntu-snap` job into `ubuntu-snap-system-docker` and `ubuntu-snap-provisions-docker`, which test the two aforementioned scenarios. Previously, `ubuntu-snap` manually installed a given snap artifact as built from CI, but with these new jobs, it instead uses `install.sh` to install the published `openshell` snap from the `latest/edge` track, thus matching the behavior of the other release canary jobs. Make corresponding changes to the `nix` guest reproducer, documentation, and tests. Signed-off-by: Oliver Calder <oliver.calder@canonical.com> * fix(snap): configure local gateway authentication The `openshell` snap runs the gateway as a systemd system service, which runs as root. Thus, the mTLS certs are generated by root and stored in a root-owned directory to which non-root users do not have access. For this reason, the snap's `openshell-gateway-wrapper` script sets `OPENSHELL_DISABLE_TLS=true`. This commit ensures that the `openshell` snap's gateway allows unauthenticated local access by writing a default `gateway.toml` configuration file during the install hook, which runs after the snap is first installed but before services are started. The config file contains sets `allow_unauthenticated_users = true`. Signed-off-by: Oliver Calder <oliver.calder@canonical.com> * fix(install): configure snap gateway authentication Recently, a new install hook was added which writes a default config file for the `openshell` snap to allow unauthenticated local access to the gateway. This is because the gateway service runs as root and the mTLS certificates are not accessible to non-root users. However, the `install.sh` script installs the `openshell` snap from the snap store, and the published version may not yet have that new install hook. Or, the user may already have the snap installed, in which case the install hook does not run. In either case, we need `install.sh` to ensure that the config file is written to set the gateway auth to `allow_unauthenticated_users = true`, and then restart the openshell gateway service. Signed-off-by: Oliver Calder <oliver.calder@canonical.com> * fixup! feat(snap): install openshell snap via install.sh when snap available Signed-off-by: Oliver Calder <oliver.calder@canonical.com> --------- Signed-off-by: Oliver Calder <oliver.calder@canonical.com> |
||
|
|
679b190677 |
feat(testing): support independent gateway and supervisor image overrides (#3341)
* feat(testing): normalize configurable test images Signed-off-by: Bobbins228 <mcampbel@redhat.com> Signed-off-by: Kris Hicks <khicks@nvidia.com> * feat(helm): add global image overrides Signed-off-by: Bobbins228 <mcampbel@redhat.com> * feat(helm): support image registry overrides Signed-off-by: Bobbins228 <mcampbel@redhat.com> Signed-off-by: Kris Hicks <khicks@nvidia.com> * refactor(helm): simplify image configuration Signed-off-by: Bobbins228 <mcampbel@redhat.com> * fix(e2e): avoid reloading reused kind sandbox image Signed-off-by: Bobbins228 <mcampbel@redhat.com> * fix(helm): default sandbox image to nvcr.io/nvidia/base/ubuntu:24.04 Signed-off-by: Kris Hicks <khicks@nvidia.com> --------- Signed-off-by: Bobbins228 <mcampbel@redhat.com> Signed-off-by: Kris Hicks <khicks@nvidia.com> Co-authored-by: Kris Hicks <khicks@nvidia.com> |
||
|
|
0b351c4a9b |
fix(helm)!: reduce gateway Secret privileges (#3616)
* fix(driver-kubernetes-secrets)!: store provider credentials in one namespace The Kubernetes Secrets credential driver now stores every credential in its configured namespace in all workspace modes and rejects handles that reference any other namespace before contacting the Kubernetes API. The gateway reaches credential Secrets through the Role in that namespace; this allows removing the Secret rules from the ClusterRole. - Remove the workspace_mode, gateway_id, and allow_reference_namespace driver settings and stop rendering them from Helm. Configurations that set them fail at startup. Existing credential state is not migrated. - Add server.credentialDrivers.kubernetesSecrets.createNamespace to provision a dedicated credential namespace. The namespace is kept on uninstall, adopted by a reinstall of the same release, and left untouched when something else owns it. - Update the gateway config reference, Kubernetes setup docs, 0.1.0 upgrade guide, compute-runtime architecture, and cluster debugging skill. Signed-off-by: Kris Hicks <khicks@nvidia.com> * fix(helm): reduce gateway Secret permissions Remove the gateway's Secret list permission in every workspace mode and grant source Secret reads through a Role in the sandbox namespace. Bootstrap Secret cleanup deletes Secrets by exact name instead of listing them. - Grant get on the copied client TLS and image-pull Secrets through a Role in the sandbox namespace. The ClusterRole keeps get and patch on those names for the ownership check and server-side apply into workspace namespaces. - Delete sandbox and supervisor bootstrap Secrets by exact name, derived from the runtime generation recorded on the Sandbox and, on restart, the target generation, tolerating 404. The generation annotation is cleared only after cleanup succeeds, and each bootstrap Secret has a Pod owner reference, so garbage collection removes any generation the driver does not name. - Drop Secret list from the ClusterRole and the shared-mode sandbox Role. - Extend the managed e2e RBAC checks to Secret list. - Update the Kubernetes setup and sandbox runtime docs, compute-runtime architecture, and cluster debugging skill. Signed-off-by: Kris Hicks <khicks@nvidia.com> * fix(driver-kubernetes): stage workspace Secrets per runtime generation Every Secret the Kubernetes driver writes into a workspace namespace is now scoped to one sandbox runtime generation, immutable, and created with create only, so the gateway never reads, patches, or adopts an existing Secret there. This removes the gateway's cluster-wide get and patch on the copied client TLS and image-pull Secret names. - In managed mode, create an immutable copy of each configured image-pull Secret per generation, named os-pull-<id>-<generation>-<n> and owned by the generation's workload and supervisor Pods. Pods and the restarted Sandbox template reference those names, and generation cleanup deletes them by name. A Secret already holding a generation name fails the create. - Outside shared mode, stage the gateway client TLS material into the supervisor bootstrap Secret instead of copying the client TLS Secret into the workspace namespace. - Remove the fixed-name TLS and image-pull copies, the target ownership read, and the ClusterRole get and patch rule on the copied names. Source reads stay in the sandbox-namespace Role. - Update the managed e2e to expect generation image-pull Secrets and client TLS material in the supervisor bootstrap Secret, and to check that the gateway cannot read the copied names in workspace namespaces. - Update the gateway config and compute driver references, Kubernetes setup and sandbox runtime docs, compute-runtime architecture, driver README, and cluster debugging skill. Signed-off-by: Kris Hicks <khicks@nvidia.com> * fix(helm)!: grant operator-mode Secret permissions through the workspace chart The operator-mode gateway ClusterRole grants no Secret permissions. The openshell-workspace chart Role, installed in each operator-managed namespace, grants the gateway create and delete on Secrets for sandbox runtime generations. Operator-managed namespaces require the workspace chart. - Fail the chart tests on any ClusterRole rule that includes Secrets in operator and shared modes. - Install the workspace chart when the operator e2e provisions a namespace, and assert that the gateway has no Secret permissions in a namespace without it. - Update the Kubernetes setup docs, 0.1.0 upgrade guide, compute-runtime architecture, and cluster debugging skill. Signed-off-by: Kris Hicks <khicks@nvidia.com> --------- Signed-off-by: Kris Hicks <khicks@nvidia.com> |
||
|
|
a408f5dd08 |
chore(kubernetes): update Agent Sandbox to v1.0.3 (#3578)
Signed-off-by: Mrunal Patel <mrunalp@gmail.com> |
||
|
|
907f894ebc |
ci(release): publish prereleases with qualification summary (#3593)
Signed-off-by: Evan Lezar <elezar@nvidia.com> |
||
|
|
88357775ad |
docs(windows): align Z3 pin with z3-sys 0.13 (#3561)
z3-sys 0.13 generates bindings for the Z3 5.x series and resolves GH_RELEASE_VERSION to 5.1.0 for the prebuilt-release path. The Windows wrapper still pinned Z3_SYS_Z3_VERSION to 4.16.0, so a prebuilt Windows build would fetch an archive that does not match the generated bindings. Move the wrapper pin to 5.1.0 and update the contributor, architecture, and skill references that still named 4.16.0. Signed-off-by: Jim Meyer <jimeyer@nvidia.com> |
||
|
|
7139df8ca5 |
fix(exec): preserve output after stdin EOF and verify stream completion (#3359)
* fix(exec): preserve output after stdin EOF and verify stream completion Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> * test(exec): cover fair duplex progress and live SDK completion Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> * test(sdk): compare large exec buffers with native equality Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> * test(exec): use workspace-scoped sandbox name in EOF regression Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> --------- Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> |
||
|
|
50230616d5 |
refactor(runtime): retire Community image dependencies (#3386)
* feat(sandbox): default to official Alpine sandbox image default_sandbox_image() now returns docker.io/library/alpine:3.22, a generic version-qualified official image, so a fresh install no longer depends on the community sandbox image catalog. All compute drivers (docker, podman, kubernetes, vm) inherit this fallback. Part of #3116. Signed-off-by: Akram Signed-off-by: Akram <akram.benaissi@gmail.com> * feat(deploy): default deployment configs to the official Alpine sandbox image Update the shared gateway default_image, Helm chart values, the standalone Kubernetes manifest, and the dev gateway task scripts to use docker.io/library/alpine:3.22 instead of the community base image, consistent with default_sandbox_image(). GPU e2e image-build base is left unchanged (CUDA needs a glibc base). Part of #3116. Signed-off-by: Akram Signed-off-by: Akram <akram.benaissi@gmail.com> * feat(driver): default to numeric non-root identity for USER-less images With the default sandbox image now Alpine, images that declare no OCI USER must start instead of being rejected. When the image declares no USER and the policy requests none, the Podman and Docker drivers now supply a numeric non-root identity (DEFAULT_SANDBOX_UID/GID = 1000) instead of rejecting, matching the numeric-identity behavior of the Kubernetes and VM drivers. The supervisor's resolved-identity path runs the sandbox as a synthesized non-root account without the account existing in the image. Images that declare a USER keep the OCI resolution path unchanged. Part of #3116. Signed-off-by: Akram <akram.benaissi@gmail.com> Signed-off-by: Evan Lezar <elezar@nvidia.com> * test(conformance): use Alpine workload image Signed-off-by: Evan Lezar <elezar@nvidia.com> * refactor(policy): drop community image /app path from default policy The restrictive default policy granted read-only access to /app, a directory that only existed in the community base image. A generic Alpine default has no /app, so remove it. Landlock best-effort already ignores absent paths; this just stops advertising a community-specific layout in the default. Part of #3116. Signed-off-by: Akram Signed-off-by: Akram <akram.benaissi@gmail.com> * docs(config): document Alpine default images Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(podman): report early sandbox termination Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(podman): initialize rootless workspace ownership Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(sandbox): qualify NVIDIA Ubuntu default Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): initialize rootful default workspace Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(sftp): add native sandbox adapter Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sftp): gate runtime helper support to Linux Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sftp): support standard OpenSSH file operations Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sftp): harden rename and special file handling Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(runtime): remove community image dependencies Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(e2e): build provider readiness tool fixture Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(e2e): use a dedicated Noble fixture for Docker tests Signed-off-by: Evan Lezar <elezar@nvidia.com> --------- Signed-off-by: Akram Signed-off-by: Akram <akram.benaissi@gmail.com> Signed-off-by: Evan Lezar <elezar@nvidia.com> Signed-off-by: Drew Newberry <anewberry@nvidia.com> Co-authored-by: Evan Lezar <elezar@nvidia.com> Co-authored-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
99ed6a9df0 |
chore(build): remove stale static-supervisor leftovers (#3520)
The opt-in glibc-static supervisor variant from #2682 is gone: the Nix release builds (#2977) removed it from CI and the supervisor Dockerfile, and the RFC 0012 sandbox split (#2942) removed it from the staging script. The split also moved the binary that runs inside workload images to openshell-sandbox; the supervisor now runs from its own image and is dynamically linked. A few places still describe the old model: - skills/debug-openshell-cluster told operators to check that supervisor_image contains a static /openshell-supervisor. Ask instead for a supervisor from the matching release whose loader and shared libraries are available inside that image, and give a `--version` check that shows both. The static requirement for /openshell-sandbox is unchanged. - verify-static-binary.sh justified its check with the supervisor and the glibc-static variant. Describe the property it enforces instead, with the musl sandbox runtime as the main example. - stage-prebuilt-binaries.sh kept an unreachable gnu-static case in target_triple. No behavior change: resolve_component only ever selects gnu or musl. Signed-off-by: Emilien Macchi <emacchi@redhat.com> |
||
|
|
0a770d9173 |
feat(kubernetes): support HA gateway rebalancing (#1868)
* feat(kubernetes): support HA gateway rebalancing Signed-off-by: Drew Newberry <anewberry@nvidia.com> * perf(server): cache peer connections, tokens, and owner lookups Every forwarded relay rebuilt its setup from scratch: an owner lookup, a blocking read of the peer token, a TLS connect to the owning replica, and a TokenReview plus Pod GET on the receiving side. Sandbox service routing does this per HTTP request, so the apiserver calls scaled with traffic. Cache all of it on ServerState: - peer channels pooled per endpoint, so relays multiplex over one connection instead of redialing - peer tokens keyed by SHA-256, expiring at min(ttl, token exp) so a hit cannot accept an expired token - owner records for 3s against a 45s ownership TTL, still freshness checked before use Entries are evicted when a relay fails. Also raise HTTP/2 max_concurrent_streams to 1024, since pooling funnels every relay between two replicas onto one connection and hyper's default of 200 sits below the 256 pending-relay budget. Signed-off-by: divesh <dgude@nvidia.com> * perf(server): pool upstream connections for sandbox services Each HTTP request to a sandbox service opened its own supervisor relay, paying a new TCP connection and HTTP/1 handshake every time. Worse, it counted against the 32 in-flight relay cap, so a service handling more than 32 concurrent requests failed outright. Pool idle upstreams per endpoint and port, up to 8 each for 15s. Reuse is safe because the pool only returns a connection hyper reports as ready, and HTTP/1 cannot start a request until the previous body has drained. Upgrades are never pooled since they take the connection over, and a failed send evicts that endpoint. Pruning is bounded per key, with the full sweep limited to once per 30s. Signed-off-by: divesh <dgude@nvidia.com> * fix(server): address HA gateway review findings (#3449) - Let a gateway own supervisor sessions without a peer endpoint. Requiring one whenever the store is PostgreSQL broke every single-instance PostgreSQL deployment, because no sandbox supervisor could connect. A cross-replica request to an owner that advertises no endpoint now fails immediately naming the cause, instead of retrying until the wait timeout. - Close a supervisor session on heartbeat only when another replica owns it, or after renewals fail for the ownership TTL. A database error no longer drops every session heartbeating during an outage. - Clamp owner record ages at zero so a skewed or corrupt stored timestamp cannot produce a negative age. - Bound the cross-object advisory lock with a lock timeout, so a stuck holder fails instead of blocking every mutation in the fleet. - Refuse to start when a peer endpoint is configured on a multi-replica backend but peer authentication is unavailable, and warn when a multi-replica backend has no peer endpoint at all. - Reject a plaintext peer endpoint when the gateway serves TLS. - Skip the sandbox watch poller on single-replica backends, where the local update bus already sees every write. - Rate-limit the peer owner cache sweep so an insert no longer scans the whole map under the lock. - Retry GET and HEAD on a pooled upstream the sandbox closed, instead of returning 502, and drop an emptied endpoint from the pool right away. - Document the gateway peer environment variables and the post-rollout ownership skew operators should expect. Signed-off-by: divesh <dgude@nvidia.com> * fix(server): harden HA supervisor ownership Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> Signed-off-by: divesh <dgude@nvidia.com> Co-authored-by: Drew Newberry <anewberry@nvidia.com> Co-authored-by: divesh <dgude@nvidia.com> Co-authored-by: Divesh Chowdary <47188680+FrostGod@users.noreply.github.com> |
||
|
|
dee4f98dda |
chore(vm): refresh runtime defaults and hardening (#3446)
Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
d6f3e1f5f2 |
fix(packaging): keep SPDX comments out of Debian control (#3483)
Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
17ce738bfb |
fix(ci)!: remove gateway callback listener dependency (#3365)
* fix(ci): repair post-merge release canary Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(packaging): bootstrap canary runtime prerequisites Signed-off-by: Drew Newberry <anewberry@nvidia.com> * ci(canary): collect macOS VM diagnostics Signed-off-by: Drew Newberry <anewberry@nvidia.com> * ci(canary): pin libkrun-compatible macOS runner Signed-off-by: Drew Newberry <anewberry@nvidia.com> * ci(canary): limit macOS smoke test to package startup Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(compute)!: remove gateway callback listeners Run Docker supervisors on host networking so they use the operator-configured primary gateway endpoint. Remove the unused compute-driver callback listener negotiation and listener-scoped routing machinery. BREAKING CHANGE: The ComputeDriver API no longer exposes GetGatewayListenerRequirements or GatewayListenerRequirement. External drivers must regenerate bindings and connect supervisors to the configured primary gateway endpoint. Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(docker): use sandbox runtime image in launcher Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(podman): exercise production endpoint selection Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(e2e): route supervisors to reachable gateways Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(e2e): align Podman endpoint fixtures Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): preserve host aliases for supervisors Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): align sandbox host gateway pin Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(e2e): address Docker fixtures by bridge IP Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(e2e): serialize sandbox lifecycle cases Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(e2e): host Docker TCP fixture with gateway Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(e2e): use loopback for host-network supervisor Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
d91b1999a0 |
feat(api)!: use sandbox names as canonical RPC references (#3272)
* feat(api)!: use sandbox names as canonical references Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(cli): update forward color fixture for workspace scope Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(supervisor): use sandbox names for settings lookup Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(e2e): use canonical sandbox request fields Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(e2e): use canonical sandbox receipt field Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(server): harden sandbox mutation handling Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(ci): update rebased sandbox references Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(api)!: standardize canonical entity references Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(api): codify protobuf API conventions Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(api): preserve workspace selector semantics Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(api): restore workspace selector parity Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(api): preserve descriptive name fields Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(api): update e2e request fixtures Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(core): omit workspace selector during bootstrap Signed-off-by: Drew Newberry <anewberry@nvidia.com> * chore(api): remove proto convention checker Signed-off-by: Drew Newberry <anewberry@nvidia.com> * chore(api): refresh schema fingerprints after rebase Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(cli): use canonical provider receipt field Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
c5a8c4d220 |
chore(vm): bump libkrun to v1.19.4 and libkrunfw to v5.6.1 (#3451)
The libkrun and libkrunfw projects moved from the containers/ GitHub org to libkrun/. Update all clone URLs and version pins accordingly. libkrun v1.19.4 introduces krun_add_virtiofs4 with a permissions semantics parameter, enabling correct file ownership mapping on host-to-guest bind mounts — required for issue #2585. - libkrun: v1.17.4 → v1.19.4 (libkrun/libkrun) - libkrunfw: 463f717b → v5.6.1 (libkrun/libkrunfw) - Centralise LIBKRUN_REF in pins.env instead of hardcoding in each build script - Source pins.env in the macOS build script for consistency with the Linux build script - Remove the init/init make target step — since libkrun/libkrun@05c4eb7 the init blob moved into the init_blob crate and is built by cargo Ref: https://github.com/NVIDIA/OpenShell/issues/2585 Signed-off-by: Florent Benoit <fbenoit@redhat.com> |
||
|
|
2263685cf3 |
test(tmachine): migrate Keycloak provider refresh coverage (#3404)
* test(tmachine): add Keycloak provider refresh suite Signed-off-by: Evan Lezar <elezar@nvidia.com> * refactor(tmachine): share container runtime detection Signed-off-by: Evan Lezar <elezar@nvidia.com> * ci(tmachine): run feature suites in GitHub Actions Signed-off-by: Evan Lezar <elezar@nvidia.com> * ci(tmachine): run conformance with Podman tests Signed-off-by: Evan Lezar <elezar@nvidia.com> * ci(tmachine): cover provider refresh with Podman Signed-off-by: Evan Lezar <elezar@nvidia.com> * ci(integration): split input preparation from runners Signed-off-by: Evan Lezar <elezar@nvidia.com> --------- Signed-off-by: Evan Lezar <elezar@nvidia.com> |
||
|
|
c108c31696 |
docs(fern): sync announcement configuration (#3436)
* docs(fern): sync announcement configuration Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(docs): scope announcements to synced channel Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(fern): use dev as source version Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(fern): use channel-neutral logo link Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> --------- Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> |
||
|
|
8b77925eb7 |
feat(installer): support prerelease installations (#3364)
* feat(installer): support prerelease installations Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(readme): clarify 0.1.0 production readiness Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(readme): refine production readiness message Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(release): publish rolling prerelease channels Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(release): keep prereleases out of GitHub releases Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(readme): refine 0.1.0 notice Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(readme): expand 0.1.0 notice Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(readme): simplify 0.1.0 notice Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(installer): simplify prerelease alias Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(installer): require successful prerelease runs Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(release): reuse exact prerelease artifacts Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(installer): include prover in prerelease bundle Signed-off-by: Drew Newberry <anewberry@nvidia.com> * ci: remove prerelease post-publish validation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(fern): sync announcement configuration Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs: remove stale prerelease canary claim Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(fern): defer announcement synchronization Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(fern): simplify release announcement Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> Co-authored-by: Piotr Mlocek <pmlocek@nvidia.com> |
||
|
|
04146692d9 |
refactor(providers)!: make provider profiles import-only (#3383)
* chore(providers): remove dead provider plugin modules Twelve modules under crates/openshell-providers/src/providers/ were never declared in providers/mod.rs, so they have not been compiled since the plugin registry was narrowed to the two adapters it still registers. Four of them (generic, gitlab, opencode, outlook) key off provider type identifiers that normalize_provider_type already retired. Keep google_cloud and vertex, which are the only plugins ProviderRegistry::new registers. Signed-off-by: Philippe Martin <phmartin@redhat.com> * test(providers): load example profiles from providers/ at test time Add an example_profiles module that reads the YAML under providers/ from the source checkout at run time and parses it with the existing profile loader. It is gated behind a new non-default example-profiles feature so it is available to this crate's own tests and, once wired into dev-dependencies, to gateway and CLI tests, while never reaching a release binary. Nothing consumes it yet; later changes move the test fixtures off the compiled catalog and onto these files. Signed-off-by: Philippe Martin <phmartin@redhat.com> * test: source profile fixtures from providers/ instead of the compiled catalog Unit and integration fixtures reached into builtin_profiles() to get a profile to work with, which ties the tests to the compiled catalog rather than to the files an operator would import. Point them at the example_profiles loader instead. The profiles.rs tests keep their coverage unchanged and become explicit golden tests over providers/*.yaml: they are what keeps those files valid once nothing compiles them. builtin_profiles_are_sorted_by_id widens into a test that the whole example set parses, sorts, and lints clean as one catalog. The CLI fake gateways serve the example profiles the way a real gateway serves what an operator imported, through a shared helper. No behavior change: the gateway still loads the built-in source by default and serves the same profiles. Signed-off-by: Philippe Martin <phmartin@redhat.com> * docs(providers): document the example provider profiles Each file under providers/ now opens with a header naming its expected client binary identities, the image layout those paths assume, the credential scope, the endpoint access it grants, and a smoke test. Add a README covering the import commands and why a profile should be copied and edited rather than imported unchanged. Several of these profiles bind network access to paths that only exist in the OpenShell Community image — /sandbox/.venv, /app/.venv, /sandbox/.cursor-server, /usr/lib/node_modules. Imported unchanged into another image the profile matches nothing: the catalog still advertises it, but the credential is never injected and the traffic is denied. The headers say so where it applies. Comments only; the profile schema has no documentation fields and all fifteen files still lint clean. Signed-off-by: Philippe Martin <phmartin@redhat.com> * test(server): seed unit-test state with the example provider profiles Server unit tests inherited the builtin + user source default from ServerState, so around a hundred and forty assertions about github, openai and the rest resolved against the compiled catalog. Point the test state at the user source alone and import the example profiles from providers/ into its store first, the way an operator would. Every one of those assertions keeps passing unchanged, which is the point: it proves the gateway behaves identically with an imported catalog before the default moves. Four tests asserted builtin-source semantics specifically. Under import-only the only profiles a gateway cannot edit are the ones a non-user source vends, so they now exercise a source-managed profile composed with the user source; the read-only guard they cover is the one that still applies to interceptor catalogs. The list test asserts the imported profile's user/platform identity instead of builtin with an empty scope. test_server_state_with_user_only_github_profile is gone: the default test state now is a user-only gateway with github imported. Signed-off-by: Philippe Martin <phmartin@redhat.com> * refactor(cli): resolve credential suggestions from the gateway catalog The --env credential warning scanned a profile table compiled into the CLI, so its suggestions described the binary rather than the gateway the user is talking to: a profile the gateway does not serve was suggested anyway, and an imported custom profile never was. Fetch the catalog from the connected gateway instead. The warning moves out of argument parsing and into sandbox create and sandbox template create, where a client already exists. A catalog fetch failure is not fatal — the warning degrades to its generic form rather than blocking sandbox creation. Extract the ListProviderProfiles paging loop from provider list-profiles into a shared fetch_provider_profile_catalog, which the profile-driven paths now share. Signed-off-by: Philippe Martin <phmartin@redhat.com> * refactor(cli)!: infer providers from the gateway catalog Command-to-provider inference went through a hardcoded alias table in openshell-providers, a second copy of the built-in catalog's identifiers. A custom profile could never be inferred no matter what binaries it declared, and the table drifted from the profiles it mirrored. Infer from the connected gateway's catalog instead: match the command's basename against each profile's ID and against the basenames of the binaries the profile authorizes. A profile that names /usr/bin/claude is the profile for running claude. The match must be unique — where several profiles claim a command, the user names one with --provider — and an empty catalog infers nothing. No command is special-cased. `binaries` is the operator's authorization statement, so a profile that declares a binary claims the command that runs it, whatever that binary is; narrowing that belongs in the profile rather than in a list compiled into the CLI, which could never cover an unbounded catalog anyway. Breaking: the retired aliases stop resolving, and commands the old table never listed can now infer. git, pip and uv are declared by the github and pypi example profiles, so they infer where they previously did not. Signed-off-by: Philippe Martin <phmartin@redhat.com> * refactor(providers,server)!: resolve provider profiles by exact ID normalize_provider_type was a hardcoded alias map — gh to github, claude to claude-code, vertex to google-vertex-ai — and a second copy of the built-in catalog's identifiers, independent of the YAML it mirrored. It let a provider type resolve to a profile the operator never named, and it made the built-in IDs behave as a reserved namespace. Remove it, along with detect_provider_from_command and the alias-normalizing ProviderRegistry::inject_env. A provider type now names a profile exactly: - the effective catalog resolves an ID or reports it absent, with no alias retry - plugins activate only for the ID of a profile the gateway resolved; a provider with no resolvable profile gets no plugin projection, instead of falling back to an alias guess - the CLI surfaces the gateway's not-found instead of retrying under an alias - telemetry buckets by profile ID, and the gitlab, opencode and outlook buckets go with the aliases that were their only source Breaking: `--type gh`, `--type claude` and the other aliases no longer resolve. Use the profile's own ID. Signed-off-by: Philippe Martin <phmartin@redhat.com> * feat(server): fail closed when a provider's profile is absent A sandbox composed from a provider whose profile the gateway cannot resolve started anyway: the credential and policy builders warned and skipped, so the sandbox came up carrying none of that provider's credentials or network policy. The operator learned about it later, as a denied connection or a missing environment variable, rather than as the configuration error it is. Check at the two composition boundaries — CreateSandbox and AttachSandboxProvider — right beside the catalog snapshot already taken there, and reject with a bounded diagnostic naming the provider, the profile it refers to, and the import command that supplies it. Scope-aware: a platform-scoped provider is told to import with --global. Read paths are untouched. ListProviders, GetProvider and profile export keep working so an operator can see and recover an affected provider, and the shared policy builders keep their warn-and-skip for the diagnostic paths that also reach them. Signed-off-by: Philippe Martin <phmartin@redhat.com> * refactor(server): scope the vendor base-URL pin to declared endpoints provider_profile_endpoints_are_active withheld a credential and the provider's policy layer when an openai or anthropic provider pointed its client somewhere other than the public vendor endpoint. The guard was keyed on profile.source == "builtin" and on those two profile IDs, so it protected only profiles OpenShell shipped. Once profiles are import-only no profile is ever builtin, and the guard would silently stop applying — including to an operator who imported providers/openai.yaml verbatim. Key it on the profile instead of on where the profile came from. A profile's endpoints are the boundary its credential is bound to, so if the provider configures a *_BASE_URL pointing at a host the profile does not declare, the profile no longer describes where that credential goes and is treated as endpointless. Host matching reuses the DNS-label-aware matcher in openshell-core, so wildcard endpoints such as Vertex's *-aiplatform.googleapis.com resolve correctly. A profile with no declared endpoints has no boundary to contradict, and a config value that names no host is not a redirect this can reason about. Both keep the profile active. The control now covers every endpoint-bearing profile, including an operator's own, and the last "builtin" string leaves the gateway. Signed-off-by: Philippe Martin <phmartin@redhat.com> * test(e2e): import example provider profiles during gateway bring-up The e2e suites create providers from github, openai, nvidia, claude-code and google-cloud, and the GitHub lane runs a real git clone through the profile's binary attribution. A gateway serves only the profiles an operator imported, so the lanes have to import them. Add e2e_import_example_provider_profiles to the shared bring-up helpers and call it from the Docker, Podman, Kubernetes and VM wrappers once the gateway is healthy and registered. It runs the same command the upgrade notes give operators, against the repository's own providers/ directory, so the lanes exercise the documented path rather than a test-only shortcut. Lands before the default changes: a same-ID user profile already shadows a built-in, so importing works today. Signed-off-by: Philippe Martin <phmartin@redhat.com> * refactor(providers)!: make provider profiles import-only OpenShell compiled fifteen provider profile YAML files into every release binary and selected them by default, so a fresh gateway published a catalog it was never configured with. Those profiles are not image-neutral: their binary selectors name paths that exist in the OpenShell Community image, so changing the sandbox image could make a profile inert while the catalog still advertised it. Every endpoint, credential name and binary path in providers/ was also effectively part of the 0.1.0 public contract. A gateway's catalog is now exactly what an operator imported: - providers/*.yaml is no longer include_str!'d, and builtin_profiles() is gone from the openshell-providers API - the default provider_profile_sources is [{ type = "user" }], and a gateway with nothing imported reaches ready and serves an empty catalog — an empty catalog is a valid state, not a startup failure - the builtin source type is removed from the configuration schema, and a gateway.toml that still names it is rejected at parse time with the import command rather than an unknown-variant error - nothing reserves the canonical identifiers any more, so github, pypi, anthropic and the rest import at their own IDs and the imported profile is the only definition for that ID; the static_fallback that kept a shipped definition resident behind an imported one is gone with the source that produced it A collision between an interceptor-vended profile and an imported one still fails closed, which is the behavior that shadowing quietly bypassed for built-ins. Operators upgrading should export the profiles their deployment relies on before upgrading, or copy them from the providers/ directory of the matching release tag, then import them at the scope their providers use. Signed-off-by: Philippe Martin <phmartin@redhat.com> * docs(providers): document import-only provider profiles Rewrite the provider profile documentation around a catalog the operator builds rather than one the gateway ships. - gateway-config: the default is [{ type = "user" }], the builtin source type is gone and rejected at startup, and an empty catalog is a valid ready state - profiles: replace the built-in profile table and its shadowing semantics with the import workflow, and say plainly that a profile whose binary paths do not match the image is inert - manage-providers: replace the two fixed provider-type tables with `provider list-profiles`, and describe catalog-driven command inference - inference-routing: drop the "still loads built-in profiles for compatibility" transition text; its migration walkthrough is now the normal path - quickstart and the Docker Compose, GitHub, AWS, Google Cloud and Vertex tutorials: import the profile before creating the provider, since nothing resolves without it - release notes: a 0.1.0 migration section covering export-before-upgrade, importing from the release tag's providers/ directory, what happens to a provider whose profile is missing, and the removal of the legacy type aliases - README, architecture, the openshell-cli and debug-openshell-cluster skills, and the governance interceptor example follow the same change; the example's smoke assertion now imports a profile to prove an authoritative interceptor hides it Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(cli,e2e): distinguish catalog failures and authenticate OIDC seeding Addresses two findings from review (GATOR-c1bd9867-02 and -03). The provider profile catalog lookup collapsed its error into an empty catalog, so an authorization, availability or transport failure on ListProviderProfiles was indistinguishable from a gateway that genuinely has no profiles. Command inference then resolved nothing and the sandbox was created without the provider it needed, deferring the failure to the workload. Keep the lookup's outcome instead of discarding it. The credential warning is advisory and still degrades to its generic form, but inference now consults the catalog only when there is a command to resolve and surfaces the lookup failure when there is, naming the gateway and pointing at explicit --provider selection. Two regression tests cover it through a fake gateway whose ListProviderProfiles returns UNAVAILABLE while every other RPC succeeds: sandbox creation fails without sending a provider-less create request, and a sandbox with no trailing command still succeeds because it needs no catalog. The e2e profile seeding also ran unauthenticated in the OIDC lanes. Those lanes deliberately skip gateway registration and start the gateway without a TLS client CA, so no mTLS identity exists and no token has been acquired when the import runs; the wrapper exited during setup. Skipping the import is not sufficient because the provider tests now require the claude-code profile. Add e2e_register_oidc_admin_session, which mints an administrator token with Keycloak's password grant — the same grant the OIDC test helpers use — and writes the gateway metadata and token bundle that an interactive login would have stored, so the import runs as an authenticated administrator. The mTLS lanes keep the direct import unchanged. Both affected wrappers are covered: with-podman-gateway.sh had the same defect as with-docker-gateway.sh. Signed-off-by: Philippe Martin <phmartin@redhat.com> * refactor(cli)!: remove command-derived provider attachment Addresses GATOR-c1bd9867-01. A profile's `binaries` list authorizes a binary to reach that profile's endpoints. It is not a statement that running the binary asks for the provider, and reading it as attachment intent let a command silently gain provider authority: the aws-s3 example declares /bin/bash, so `sandbox create -- bash` resolved an existing aws-s3 provider and attached its credential-backed capability and network policy to the shell and its descendants. Attachment of an already-created provider never prompted, so the confirmation flow did not guard it. The distinction that would make inference sound — whether a declared binary is a profile's client or merely a permitted runtime — cannot be expressed: NetworkBinary carries only a path, and the field that encoded it was removed in 0.1.0. Any substitute is a guess. Restricting the guess to a unique claimant does not help, because uniqueness measures how sparse the catalog is rather than what the user intended, and a compiled list of "generic" commands could never cover an unbounded operator catalog. Remove trailing-command inference rather than approximate it. A provider is attached only when named with --provider, which still creates a missing provider from local discovery when the name matches an imported profile ID. Sandboxes with no providers remain a normal, fully supported state. Removing inference also settles GATOR-c1bd9867-02: with no consumer deriving authority from the catalog, its only remaining use is the advisory credential warning, so a failed lookup degrades that warning instead of blocking creation. The regression test now asserts that an unreachable catalog still creates the sandbox and attaches nothing. While repurposing the deduplication test, a pre-existing defect surfaced: repeating a name in --provider auto-created it twice and the second attempt failed with "provider already exists", because the explicit pass never consulted the set of names it had already handled. Guard it. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(e2e): load the OIDC token and trust the gateway when seeding profiles Addresses the carried finding GATOR-c1bd9867-03. e2e_register_oidc_admin_session established a session the CLI never used. The CLI decides whether to load a stored bearer token by matching on the auth_mode field of the gateway metadata alone; the helper omitted that field, so the metadata fell through to the default arm and oidc_token.json stayed on disk unread. The profile import went out unauthenticated exactly as it had before the helper existed. The helper also installed no trust anchor, leaving the CLI unable to verify the gateway's self-signed serving certificate. Write auth_mode = "oidc" so the stored token is loaded, and install the CA at <gateway>/mtls/ca.crt. Only the CA is installed: with no client certificate or key on disk the CLI falls back to CA-only server verification and authenticates with the bearer token, which is what these lanes need because they start the gateway without --tls-client-ca. Certificate verification stays on; no insecure transport override is introduced. Assert the session before anything depends on it. ListProviderProfiles is annotated auth_mode: "bearer", so it cannot succeed unless the token was loaded and accepted. The helper now fails at that point, naming the gateway config directory and echoing the CLI output, rather than letting the defect surface later as an opaque profile import error. Both affected wrappers pass the PKI directory and CLI binary the helper needs; the mTLS lanes are untouched. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(e2e): request the OpenShell scopes when minting the admin token The OIDC lanes failed at the session assertion added in b25dd0298 with PERMISSION_DENIED and "scope 'provider:read' required". Authentication was working -- the gateway logged ListProviderProfiles at gRPC status 7, which it can only reach once the bearer token has been loaded and accepted. The token simply carried no OpenShell scope. sandbox:*, provider:*, config:*, workspace:* and openshell:all are optional client scopes on the openshell-cli client in scripts/keycloak-realm.json, so Keycloak mints them only when the request asks for them. The password grant here asked for nothing, leaving the realm defaults (openid, profile, email, roles, web-origins, acr) and an access token that authorizes no RPC. Request "openid openshell:all", as e2e/python/oidc/oidc_auth_test.py already does for its administrator tokens. openshell:all is SCOPE_ALL in crates/openshell-server/src/auth/authz.rs, so one scope covers the setup calls without enumerating them. Verified against the realm: the token goes from "email profile openid" to "email profile openshell:all openid". Record the same scopes in metadata.json. The helper writes the bundle openshell gateway login would have stored, and oidc_scopes is the field that login path reads back, so leaving it out would misdescribe the stored token. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(e2e): seed provider profiles once per VM lane gateway The VM lanes failed on the second test target with "custom provider profile 'anthropic' already exists" for all fifteen example profiles, and the import exited non-zero. e2e_import_example_provider_profiles was called from run_e2e_test, so it ran once per target -- four times against one long-lived gateway. That was harmless while import overwrote silently, but profiles are now import-only: ImportProviderProfiles is create-only and reports an existing id as an error-severity diagnostic, with no overwrite flag on the request. The first import therefore succeeds and every later one fails. Hoist the call to just after the conformance run, which is where the docker, podman and kube lanes already seed their catalogs. The profiles persist for the gateway's lifetime, so every target still finds them, and the E2E_TEST_OVERRIDE path is covered by the same single call. Signed-off-by: Philippe Martin <phmartin@redhat.com> * test(e2e): name the provider explicitly in the OIDC workspace-user test user_can_create_sandbox_with_inferred_provider_command reached the gateway once the admin session was fixed, and then failed: it asserted a missing-provider error, but the sandbox was created and died provisioning with "failed to spawn sandbox entrypoint process 'claude-code'". The test drove provider resolution by passing claude-code as the trailing command and relying on the CLI to infer the provider type from it. That inference is what this branch removed, so the trailing word is now nothing but an entrypoint, and the image has no such binary. The regression the test guards is not inference itself: it is that a workspace user resolving a provider is not gated behind Platform Admin. Name the provider with --provider, the only remaining way to attach one. The CLI still has to fetch the claude-code profile before it can auto-create the provider, so the lookup a workspace user must be allowed to make still happens, and auto-creation still stops at the non-interactive branch with "missing required provider". Both assertions therefore keep their meaning. Rename the test and rework its comments to describe what it now exercises; the old name would otherwise outlive the behavior it was named for. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(providers): reconcile import-only profiles with main Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> --------- Signed-off-by: Philippe Martin <phmartin@redhat.com> Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> Co-authored-by: John Myers <9696606+johntmyers@users.noreply.github.com> |
||
|
|
769273f096 |
feat(middleware): add a hook to inspect HTTP responses (#3074)
* feat(middleware): implement HTTP response processing Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(middleware): allow one-byte response stream units Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * feat(examples): separate content guard from middleware protocol demos Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(middleware): address HTTP response review findings Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * refactor(examples): defer protocol demo to a separate PR Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * refactor(middleware): keep response body timeout local Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(middleware): honor fail-open for unrepresentable responses Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(middleware): trim runtime docs and extract troubleshooting reference Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * chore(middleware): split build fix and simplify test and skill guidance Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * refactor(middleware): isolate HTTP response processing Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(middleware): hide response credential headers Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(middleware): end invalid preflight streams Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(middleware): remove hard-wrapped prose Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(middleware): use protobuf request timeouts Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> --------- Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> |
||
|
|
9b9f7905e1 |
fix(dev): extract Docker sandbox runtime from sandbox image (#3422)
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> |
||
|
|
58b5f8f976 |
feat(prover): add standalone policy boundary checker (#3289)
* feat(prover): add standalone policy maximum checker Signed-off-by: Johnny Greco <jogreco@nvidia.com> * refactor(prover): simplify check scope schema Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(prover): align containment and cancellation with runtime Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(prover): stabilize containment checks in CI Signed-off-by: Johnny Greco <jogreco@nvidia.com> * test(prover): avoid solver in fast-path guard test Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(prover): align string containment with runtime Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(prover): reject ambiguous z3 string escapes Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(prover): align containment with runtime boundaries Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(prover-cli): harden cancellation and invalid input Signed-off-by: Johnny Greco <jogreco@nvidia.com> * feat(packaging): install policy prover with OpenShell Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(prover): clarify installation and check results Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(build): describe prover distribution directly Signed-off-by: Johnny Greco <jogreco@nvidia.com> * refactor(prover): rename maximum policy to boundary Signed-off-by: Johnny Greco <jogreco@nvidia.com> * refactor(prover): localize fail-closed validation Signed-off-by: Johnny Greco <jogreco@nvidia.com> * test(prover): cover fail-closed CLI surfaces Signed-off-by: Johnny Greco <jogreco@nvidia.com> * test(prover): allow CI load for REST solver proof Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(prover): use canonical policy schema for containment Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(prover): preserve uncertainty for runtime binary globs Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(prover): bound policy validation work Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(prover): document validation resource limits Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(prover): make containment API extensible Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(prover): define containment API contract Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(ci): integrate prover with consolidated builds Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(ci): declare release packaging dependency Signed-off-by: Johnny Greco <jogreco@nvidia.com> --------- Signed-off-by: Johnny Greco <jogreco@nvidia.com> |
||
|
|
f419b9c1df |
fix(python): use portable empty-array expansion for macOS Bash 3.x (#3413)
macOS ships Bash 3.2 which errors on `"${arr[@]}"` when the array is
empty under `set -u`. Use `${arr[@]+"${arr[@]}"}` to safely expand
VERIFY_ARGS only when it has elements.
Signed-off-by: Florent Benoit <fbenoit@redhat.com>
|
||
|
|
9c41f057c3 |
fix(python): stabilize development tasks under jj (#3354)
setuptools-scm can select jj's Git-bridge dev tag when uv resolves the local package. Pin a development-only version for every Python task that performs that resolution. The 0.0.0 value is valid local editable-package metadata for development tasks; it is neither a release version nor a Git tag. Wheel builds remain unpinned and derive their published version from the actual release tag. Signed-off-by: Kris Hicks <khicks@nvidia.com> |
||
|
|
b3e4ad4579 |
fix(ci): repair RFC 0012 post-merge checks (#3360)
* fix(ci): lock nextest for Windows ARM64 Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): preserve GPU access in sandbox runtime Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(windows): restore host proxy identity binding Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(windows): satisfy cross-platform network lint Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(windows): satisfy MXC lint Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(windows): exclude Unix supervisor runtime Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(mxc): satisfy Windows test lint Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(windows): escape proxy test policy paths Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): accept admitted GPU runtime groups Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(docker): serialize proxy pipeline scenarios Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
314c73343c |
fix(ci): restore prebuilt Z3 on Windows (#3353)
Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com> |
||
|
|
c1f2e7189f |
feat(isolation): implement the RFC 0012 sandbox architecture (#2942)
* feat(isolation): add RFC 0012 backend contract Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> * refactor(isolation): name the interface crate explicitly Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): expose trusted host gateway Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(agents): inventory the MXC driver Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): add mediated DNS transport Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): tighten interface error and digest contracts Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(isolation): remove unrelated driver inventory Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): define capability-free launch contract Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): seal confirmed boundary state Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): validate confirmation for external backend implementations Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): clarify mediated DNS identity Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(isolation): generalize loopback connector Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(isolation): unify typed network mediation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): bind launches to sandbox sessions Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(mxc): initialize extended sandbox status Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): add boundary protocol and Linux primitives Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): harden signals and separate process status from transport Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): validate remote confirmation through public contract Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): validate wire state and propagate snapshot failures Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(isolation): import owned agent specification explicitly Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(isolation): describe mediated DNS channel Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): bound mediation attach without nested retries Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(isolation): generalize loopback protocol Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): add transport-neutral session authentication Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(isolation): separate sandbox backend protocol Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): harden runtime boundary controls Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): add terminal boundary operation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): split supervisor and sandbox runtimes Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): harden boundary isolation and lifecycle ownership Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): reject private root redirects and adopt typed errors Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): preserve accept thread ownership on musl Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(sandbox): isolate credential probes from filtered threads Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): return retained exec exit status to independent waiters Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): bound network mediation and preserve socket authorization Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): bound control admission and retire stale mediation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * ci(e2e): select migrated drivers per stack layer Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(sandbox): implement loopback connector Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): authenticate the Sandbox Protocol Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(supervisor): rotate launch-scoped authentication Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(sandbox): consume dedicated backend crate Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(sandbox): align topology session fixture Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): align projected bootstrap bundle Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(auth): validate refreshed credentials before rotation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): fail closed across supervisor disconnects Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): repair rebased sandbox CI Signed-off-by: Drew Newberry <anewberry@nvidia.com> * build(runtime): publish separate sandbox and supervisor images Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(config): configure the sandbox runtime image Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(ci): validate sandbox binary linkage Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(isolation): use backend and runtime terminology Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(sandbox): use a scratch runtime image Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(ci): refresh schema and dependency policy Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): bind reconnects to supervisor process Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs: align runtime split operational guidance Signed-off-by: Drew Newberry <anewberry@nvidia.com> * chore(security): document Kubernetes runtime RBAC Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): enforce runtime lifecycle invariants Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(compute): identify sandbox start generations Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(server): restore sandbox launch sessions Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): support authenticated runtime replacement Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(auth): bind sandbox session successors Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(auth): retry pending sandbox successors Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(vm): run the supervisor outside the guest workload Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(vm): use sandbox backend protocol Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): repair rebase integration Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): use unified build toolchain Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(vm): use sandbox runtime terminology Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(vm): own guest network bootstrap Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): expose guest init version Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): select native supervisor artifacts Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): guard guest init Linux symbols Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): scope Linux test imports Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): avoid guest interface casts Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): reconcile admitted sandbox identity Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): share resolved sandbox identity Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): surface host supervisor failures Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): include guest logs on supervisor exit Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): rotate and clean runtime generations Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): make sandbox starts generation-aware Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): rotate restored sandbox sessions Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(vm): keep shared paths in the base layer Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): bind sandbox session lineage Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(docker): isolate workloads behind the host supervisor Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(docker): rotate launch-scoped authentication Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(docker): use sandbox backend protocol Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(docker): use host networking for supervisor Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): preserve host gateway alias resolution Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(docker): use separate sandbox and supervisor images Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): restore startup validation after rebase Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(docker): name the sandbox runtime directly Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): narrow supervisor CA runtime storage Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): close companion isolation gaps Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(docker): align mediated network expectations Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(docker): exercise mediated network paths Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): attach supervisor to managed network Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): defer supervisor recovery until gateway is ready Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): make sandbox starts generation-aware Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): rotate restored sandbox sessions Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): preserve workloads during session rotation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(docker): remove unrelated configuration RFC changes Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): bind sandbox session lineage Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(kubernetes): add proxy-pod isolation topology Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(kubernetes): use sandbox backend protocol Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): use stable sandbox service authority Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(kubernetes): split sandbox and supervisor images Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): adapt proxy pods to current runtime APIs Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(kubernetes): describe the single runtime placement Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(kubernetes): simplify sandbox orchestration Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): validate deployment prerequisites Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): update Trivy Helm profile inventory Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(kubernetes): update Trivy scan inventory count Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): reuse preloaded runtime images in e2e Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): type and clean runtime resources Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): make sandbox restarts recoverable Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): rotate restored sandbox sessions Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): preserve supervisor egress Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): bind sandbox session lineage Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(podman): adopt isolated sandbox and supervisor containers Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): stage bootstrap archives at named volume destinations Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(podman): rotate launch-scoped authentication Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(podman): use sandbox backend protocol Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(podman): use host networking for supervisor Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(podman): split sandbox and supervisor images Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): repair rebase integration Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(podman): name the sandbox runtime directly Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): provision supervisor CA runtime storage Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): address isolation review findings Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): inspect Debian supervisor provenance Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): use libpod-compatible tmpfs options Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): bind verified sandbox runtime binary Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): provide external driver data directory Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): start sandbox before joining user namespace Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): separate supervisor user namespace Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): make sandbox starts generation-aware Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): rotate restored sandbox sessions Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): bind sandbox session lineage Signed-off-by: Drew Newberry <anewberry@nvidia.com> * perf(isolation): add TCP and DNS benchmark harnesses Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(perf): align benchmark timing and supported protocols Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(perf): report TCP benchmark metrics accurately Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(perf): cancel failed worker startup Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): build matching local supervisor image Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): make local sandbox smoke test runnable Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): wire local sandbox runtime image Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): narrow sandbox service RBAC Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(ci): validate split runtime artifacts Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): harden runtime session handling Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(supervisor): add standalone network proxy role Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(rfc): remove implementation companion notes Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(vm): standardize runtime release name Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): pin renamed runtime artifacts Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(auth): persist sandbox runtime identity Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(runtime): restore branch validation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): reconcile main after rebase Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(network): close unframed HTTP 1.0 responses Signed-off-by: Drew Newberry <anewberry@nvidia.com> * chore(isolation): preserve upstream OCSF updates Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(security): close credential and TLS replay paths Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(auth): make sandbox refresh retries idempotent Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
d95bab5436 |
fix(ci): upstream Windows SDK validation support (#3327)
* fix(ci): upstream Windows SDK validation support Port remaining Windows validation tooling from GitLab independently of the combined MXC runtime port. Preserve locked SDK dependencies and guard temporary launcher cleanup. Co-authored-by: Shailendra Singh <shailendras@nvidia.com> Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com> * fix(ci): install native TypeScript dependencies before validation Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com> --------- Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com> Co-authored-by: Shailendra Singh <shailendras@nvidia.com> |
||
|
|
b799fccb8b |
fix(auth): harden OIDC trust root retrieval (#3332)
* fix(auth): harden OIDC trust root retrieval Signed-off-by: Mrunal Patel <mrunalp@gmail.com> * fix(e2e): pass OIDC HTTP acknowledgement value Signed-off-by: Mrunal Patel <mrunalp@gmail.com> --------- Signed-off-by: Mrunal Patel <mrunalp@gmail.com> |
||
|
|
26f2f96393 |
feat(mxc): add Windows host proxy for MXC sandbox network egress (#3163)
* Implement Windows host proxy integration and update dependencies for OpenShell
* Update README and gateway config to clarify egress proxy address handling and allocation
* Refactor ProxyIdentityMode to return Result for static_binary and add tests for binary path and SHA256 hash
* Enhance platform_hosts_path for Windows to use SystemRoot and improve error handling for hosts file reading
* Refactor FileFingerprint to use Option for mtime and ctime, simplifying metadata handling
* Add conditional compilation for Windows host module
* add unit tests for OPA policy evaluation and identity handling
* remove openshell-supervisor-network from unsupported driver package test exclusion list
* feat(mxc): enable host proxy TLS state generation
Generate per-sandbox TLS state for the MXC host proxy so HTTPS L7 enforcement can use the same MITM path as Linux. Grant generated CA material to the MXC process and inject standard trust env vars, while matching Linux behavior by disabling TLS termination on CA setup failure and relying on proxy fail-closed handling.
* fix(docs): remove outdated notes on governed egress from docs
* fix(tests): update TLS environment variable paths to use temporary directory
* fix(examples): make run-mxc-e2e harness correct and orphan-free
The MXC e2e harness never actually exercised the fs scenarios: it started
the gateway once and patched agent_command per scenario AFTERWARDS, so the
running gateway kept launching the default demo agent (not shipped in the
kit) and every fs scenario failed with CreateProcessW error:2. It also
scored on the `sandbox create` exit code (non-zero due to the harmless
interactive attach), wrote sandbox records to the persistent gateway DB
(leaving orphans that collided on later runs), and its deny scenarios never
proved denial.
Changes:
- Start a FRESH gateway per scenario so each scenario's agent_command is
actually loaded (root cause of CreateProcessW error:2).
- Score by on-disk artifact / expected outcome, not `sandbox create` exit.
- Real deny assertions: a control write to a granted path must succeed
(proves the agent ran) while the denied write must be absent. fs-empty
probes an ungranted out-of-share path (share_dir is mapped rw by design).
- Run the gateway on an ephemeral in-memory DB (sqlite::memory:) so the
harness never writes to the persistent store and cannot leave orphan
sandbox records; also use unique per-run sandbox names + pre-delete.
- Fix the process_container probe: use a real cwd + absolute cmd.exe
(canonical wxc-exec does not expand %TEMP% -> 0x8007010B).
- Fix summary counts (@() so a single FAIL is counted and exit is non-zero).
Verified PASS=4 FAIL=0 on 7F203-MXC-003 (no BaseContainer velocity keys)
using a canonical wxc-exec build (AppContainer fallback).
Signed-off-by: Akber Raza <akberr@nvidia.com>
* fix(e2e): probe timeout is milliseconds (10ms->30000ms)
MXC process.timeout is wall-clock ms (wire.rs). The 10 value meant 10ms,
which the base-container tier (7F203-MXC-001/.181) enforced strictly and
timed the probe out. AppContainer path (.18/-003) happened to slip under
it. Bump to 30000ms so the process_container preflight probe is reliable
across both tiers.
Signed-off-by: Akber Raza <akberr@nvidia.com>
* fix(mxc): use native paths in real runtime probes
Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>
* fix(mxc): make processcontainer work with mxc-latest-released wxc-exec
Three fixes to support the release wxc-exec binary (BaseContainer dispatcher)
in addition to mxc-fixes-env-vars:
1. Seed process env from host (driver.rs)
ProcessContainer starts with a completely blank environment -- no PATH,
SystemRoot, or anything. Seed the process env from the gateway host
environment so the agent binary can locate DLLs and run. Skip internal
Windows drive-letter variables (keys starting with '=') which cause
CreateProcessW to return ERROR_ENVVAR_NOT_FOUND. User agent_env entries
and TLS CA vars are applied as overrides on top of the host env.
2. Remove TLS readonly_paths grant (driver.rs)
The release wxc-exec (BaseContainer dispatcher) requires write-DAC
permission on every path in readonly_paths to set up AppContainer ACLs.
Adding the proxy's temp TLS directory caused a DACL error and exit -1.
The CA cert paths remain available to the agent via TLS env vars.
3. Remove allowedHosts from network JSON (mxc.rs)
The release wxc-exec rejects network.allowedHosts / network.blockedHosts
on Windows with "not yet supported". Removed the loopback exemption
attempt (127.0.0.1, ::1, localhost) from the network section.
Intra-container loopback works natively in the release binary without
it -- the spawner can connect to the server at 127.0.0.1:22000 directly.
Additional changes:
- mxc-ws-agent.rs: add relay-debug.txt error capture and relay-ready.txt
marker for reliable timing of host client connections.
- mxc-ws-gateway.toml: debug = true for JSON config dump during diagnosis.
- run-ws-agent-test.ps1: default port changed to 17670 (gateway default);
relay-ready.txt polling before ws-echo to avoid connecting before the
spawner has established the proxy bridge.
Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>
* fix(e2e): address CodeRabbit review on run-mxc-e2e.ps1 (MR !46)
Four robustness/correctness fixes from CodeRabbit:
1. Start-Gw: kill the spawned gateway before the "did not start within 30s"
throw. If the process is alive but never binds the port, $gw is not yet
assigned in the caller, so the finally block cannot reap it -> orphan
gateway holding the port for the next run.
2. create-fail scoring: a non-zero `sandbox create` exit alone is not proof
of a policy rejection (gateway-registration/transport/fixture errors also
exit non-zero and would false-pass). PASS now requires a genuine
rejection signal (network / invalid_argument / network_policies) AND that
it is not an infrastructure failure; other non-zero exits go to FAIL with
output captured.
3. deny scenarios (ControlTarget path): snapshot the deny target AFTER
Wait-File lands the control artifact, so a late denied write (enforcement
regression racing the control write) can no longer be recorded as PASS.
4. -KeepRunning: break out of the scenario loop after the first scenario so
a later scenario does not start a second gateway on the same port
(previously a reliable port collision instead of a usable debug mode).
Re-verified PASS=4 FAIL=0 on both boxes (7F203-MXC-001 base-container and
7F203-MXC-003 AppContainer fallback); network-policy-rejected correctly
scores as "policy rejection".
Signed-off-by: Akber Raza <akberr@nvidia.com>
* feat(mxc-e2e): collect run-mxc-e2e output into a results bundle
Mirror the sibling run-*.ps1 scripts by collecting every run's logs into a
timestamped results-e2e-<stamp>\ folder and zipping it. The bundle contains the
console transcript, per-scenario gateway stdout/stderr, the exact TOML rendered
for each scenario, the policy fixture used, and a summary.txt with the verdict
table.
Per-scenario gateway logs now land in gateway.<scenario>.log/.err.log inside the
bundle instead of a single fixed gateway.e2e.log in the script directory.
Wrap pre-flight, mode setup, scenario definitions, and the scenario loop in a
single try/catch/finally so the finally always writes the summary, stops the
transcript, and zips the bundle -- even on a pre-flight failure. The existing
per-scenario gateway-cleanup try/finally stays nested inside. All scenario
logic, scoring rules, and comments are preserved.
Signed-off-by: Prashant S Khodade <pkhodade@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>
* fix(mxc-e2e): address CodeRabbit review on run-mxc-e2e.ps1
- Require -Scenario when -KeepRunning: the loop breaks after the first
scenario, so a full-suite run would execute only one scenario yet still
report the suite as PASS. Fail fast so a partial run can't be mislabeled
complete.
- Start-Transcript now runs inside the guarded try block with a
$transcriptStarted flag; Stop-Transcript is only called when it actually
started, so a Start-Transcript failure still yields the results bundle.
- Wrap the -Scenario filter in @() so a single exact match stays an array
(reliable .Count and a proper array for the scenario loop on PS 5.1).
Signed-off-by: Prashant S Khodade <pkhodade@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>
* fix(examples): pass gateway config via OPENSHELL_GATEWAY_CONFIG for spaced paths
Start-Process -ArgumentList does not quote array elements, so launching the
gateway with a bare --config <path> token split on any space in the install
path (e.g. C:\Users\First Last\...), and clap rejected the fragment with
'unrecognized subcommand'. Every MXC example launcher that started the gateway
hit this when the kit was unzipped under a path containing a space.
Pass the config path through the OPENSHELL_GATEWAY_CONFIG env var (which the
gateway already reads via clap) and drop the --config token. Env vars carry
spaces safely.
Affected: run-ocsf-audit, run-mxc-e2e, run-demo, run-inference-test,
run-ollama-test. run-mtls-test was not affected (its launch passes no config
path). Root-caused and fix-verified on 7F203-MXC-003 from a spaced path.
Signed-off-by: Akber Raza <akberr@nvidia.com>
* fix(run-mxc-e2e): improve scoring logic and enhance command execution handling
* fix(mxc): reconcile proxy support after rebase
Restore the proxy-enabled OCSF audit example removed by
|
||
|
|
5b9daab935 |
fix(ci): restore mise run ci on macOS (#3294)
- Replace BSD-incompatible in-place sed calls with portable temp-file rewrites. - Remove test-only shell interception and capture generated gateway config directly. - Allow parity tests to use supplied supervisor binaries without resolving a Linux target. - Normalize temporary-directory paths and use portable RPM config installation. - Set a valid setuptools-scm version for Python protobuf generation in Jujutsu checkouts. Signed-off-by: Kris Hicks <khicks@nvidia.com> |
||
|
|
5b57f0d154 |
fix(ci): restore Windows test portability (#3288)
* fix(ci): restore Windows test portability Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(tasks): skip Unix lockfile check on Windows Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(tasks): use buf shim on Windows Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> --------- Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> |
||
|
|
b92620e838 |
ci(rust): reject stale Cargo lockfiles (#3227)
* ci(rust): reject stale Cargo lockfiles Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(ci): clarify lockfile validation policy Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * refactor(ci): structure and test Cargo lockfile validation Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * chore(ci): remove lockfile validator regression tests Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * ci(rust): complete locked validation and lint examples Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * ci(rust): check lockfile diffs after validation Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(build): shorten lockfile validation notes Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * ci(rust): skip lockfile check after failures Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> --------- Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> |
||
|
|
02b664bb0d |
refactor(config): normalize and enforce gateway schema v2 (#2814)
* refactor(config): normalize compute driver field names Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * refactor(config): introduce canonical gateway fields Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * refactor(config): enforce gateway schema version 2 Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): preserve compute driver runtime guarantees Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): address schema v2 review regressions Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): complete schema v2 migration safeguards Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(config): expand schema v2 regression coverage Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(config): add schema v2 parity manifest Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): correct parity manifest inventory Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * docs(config): record schema v2 intentional changes Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * docs(config): disposition schema v2 parity gaps Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): add dual schema parity harness Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): establish compute lifecycle parity baseline Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): preserve gateway option compatibility Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): record gateway option parity Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * docs(config): close gateway-wide parity gaps Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(podman): apply configured pids limit Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): validate Podman option parity Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): add Kubernetes option parity harness Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): record Kubernetes option parity Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): disposition VM parity lanes Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): add external driver parity lane Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(e2e): preserve external driver pull policy Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): attest parity artifacts and launches Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): require clean parity build sources Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): bind parity runtime artifacts Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(e2e): use isolated supervisor tags Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(e2e): qualify parity image tags Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(e2e): serve parity supervisor locally Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): isolate parity podman services Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): harden parity evidence provenance Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): pin parity sandbox artifacts Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): attest parity runtime inputs Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): bind parity runtime evidence Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): record compute boundary parity Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): disposition cross-cutting parity lanes Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(packaging): preflight gateway config upgrades Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): preserve rebase integration guarantees Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(ci): isolate temporary git signing config Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): update remaining schema v2 consumers Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(ci): provide e2fs tools to VM tests Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): align preflight with gateway startup Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(vm): preserve rootfs tar configuration Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * chore(config): adopt duration unit constructors Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(packaging): preflight RPM gateway config Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): address driver review findings Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(e2e): require fresh semantic parity evidence Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(docker): update tests for renamed sandbox label Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(gateway): preserve selective driver coverage after rebase Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> Signed-off-by: Drew Newberry <anewberry@nvidia.com> Co-authored-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
38f2aef930 |
feat(gateway): support selective compute driver builds (#3118)
* feat(gateway): support selective compute driver builds Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(gateway): support selective Windows MXC builds Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
ddc8bba967 |
ci(windows): add Windows MSVC CI jobs (#2738)
* fix(ci): preserve Windows Rust build cache Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(ci): invalidate empty Windows caches Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * perf(ci): cache Windows builds with sccache Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * ci(windows): restore target directory caching Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * perf(ci): use prebuilt Z3 on Windows Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * perf(ci): layer sccache on Windows target cache Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * ci(windows): split PR checks from main validation Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * ci(windows): separate checks builds and cache seeding Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(ci): simplify Windows build dependency Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(ci): rely on Windows job dependency status Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(ci): use valid opt-in Windows ARM runner Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * ci(windows): keep ARM64 validation local Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(ci): install Clippy for Windows validation Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * ci(windows): focus platform lint coverage Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(licenses): explain bzip2 allowance Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * ci(windows): simplify workflow name Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(windows): allow async platform stub Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(network): make file fingerprints portable Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * ci(windows): lint supported deliverables Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * ci(windows): allow platform-gated lint Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * chore(ci): align Windows cache action with main Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(ci): align Windows validation with prerequisites Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * ci(windows): pin Rust toolchain action Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(ci): use enterprise-approved Windows actions Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(windows): restore strict MSVC validation Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * ci(windows): run Rust tests with nextest Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(ci): normalize nextest lock provenance Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * ci(windows): add native arm64 validation Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(ci): lock nextest for Windows ARM64 Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(windows): resolve duplicate MXC authentication method Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * test(conformance): use native absolute paths on Windows Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(windows): address MSVC review feedback Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * ci(windows): simplify cache key names Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(ci): isolate Windows Rust toolchains for stable caches Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(ci): configure Rustup home in runner setup Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * ci(windows): surface sccache server write diagnostics Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * ci(windows): remove temporary cache diagnostics Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * ci(windows): address review feedback Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(windows): reconcile merged driver capabilities Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(deps): preserve AWS-LC-only lockfile Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * chore(deps): allow z3 prebuilt TLS wrapper Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> --------- Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> |
||
|
|
ce25acca5a |
fix(build): honor Cargo target directory when staging binaries (#3262)
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> |