Files
50230616d5 refactor(runtime): retire Community image dependencies (#3386)
* feat(sandbox): default to official Alpine sandbox image

default_sandbox_image() now returns docker.io/library/alpine:3.22, a generic
version-qualified official image, so a fresh install no longer depends on the
community sandbox image catalog. All compute drivers (docker, podman,
kubernetes, vm) inherit this fallback.

Part of #3116.

Signed-off-by: Akram
Signed-off-by: Akram <akram.benaissi@gmail.com>

* feat(deploy): default deployment configs to the official Alpine sandbox image

Update the shared gateway default_image, Helm chart values, the standalone
Kubernetes manifest, and the dev gateway task scripts to use
docker.io/library/alpine:3.22 instead of the community base image, consistent
with default_sandbox_image(). GPU e2e image-build base is left unchanged (CUDA
needs a glibc base).

Part of #3116.

Signed-off-by: Akram
Signed-off-by: Akram <akram.benaissi@gmail.com>

* feat(driver): default to numeric non-root identity for USER-less images

With the default sandbox image now Alpine, images that declare no OCI USER
must start instead of being rejected. When the image declares no USER and
the policy requests none, the Podman and Docker drivers now supply a numeric
non-root identity (DEFAULT_SANDBOX_UID/GID = 1000) instead of rejecting,
matching the numeric-identity behavior of the Kubernetes and VM drivers. The
supervisor's resolved-identity path runs the sandbox as a synthesized
non-root account without the account existing in the image. Images that
declare a USER keep the OCI resolution path unchanged.

Part of #3116.

Signed-off-by: Akram <akram.benaissi@gmail.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(conformance): use Alpine workload image

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* refactor(policy): drop community image /app path from default policy

The restrictive default policy granted read-only access to /app, a directory
that only existed in the community base image. A generic Alpine default has no
/app, so remove it. Landlock best-effort already ignores absent paths; this
just stops advertising a community-specific layout in the default.

Part of #3116.

Signed-off-by: Akram
Signed-off-by: Akram <akram.benaissi@gmail.com>

* docs(config): document Alpine default images

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(podman): report early sandbox termination

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(podman): initialize rootless workspace ownership

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(sandbox): qualify NVIDIA Ubuntu default

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): initialize rootful default workspace

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(sftp): add native sandbox adapter

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sftp): gate runtime helper support to Linux

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sftp): support standard OpenSSH file operations

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sftp): harden rename and special file handling

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(runtime): remove community image dependencies

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(e2e): build provider readiness tool fixture

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(e2e): use a dedicated Noble fixture for Docker tests

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Akram
Signed-off-by: Akram <akram.benaissi@gmail.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Evan Lezar <elezar@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
2026-09-22 14:43:51 +02:00
..

GPU workload images

This directory defines workload test images currently used by the OpenShell GPU e2e suite.

Contract

Each workload image must:

  • Use the standard OpenShell sandbox base image as its final-stage base or ensure that the requirements for a sandbox image are met.
  • Provide a manifest command that runs the workload inside the sandbox image.
  • Run the same workload as the image default entrypoint for direct container-engine validation.
  • Require no network access after the image is pulled.
  • Print OPENSHELL_GPU_WORKLOAD_SUCCESS only when validation succeeds.
  • Print OPENSHELL_GPU_WORKLOAD_FAILURE and exit non-zero when validation fails.
  • Be usable as an OpenShell sandbox image when OpenShell invokes the manifest command explicitly.

OpenShell sandbox creation replaces the image entrypoint with the supervisor and does not run the OCI image CMD. E2e tests that use these images through OpenShell run the command from each manifest entry explicitly.

The test harness is manifest-driven. Each workload entry carries:

  • name
  • image
  • command
  • expect
  • requirements

Images

Source directory Image name Purpose
smoke-pass gpu-workload-smoke-pass Always succeeds and prints the success marker.
smoke-fail gpu-workload-smoke-fail Always fails and prints the failure marker.
cuda-basic gpu-workload-cuda-basic Runs CUDA deviceQuery and vectorAdd validation.

Build

Build all workload images:

mise run e2e:workloads:build

Build a subset by source directory name:

OPENSHELL_GPU_WORKLOAD_IMAGES=smoke-pass,smoke-fail \
mise run e2e:workloads:build

The build task uses tasks/scripts/container-engine.sh. Set CONTAINER_ENGINE=docker or CONTAINER_ENGINE=podman to choose an engine explicitly. When unset, the helper uses its existing auto-detection behavior.

Local tags use a short SHA-256 fingerprint of the selected workload contexts and external build inputs. Set OPENSHELL_GPU_WORKLOAD_IMAGE_TAG=<tag> to override the tag.

The task writes the latest build refs to:

e2e/gpu/images/.build/latest.env

The task also writes the local workload manifest used by the Rust e2e runner:

e2e/gpu/images/.build/workloads.yaml

That local manifest is created by mise run e2e:workloads:build. It contains the full image reference, command, expected outcome, and requirements for each selected workload. It also records the external build inputs used to produce the workload images.

Use the env file in later commands:

source e2e/gpu/images/.build/latest.env

That env file exports OPENSHELL_E2E_WORKLOAD_MANIFEST pointing at the local manifest. The per-image refs remain available as a convenience for direct container-engine validation.

Direct Validation

Validate smoke pass:

docker run --rm "${OPENSHELL_E2E_GPU_SMOKE_PASS_IMAGE}"

Validate smoke fail:

docker run --rm "${OPENSHELL_E2E_GPU_SMOKE_FAIL_IMAGE}"

The smoke fail command should exit non-zero and print OPENSHELL_GPU_WORKLOAD_FAILURE.

Validate CUDA with Docker CDI:

docker run --rm --device nvidia.com/gpu=all \
  "${OPENSHELL_E2E_GPU_CUDA_WORKLOAD_IMAGE}"

Use podman run with the same --device nvidia.com/gpu=all option on hosts where Podman CDI is configured.

Direct container-engine validation catches image, CDI, CUDA, and host GPU setup issues before OpenShell sandbox behavior is involved.

Manifest-Driven Validation

Run manifest-driven GPU validation through the e2e tasks so the workload images, manifest, gateway, and container-engine environment match CI:

mise run e2e:workloads:build
mise run e2e:docker:gpu

For Podman GPU validation, build the manifest with CONTAINER_ENGINE=podman mise run e2e:workloads:build, then run mise run e2e:podman:gpu.

The workload validation path reads:

OPENSHELL_E2E_WORKLOAD_MANIFEST

When that variable is unset, the runner uses the default local manifest path:

e2e/gpu/images/.build/workloads.yaml

If neither path exists, the workload validation test prints a clear skip message telling you to run:

mise run e2e:workloads:build

or to set OPENSHELL_E2E_WORKLOAD_MANIFEST to an external manifest.

Each manifest entry supplies the sandbox image and command. OpenShell runs that command through openshell sandbox create --gpu --from <image> -- <command>. The test runner iterates all GPU-tagged workload entries and enforces each entry's declared expectation:

  • expect: pass requires OPENSHELL_GPU_WORKLOAD_SUCCESS
  • expect: fail requires OPENSHELL_GPU_WORKLOAD_FAILURE

The current local manifest includes three workloads:

  • smoke-pass expected to pass
  • smoke-fail expected to fail
  • cuda-basic expected to pass

External Manifests

External workload catalogs can use the same schema. Point the runner at one with:

export OPENSHELL_E2E_WORKLOAD_MANIFEST=/abs/path/to/workloads.yaml

That lets alternate workload manifests use the same test runner without introducing per-workload env vars.