feat: CDI-based NVIDIA GPU passthrough into gVisor actor containers (#502)

## Summary

- `atecontroller` propagates a pool's `nvidia.com/gpu` request onto the
`ateom` container and mounts the host NVIDIA toolkit read-only (path
overridable via `ATE_NVIDIA_TOOLKIT_HOST_PATH`)
- `ateom-gvisor` generates a CDI spec with `nvidia-ctk` and injects the
device nodes, driver-library mounts, and env into each actor container's
OCI spec
- runs the CDI `createContainer` hooks except `update-ldcache` which
needs a privileged ateom, staging the SONAME symlinks it would create
from each library's ELF `DT_SONAME`
- enables `runsc --nvproxy` at sandbox creation

Requesting `nvidia.com/gpu` on the pool is the only configuration
needed; a pool that requests N GPUs makes all N usable.

Two details of the CDI spec are worth calling out, because getting
either wrong fails at runtime rather than at parse time. `nvidia-ctk`
leaves `major`/`minor` unset — CDI delegates that to the OCI runtime —
so each device node is resolved by stat-ing the host; without it the
actor gets `0,0` char devices and NVML reports it cannot communicate
with the driver. And it emits per-index, per-UUID, and `all` devices
that repeat the same nodes, so only `all` is applied. The spec is plain
JSON, so `encoding/json` suffices and no CDI library is vendored.

`update-ldcache` is the one hook that cannot run here: its `ldconfig`
unshares a mount namespace and mounts a private `/proc`, which
`mount_too_revealing()` rejects under the pod's masked `/proc`.
Permitting it would need `procMount: Unmasked`, which Kubernetes only
allows with `hostUsers: false`, and that user namespace breaks the
per-actor cgroup delegation from #496. Skipping it avoids the whole
chain, so a GPU worker keeps the same posture as any other unprivileged
gVisor worker. `create-symlinks` and `enable-cuda-compat` still run
unmodified.

`--nvproxy` must be set when the sandbox is created — the `pause`
container, which holds no GPU devices — so runsc's auto-detection never
fires on its own; without the flag the GPU subcontainer crashes the
sentry on start. GPU detection matches any device index rather than
assuming `/dev/nvidia0`, since a worker sharing a multi-GPU node can be
assigned `/dev/nvidia2` and `/dev/nvidia3`.

GPU pools must set `spec.ateomImage` to a glibc build
(`KO_DEFAULTBASEIMAGE=debian:stable-slim ko build ./cmd/ateom-gvisor`)
because the distroless default cannot exec `nvidia-ctk`; the default
base is unchanged for every other pool. `atelet` also has to run on the
GPU nodes to restore actors there, so its DaemonSet needs a toleration
for whatever taint they carry. Both are documented in the API guide
rather than defaulted.

## Testing

- `make test`
- `env -u NO_COLOR make verify`
- Real GPU, Tesla T4 / driver 580.65.06, through the full actor flow:
`nvidia-smi`, `vectorAdd`, `nbody` at 3.77 TFLOP/s, PyTorch matmul via
cuBLAS at 3.5 TFLOP/s (T4 peak FP32 is ~8.1, so no measurable sandbox
penalty)
- Actor whose entrypoint runs the CUDA sample directly exits 0,
confirming the injected env reaches the workload without help from the
test harness
- Two workers holding two GPUs each on one 4-GPU node see disjoint
device sets

Snapshot and restore work when the workload holds no CUDA context. A
live CUDA context cannot be checkpointed - gVisor fails with `can't save
with live nvproxy clients` and the failed checkpoint terminates the
sandbox so a GPU actor can only be suspended between CUDA workloads.
Documented as a known limitation in the API
guide; a follow-up issue will track lifting it via `cuda-checkpoint`.

Fixes #627 

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR

---------

Signed-off-by: Eliran Wolff <eliranw@nvidia.com>
This commit is contained in:
eliranw
2026-08-07 15:08:44 -07:00
committed by GitHub
parent 5f64ba4c26
commit 934c7418d0
11 changed files with 1357 additions and 21 deletions
+54 -9
View File
@@ -43,7 +43,13 @@ spec:
# gvisor SandboxConfig unless sandboxConfigName is set.
```
### Example with GPU node scheduling
### GPU worker pools
A GPU pool needs two things: (1) scheduling onto GPU nodes, and (2) a
`nvidia.com/gpu` request in `template.resources`. The request does double duty —
it makes the device plugin assign a GPU to the worker pod **and** triggers
Substrate to pass that GPU **through to each actor's sandbox**. No per-actor
configuration is needed.
```yaml
apiVersion: ate.dev/v1alpha1
@@ -53,8 +59,10 @@ metadata:
namespace: ate-demo
spec:
replicas: 5
ateomImage: ko://github.com/agent-substrate/substrate/cmd/ateom-gvisor
# GPU pools need a glibc ateom-gvisor build — see Requirements below.
ateomImage: <your-registry>/ateom-gvisor-glibc@sha256:...
template:
# (1) schedule onto GPU nodes
nodeSelector:
cloud.google.com/gke-accelerator: nvidia-tesla-t4
tolerations:
@@ -62,13 +70,6 @@ spec:
operator: Exists
effect: NoSchedule
priorityClassName: substrate-workers
nodeAffinity:
requiredDuringSchedulingIgnoredDuringExecution:
nodeSelectorTerms:
- matchExpressions:
- key: workload
operator: In
values: [substrate]
resources:
requests:
cpu: 500m
@@ -76,8 +77,52 @@ spec:
limits:
cpu: "1"
memory: 2Gi
# (2) claim a GPU — this request is what triggers GPU passthrough
nvidia.com/gpu: "1"
```
`atecontroller` propagates the request onto the `ateom` container and mounts the
host NVIDIA toolkit into the pod. `ateom-gvisor` then generates a CDI spec with
`nvidia-ctk` and injects the GPU device nodes, driver libraries, and env into
each actor container's OCI spec, and runs `runsc` with `--nvproxy` so CUDA and NVML
work inside the sandbox. A worker requesting `nvidia.com/gpu: N` passes all N
through.
Every container in the actor gets the GPU. An actor's containers share one sandbox
and the worker's whole device set, and `ActorTemplate` has no per-container resource
fields, so GPUs are shared at the actor level rather than assigned to one container,
the same as cpu and memory. This differs from a Kubernetes Pod, where the GPU goes
only to the container that requests it. If per-container resource limits are added
later, GPU assignment should follow them.
The driver library directory is prepended to each container's `LD_LIBRARY_PATH` so an
image does not have to set it to find `libcuda.so.1`; any existing value is kept after
it rather than replaced.
**Requirements**
- **A glibc `ateom-gvisor` image**, set as `spec.ateomImage`. The distroless default
cannot exec `nvidia-ctk`. Build one with
`KO_DEFAULTBASEIMAGE=debian:stable-slim ko build ./cmd/ateom-gvisor`.
- **`nvidia-ctk` on the node**, at the path mounted into the worker — by default
`/usr/local/nvidia/toolkit`, overridable with the controller's
`ATE_NVIDIA_TOOLKIT_HOST_PATH`. gpu-operator installs it; GKE's built-in GPU
support does not.
- **The driver mounted into the pod by the device plugin**, at `/usr/local/nvidia`
by default. `nvidia-ctk` needs its libraries to enumerate the GPUs at all, so a
cluster whose plugin uses a different layout must set the controller's
`ATE_NVIDIA_DRIVER_ROOT`.
- **`atelet` must run on the GPU nodes**, so add a matching toleration to its
DaemonSet if those nodes are tainted (for example `nvidia.com/gpu`).
- **gVisor only.** `microvm` pools would need VFIO PCI passthrough instead.
**Known limitation: a GPU actor can only be suspended while no CUDA context is
open.** gVisor cannot serialize GPU state, so a checkpoint taken while the workload
holds a context fails with an nvproxy encoding error and terminates the
sandbox. Workloads that run CUDA and exit snapshot normally; one that keeps a
context alive (a model resident in device memory, say) cannot be suspended or have a
golden snapshot taken.
### Status (`WorkerPoolStatus`)
| Field | Type | Description |