mirror of
https://github.com/agent-substrate/substrate.git
synced 2026-10-02 03:24:42 +08:00
feat: CDI-based NVIDIA GPU passthrough into gVisor actor containers (#502)
## Summary - `atecontroller` propagates a pool's `nvidia.com/gpu` request onto the `ateom` container and mounts the host NVIDIA toolkit read-only (path overridable via `ATE_NVIDIA_TOOLKIT_HOST_PATH`) - `ateom-gvisor` generates a CDI spec with `nvidia-ctk` and injects the device nodes, driver-library mounts, and env into each actor container's OCI spec - runs the CDI `createContainer` hooks except `update-ldcache` which needs a privileged ateom, staging the SONAME symlinks it would create from each library's ELF `DT_SONAME` - enables `runsc --nvproxy` at sandbox creation Requesting `nvidia.com/gpu` on the pool is the only configuration needed; a pool that requests N GPUs makes all N usable. Two details of the CDI spec are worth calling out, because getting either wrong fails at runtime rather than at parse time. `nvidia-ctk` leaves `major`/`minor` unset — CDI delegates that to the OCI runtime — so each device node is resolved by stat-ing the host; without it the actor gets `0,0` char devices and NVML reports it cannot communicate with the driver. And it emits per-index, per-UUID, and `all` devices that repeat the same nodes, so only `all` is applied. The spec is plain JSON, so `encoding/json` suffices and no CDI library is vendored. `update-ldcache` is the one hook that cannot run here: its `ldconfig` unshares a mount namespace and mounts a private `/proc`, which `mount_too_revealing()` rejects under the pod's masked `/proc`. Permitting it would need `procMount: Unmasked`, which Kubernetes only allows with `hostUsers: false`, and that user namespace breaks the per-actor cgroup delegation from #496. Skipping it avoids the whole chain, so a GPU worker keeps the same posture as any other unprivileged gVisor worker. `create-symlinks` and `enable-cuda-compat` still run unmodified. `--nvproxy` must be set when the sandbox is created — the `pause` container, which holds no GPU devices — so runsc's auto-detection never fires on its own; without the flag the GPU subcontainer crashes the sentry on start. GPU detection matches any device index rather than assuming `/dev/nvidia0`, since a worker sharing a multi-GPU node can be assigned `/dev/nvidia2` and `/dev/nvidia3`. GPU pools must set `spec.ateomImage` to a glibc build (`KO_DEFAULTBASEIMAGE=debian:stable-slim ko build ./cmd/ateom-gvisor`) because the distroless default cannot exec `nvidia-ctk`; the default base is unchanged for every other pool. `atelet` also has to run on the GPU nodes to restore actors there, so its DaemonSet needs a toleration for whatever taint they carry. Both are documented in the API guide rather than defaulted. ## Testing - `make test` - `env -u NO_COLOR make verify` - Real GPU, Tesla T4 / driver 580.65.06, through the full actor flow: `nvidia-smi`, `vectorAdd`, `nbody` at 3.77 TFLOP/s, PyTorch matmul via cuBLAS at 3.5 TFLOP/s (T4 peak FP32 is ~8.1, so no measurable sandbox penalty) - Actor whose entrypoint runs the CUDA sample directly exits 0, confirming the injected env reaches the workload without help from the test harness - Two workers holding two GPUs each on one 4-GPU node see disjoint device sets Snapshot and restore work when the workload holds no CUDA context. A live CUDA context cannot be checkpointed - gVisor fails with `can't save with live nvproxy clients` and the failed checkpoint terminates the sandbox so a GPU actor can only be suspended between CUDA workloads. Documented as a known limitation in the API guide; a follow-up issue will track lifting it via `cuda-checkpoint`. Fixes #627 - [x] Tests pass - [x] Appropriate changes to documentation are included in the PR --------- Signed-off-by: Eliran Wolff <eliranw@nvidia.com>
This commit is contained in:
+54
-9
@@ -43,7 +43,13 @@ spec:
|
||||
# gvisor SandboxConfig unless sandboxConfigName is set.
|
||||
```
|
||||
|
||||
### Example with GPU node scheduling
|
||||
### GPU worker pools
|
||||
|
||||
A GPU pool needs two things: (1) scheduling onto GPU nodes, and (2) a
|
||||
`nvidia.com/gpu` request in `template.resources`. The request does double duty —
|
||||
it makes the device plugin assign a GPU to the worker pod **and** triggers
|
||||
Substrate to pass that GPU **through to each actor's sandbox**. No per-actor
|
||||
configuration is needed.
|
||||
|
||||
```yaml
|
||||
apiVersion: ate.dev/v1alpha1
|
||||
@@ -53,8 +59,10 @@ metadata:
|
||||
namespace: ate-demo
|
||||
spec:
|
||||
replicas: 5
|
||||
ateomImage: ko://github.com/agent-substrate/substrate/cmd/ateom-gvisor
|
||||
# GPU pools need a glibc ateom-gvisor build — see Requirements below.
|
||||
ateomImage: <your-registry>/ateom-gvisor-glibc@sha256:...
|
||||
template:
|
||||
# (1) schedule onto GPU nodes
|
||||
nodeSelector:
|
||||
cloud.google.com/gke-accelerator: nvidia-tesla-t4
|
||||
tolerations:
|
||||
@@ -62,13 +70,6 @@ spec:
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
priorityClassName: substrate-workers
|
||||
nodeAffinity:
|
||||
requiredDuringSchedulingIgnoredDuringExecution:
|
||||
nodeSelectorTerms:
|
||||
- matchExpressions:
|
||||
- key: workload
|
||||
operator: In
|
||||
values: [substrate]
|
||||
resources:
|
||||
requests:
|
||||
cpu: 500m
|
||||
@@ -76,8 +77,52 @@ spec:
|
||||
limits:
|
||||
cpu: "1"
|
||||
memory: 2Gi
|
||||
# (2) claim a GPU — this request is what triggers GPU passthrough
|
||||
nvidia.com/gpu: "1"
|
||||
```
|
||||
|
||||
`atecontroller` propagates the request onto the `ateom` container and mounts the
|
||||
host NVIDIA toolkit into the pod. `ateom-gvisor` then generates a CDI spec with
|
||||
`nvidia-ctk` and injects the GPU device nodes, driver libraries, and env into
|
||||
each actor container's OCI spec, and runs `runsc` with `--nvproxy` so CUDA and NVML
|
||||
work inside the sandbox. A worker requesting `nvidia.com/gpu: N` passes all N
|
||||
through.
|
||||
|
||||
Every container in the actor gets the GPU. An actor's containers share one sandbox
|
||||
and the worker's whole device set, and `ActorTemplate` has no per-container resource
|
||||
fields, so GPUs are shared at the actor level rather than assigned to one container,
|
||||
the same as cpu and memory. This differs from a Kubernetes Pod, where the GPU goes
|
||||
only to the container that requests it. If per-container resource limits are added
|
||||
later, GPU assignment should follow them.
|
||||
|
||||
The driver library directory is prepended to each container's `LD_LIBRARY_PATH` so an
|
||||
image does not have to set it to find `libcuda.so.1`; any existing value is kept after
|
||||
it rather than replaced.
|
||||
|
||||
**Requirements**
|
||||
|
||||
- **A glibc `ateom-gvisor` image**, set as `spec.ateomImage`. The distroless default
|
||||
cannot exec `nvidia-ctk`. Build one with
|
||||
`KO_DEFAULTBASEIMAGE=debian:stable-slim ko build ./cmd/ateom-gvisor`.
|
||||
- **`nvidia-ctk` on the node**, at the path mounted into the worker — by default
|
||||
`/usr/local/nvidia/toolkit`, overridable with the controller's
|
||||
`ATE_NVIDIA_TOOLKIT_HOST_PATH`. gpu-operator installs it; GKE's built-in GPU
|
||||
support does not.
|
||||
- **The driver mounted into the pod by the device plugin**, at `/usr/local/nvidia`
|
||||
by default. `nvidia-ctk` needs its libraries to enumerate the GPUs at all, so a
|
||||
cluster whose plugin uses a different layout must set the controller's
|
||||
`ATE_NVIDIA_DRIVER_ROOT`.
|
||||
- **`atelet` must run on the GPU nodes**, so add a matching toleration to its
|
||||
DaemonSet if those nodes are tainted (for example `nvidia.com/gpu`).
|
||||
- **gVisor only.** `microvm` pools would need VFIO PCI passthrough instead.
|
||||
|
||||
**Known limitation: a GPU actor can only be suspended while no CUDA context is
|
||||
open.** gVisor cannot serialize GPU state, so a checkpoint taken while the workload
|
||||
holds a context fails with an nvproxy encoding error and terminates the
|
||||
sandbox. Workloads that run CUDA and exit snapshot normally; one that keeps a
|
||||
context alive (a model resident in device memory, say) cannot be suspended or have a
|
||||
golden snapshot taken.
|
||||
|
||||
### Status (`WorkerPoolStatus`)
|
||||
|
||||
| Field | Type | Description |
|
||||
|
||||
Reference in New Issue
Block a user