Files
substrate/benchmarking/automation/README.md
T
Haven Xia bdf494999b api: rename ateomImage to workerImage and make it optional (#1210)
Rename `WorkerPoolSpec.AteomImage`to `WorkerImage` and drop the
constraints so the field can be left unset.

This is the basis for let an empty workerImage lets the controller
inject a versioned default image chosen by the pool's sandbox class.

Part of #861

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-28 12:46:00 -07:00

7.5 KiB

Substrate benchmark automation

Scheduled, repeatable benchmark runs of a substrate branch. A CronJob on an orchestration cluster drives the full build/deploy/run/teardown cycle against a separate test cluster, once per entry in tests.yaml. Each run uploads results to GCS via benchmarking/locust/runner.py (type: locust) or benchmarking/nighthawk-ingress/runner.py (type: nighthawk-ingress, the router capacity benchmark — see benchmarking/nighthawk-ingress/README.md).

How it works

  1. CronJob fires on the orchestration cluster; pod starts (orchestrator + DIND sidecar).
  2. Orchestrator waits for DIND, copies the appropriate ate-dev-env.sh, runs gcloud container clusters get-credentials for the test cluster.
  3. Shallow-clones --repo at --branch, captures the commit hash.
  4. docker build && docker push builds the runner image for each test type in use, tagged with the commit hash: ${KO_DOCKER_REPO}/locust-test:<commit> and/or ${KO_DOCKER_REPO}/nighthawk-ingress-test:<commit>.
  5. hack/install-ate.sh --deploy-ate-system + benchmarking/workloads/deploy.sh --deploy --sandbox-class <class> (these build & push substrate / workload images via ko as part of their deploy steps — there's no separate make build-images step). For a microvm test the orchestrator also runs hack/install-microvm-deps.sh --install between the two, which stages kata + cloud-hypervisor + virtiofsd assets to the cluster's object store bucket and applies the cluster-wide microvm SandboxConfig. For a nighthawk-ingress test the orchestrator additionally patches the atenet-router Deployment right after deploy_substrate: envoyCpu is the benchmark's independent variable and the shipped manifest sets no cpu resources or --concurrency, so each test pins cpu requests=limits and Envoy's thread count to its own envoyCpu, then waits for the rollout before deploying workloads. The teardown after each test redeploys substrate, so the pin never outlives its run.
  6. For each test in tests.yaml:
    • Submits a Job using the just-built image for the test's type (runner-job.yaml.tmpl for locust, nighthawk-ingress-runner-job.yaml.tmpl for nighthawk-ingress).
    • Polls until complete/failed/timeout; tails logs; deletes the Job.
    • Tears down workloads + micro-VM deps (if any) + substrate.
    • If not the last test, redeploys them so the next run starts clean.

Choosing a sandbox class

Each entry in tests.yaml may set sandboxClass: gvisor | microvm (default gvisor). This controls both spec.sandboxClass on the benchmark WorkerPool and its workerImage (ateom-gvisor vs ateom-microvm).

For microvm tests the target cluster must have KVM-capable nodes and the object store bucket named in its .ate-dev-env.sh must be writable by the orchestrator's Workload Identity principal.

Setup

./benchmarking/automation/setup.sh

The wizard treats the repo's .ate-dev-env.sh as the source of truth for the target cluster's environment and snapshots it to scratch/target-clusters/<name>.sh. It only prompts for the target cluster name (the routing key in tests.yaml) and the GCP project ID of the orchestrator image registry. It then builds + pushes the orchestrator image to gcr.io/<ORCH_PROJECT_ID>/ate-images/substrate-benchmark-orchestrator:<short-commit> (with a -dirty suffix if benchmarking/automation/ has uncommitted changes) and renders scratch/cronjob.yaml, scratch/test-list.yaml, and scratch/target-clusters.yaml. You can edit the .ate-dev-env.sh for your workload cluster directly in the config map.

Then edit the --repo / --branch / --dest args in scratch/cronjob.yaml and apply:

kubectl --context=<orchestration-cluster> apply -f scratch/cronjob.yaml

To trigger immediately instead of waiting for the schedule:

kubectl --context=<orchestration-cluster> -n substrate-benchmark \
  create job --from=cronjob/substrate-benchmark manual-$(date +%s)

To change the schedule, edit spec.schedule in scratch/cronjob.yaml (the default is 0 3 * * *, 3am UTC).

Test cluster prerequisites

Create the test cluster with the substrate-required beta APIs, Workload Identity, and Managed OpenTelemetry enabled. The control plane must be on Kubernetes 1.36+ so certificates.k8s.io/v1beta1 is available:

gcloud container clusters create <CLUSTER_NAME> \
  --location=<CLUSTER_LOCATION> \
  --num-nodes=5 \
  --workload-pool=<PROJECT_ID>.svc.id.goog \
  --managed-otel-scope=COLLECTION_AND_INSTRUMENTATION_COMPONENTS \
  --enable-kubernetes-unstable-apis=certificates.k8s.io/v1beta1/podcertificaterequests,certificates.k8s.io/v1beta1/clustertrustbundles

The orchestration cluster needs Workload Identity but no special APIs. It only ever runs one pod (the orchestrator + DIND sidecar), so a single zonal node keeps costs to the minimum:

gcloud container clusters create <ORCH_CLUSTER_NAME> \
  --location=<ORCH_ZONE> \
  --workload-pool=<ORCH_PROJECT_ID>.svc.id.goog \
  --num-nodes=1

IAM prerequisites

This setup assumes both clusters and the destination GCS bucket already exist. Two Workload Identity bindings are needed.

Both ServiceAccounts are created by the manifests (cronjob.yaml for the orchestrator KSA, runner-job.yaml.tmpl for the runner KSA — applied by orchestrator.py at runtime), so no kubectl create serviceaccount steps are needed. Grant IAM roles directly to each KSA's Workload Identity principal (no GSA / annotation required). The principal format is:

principal://iam.googleapis.com/projects/<PROJECT_NUMBER>/locations/global/workloadIdentityPools/<PROJECT_ID>.svc.id.goog/subject/ns/<NAMESPACE>/sa/<KSA>

Orchestrator pod (KSA substrate-benchmark-orchestrator in namespace substrate-benchmark on the orchestration cluster's project) needs:

  • roles/container.admin on the test cluster's project — required to manage cluster-scoped resources (CRDs, ClusterRoles, ClusterRoleBindings, Namespaces) that hack/install-ate.sh --deploy-ate-system creates. container.developer is not enough — it intentionally omits the container.clusterRoles.* and container.customResourceDefinitions.* permissions.
  • roles/artifactregistry.writer on KO_DOCKER_REPO — for ko (substrate images) and docker push (locust image).
ORCH_PRINCIPAL="principal://iam.googleapis.com/projects/<ORCH_PROJECT_NUMBER>/locations/global/workloadIdentityPools/<ORCH_PROJECT_ID>.svc.id.goog/subject/ns/substrate-benchmark/sa/substrate-benchmark-orchestrator"
gcloud projects add-iam-policy-binding <TEST_PROJECT_ID> \
  --role=roles/container.admin --member="${ORCH_PRINCIPAL}"
gcloud projects add-iam-policy-binding <TEST_PROJECT_ID> \
  --role=roles/artifactregistry.writer --member="${ORCH_PRINCIPAL}"

Runner Job pod (KSA benchmark-runner in namespace benchmarking on the test cluster's project) needs roles/storage.objectUser on the destination bucket so runner.py can upload results:

RUNNER_PRINCIPAL="principal://iam.googleapis.com/projects/<TEST_PROJECT_NUMBER>/locations/global/workloadIdentityPools/<TEST_PROJECT_ID>.svc.id.goog/subject/ns/benchmarking/sa/benchmark-runner"
gcloud storage buckets add-iam-policy-binding gs://<DEST_BUCKET> \
  --role=roles/storage.objectCreator --member="${RUNNER_PRINCIPAL}"

Updating tests

tests.yaml is delivered to the orchestrator via a ConfigMap mounted at /etc/orchestrator/tests.yaml, so the image doesn't need to be rebuilt when the test list changes. Just reapply the config map.