Follow up on #2007
This switches the remaining two `WorkerService` RPCs
(`RequestActorSuspend` and `SetWorkerCapacity`) to declarative
validation via `cmd/ateapi/internal/apivalidation`, completing
declarative request validation across `ateapi`.
Removes the hand-written checks in `workerservice` (`validateActorRef`,
`validateReportedCapacity`) and the now-unused
`resources.ValidateGlobalObjectRef`.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Fixes a part of #1709
This adds declarative validation to atelet's `AteomSupport` RPCs.
Malformed requests are now rejected at atelet instead of being forwarded
to the control plane. Validation runs after authentication.
| RPC | Rules |
|---|---|
| `MintActorCertificate` | actor atespace/name are required short names;
actor UID is a required UUID; the CSR is required and at most 16 KiB |
| `RequestActorSuspend` | actor atespace/name are required short names;
actor UID is a required UUID |
| `SetWorkerCapacity` | `capacity` is required and must follow the
`WorkerResources` rules ateapi declares, shared through
`internal/resources` |
The `AteomHerder` RPCs and `WorkloadSpec` get their validation in a
follow-up PR.
- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
Fixes#1835
`atelet` chose the sandbox binaries and pause image for a restore from
the snapshot's `manifest.json` rather than from the control plane. This
PR makes ateapi send them on every restore, the way it already does for
Run.
**Changes**
- **proto:** add `sandbox_assets` to `atelet.RestoreRequest`.
- **ateapi:** `ensureAteletRestored` resolves the ActorTemplate's
SandboxConfig once and sends it on the local restore, the external
restore and Run. A missing SandboxConfig now fails before any atelet
call.
- **atelet:** `sandbox_assets` is required on Restore. Restore selects
the binaries and pause image with `recordFromRequest`, the same path Run
uses. The manifest still supplies the snapshot files, sandbox class and
metrics labels.
Note - The field is required, so after atelet is upgraded, resumes on
that node fail until ate-api-server is upgraded too.
**Testing**
- ateapi: every Run and Restore sends the resolved assets; a missing
SandboxConfig fails on both restore paths.
- atelet: `TestRecordFromRequest`, a validation case for a missing
field, and `TestRestoreUsesRequestSandboxAssets` (checkpoint with pause
image v1, restore with v2, and the actor runs with v2).
- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
Follow up on the comment -
https://github.com/agent-substrate/substrate/pull/1675#discussion_r4031619230
Stamps `actor_template_uid` onto `ExternalSnapshot` so template
provenance travels with the snapshot itself, rather than being inferred
from the single `ActorStatus.current_actor_template_uid`.
#### Motivation
`current_actor_template_uid` records the template the **last sprint
booted with** (`finalizeRunning` stamps it on every resume). It does
*not* record the template the **snapshot** was captured under. Those two
diverge as soon as an actor runs a sprint that does not produce a new
durable snapshot, and `loadActorForResume` was using the former to
decide whether to force a `DATA`-only restore instead of `FULL`.
Walking an actor through a repoint and a crash:
| Step | Action | `current_…_uid` | `ExternalSnapshot` |
|---|---|---|---|
| **a** | Suspend on template **A** | A | captured on **A** |
| **b** | `UpdateActor` repoints the spec to **B** | A | captured on
**A** |
| **c** | Resume → `A != B`, restores `DATA`, sprint boots on **B** |
**B** | still on **A** |
| **d** | Actor crashes (or is reverted) — no new snapshot taken | B |
still on **A** |
| **e** | Resume → `current == target == B`, so **no repoint is
detected** | B | still on **A** |
At step **e** the old check reports "template not replaced" and restores
the external snapshot in `FULL` — replaying a memory image captured
under **A**'s sandbox on **B**. That is exactly the case the guard
exists to prevent; it was silently defeated by the sprint at step **c**
advancing `current_actor_template_uid` past the snapshot.
*(Note: Paused actors do not face this divergence. Because `UpdateActor`
is only allowed in the `SUSPENDED` state, a paused actor cannot have its
template updated. Therefore, local checkpoints do not need separate
template provenance tracking).*
#### Changes
- **Data Model Updates:** Added the `actor_template_uid` field to
`ExternalSnapshot`.
- **Snapshot Stamping:** Updated the capture logic to stamp external
snapshots with the active sandbox's `ActorTemplate` UID when suspending
an actor or creating an actor from a tag.
- **Resume Evaluation Logic:** Modified the external resume logic to
compare the target template UID against the specific snapshot's UID
(`ExternalSnapshot.actor_template_uid`) rather than the actor's overall
status. Local restores bypass this check entirely since templates cannot
change during a pause.
**Tests**
- `TestResumeActor_AteletWireRequest` (unit): snapshot fixtures now
carry `actor_template_uid`, since provenance is read from the snapshot
rather than from actor status.
- Functional expectations for create/update/resume/pause updated for the
new field on the external snapshot golden files.
- `TestUpdateTemplateLifecycle` (e2e): asserts
`external_snapshot.actor_template_uid` after suspend and re-suspend.
- **e2e (Update Template)**: Added coverage for repoint detection after
a revert (verifying the `resume` → `revert` → `resume` path described in
the table above).
- **e2e (Combined Volumes)**: Added coverage to explicitly verify which
volumes a revert rewinds.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
`RevertActor` returns a RUNNING, PAUSED, or CRASHED actor to SUSPENDED
at its last external snapshot, so CRASHED is no longer a dead end that
only `DeleteActor` can clear. Docs and code comments still described it
as terminal and told operators to delete and recreate the actor, losing
its state.
Update the api-guide, architecture, upgrade guide, and kubectl-ate
README to cover the new verb, and correct the comments that justified
keeping a partial external snapshot by naming actor deletion as the only
remaining collector -- revert collects it too.
Follow up for the PR - #1675
Issue - #1556
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
#### Summary
This PR introduces the `RevertActor` RPC for actor lifecycle management.
It includes the API definition, the corresponding workflow execution
logic, observability metrics, and updates to the authorization model to
support reverting actors.
Fixes#1556
Docs updated in #1711
#### Commit-wise Changes
**1. Add RevertActor RPC (`1145d99`)**
* Introduces the new `RevertActor` RPC to the API definitions.
* Updates the corresponding protobuf bindings (affecting files like
`ateapi.pb.go` and `ateapi_pb2.py`).
**2. Add the revertActor workflow and observability metrics
(`dcb1fae`)**
* Implements the core `revertActor` workflow logic, designed to be
idempotent and re-enterable. It progresses through the following steps:
* **Mark Reverting:** Validates that the actor is in a revertable state
(`RUNNING`, `PAUSED`, or `CRASHED`) and transitions its state to
`REVERTING`.
* **Discard Worker:** Safely tears down the execution environment by
terminating the workload, detaching volumes, and releasing the assigned
worker.
* **Collect In-Progress Snapshot:** Cleans up external object storage by
deleting any objects a previous suspend operation was partway through
writing.
* **Finalize:** Commits the actor to `SUSPENDED` and strips all
node-local and in-progress state pointers (clearing `WorkerAssignment`,
`LocalSnapshotInfo`, etc.), returning the actor to its untouched
external snapshot.
* Instruments the workflow with lifecycle operation metrics (e.g.,
updating `ate.actor.lifecycle.operation.duration` to track `revert`
operations).
* *Note/TODO:* Currently, when reverting a paused actor, the workflow
drops the pointer to the node-local state but does *not* actually prune
the local checkpoint bytes from the node (this is tracked in #641).
**3. Add `can_revert` to the authorization model (`2fc53ae`)**
* Adds the `can_revert` permission to the auth model, mirroring the
shape of `can_suspend` (editor tier of the parent atespace, plus a
direct grant so a machine identity can revert the actor it drives
without holding an atespace role).
**4. Serve RevertActor and add the CLI verb (`f04fab4`)**
* Wires the `Control.RevertActor` service method to the workflow
(replacing the generated stub that previously answered `Unimplemented`).
* Adds the `"ate revert actor"` CLI command, making the feature usable
end-to-end.
* Implements `Terminate` for the fake atelet. This was necessary because
reverting an actor from the `RUNNING` state is the first path to reach
this call in functional tests (previously, delete tests skipped this
step as they ran against actors with no worker assignment).
**5. Add a manual verify script for RevertActor (`8022db3`)**
* Adds a script to manually exercise `RevertActor` against a real
control plane, since unit and functional tests only run against a fake
atelet.
* Tests reverting from `CRASHED`, `RUNNING`, and `PAUSED` states, and
verifies that attempting to revert a `SUSPENDED` actor is properly
rejected.
* Simulates a crash by deleting the worker pod the actor runs on to
verify the workflow can handle the absence of a worker to terminate.
- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
When a worker pod crashes or is deleted out-of-band, `ateapi`
transitions the actor to `CRASHED` and clears `worker_assignment`
without calling `atelet.Terminate`. If the actor is later recovered and
scheduled onto a new worker pod on the same node,
`systemInfoVolumeRefresher.Register` panicked on the leftover entry in
r.actors.
Supersede any existing registration for the same actor UID by marking
the previous entry stale under its mutex and replacing it in r.actors.
Temporary fix for #1710
- [X] Tests pass
- [X] Appropriate changes to documentation are included in the PR
Fixes#1370
**Summary**
`atepg` keeps a registry of the live `WatchWorkers` channels and
publishes each committed worker event onto them at commit, about one
poll interval ahead of the outbox row carrying the same event. Replicas
see their own writes sooner without a second notification API.
**Why**
Replicas previously learned of their own writes only from the outbox
watch, ~50ms later. In that window the scheduler re-picks a worker it
just assigned and retries on the precondition check, and capacity a
replica just freed stays invisible to its own placement. Cross-replica
propagation is unchanged.
**Key changes**
- `WatchWorkers` enrolls its channel; the poller deregisters before
closing it, which is what makes a send on a closed channel impossible.
- `writeAndAppendEvent` publishes only after `tx.Commit` returns.
- The commit-time copy is decoded from the same payload the outbox
stores, so both deliveries are identical, including unknown-field
pruning.
- Sends are non-blocking and never advance the poll cursor, so a full
buffer costs a watcher nothing and gap-free delivery is unchanged.
- Consumers now also see their own writes out of xid order and again on
the poll. `workercache` already fences creates and updates on version,
so duplicates and stale events are no-ops.
**Testing**
`TestLocalPublishReachesWatchers`, `TestLocalPublishSurvivesWatchClose`,
and `TestCache_WatchEventsAreFenced` are new.
`TestWatchWorkers_DeliveryFencedByOldestTransaction` now consumes the
commit-time copy first and asserts the xmin fence on the outbox copy;
the fence governs the poll, and a local publish bypasses it by
construction.
Full `atepg` and `workercache` suites pass under `-race`.
- [X] Tests pass
- [X] Appropriate changes to documentation are included in the PR
Fixes a few todos in the `ateapi.proto`
* Adds custom validation for `create_time` and `update_time` fields.
* Adds better validation method for the `Container.image` field.
- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
### Description
This PR introduces native, secure support for using Cloud SQL as the
PostgreSQL store backend for `ate-api-server`.
To ensure the highest level of security and ease of use in GCP
environments, this integration leverages the Cloud SQL Auth Proxy
sidecar with automatic IAM database authentication. This means transport
security (TLS 1.3 tunnel) is handled automatically, and database
sessions are authenticated using Workload Identity via short-lived OAuth
tokens, completely eliminating the need for database passwords.
### Key Changes
* **Cloud SQL Auth Proxy Sidecar:** Added
`manifests/ate-install/cloudsql-proxy-patch.yaml` to patch the sidecar
into the `ate-api-server` deployment when a Cloud SQL instance is
configured.
* **Automated Provisioning:** Extended `tools/setup-gcp` with a new
`cloudsql` command. This handles the idempotent creation of the Cloud
SQL instance, Google Service Accounts (GSA), IAM bindings, and Workload
Identity bindings.
* **Installation Script Updates:** Updated `hack/install-ate.sh` to
parse new environment variables (e.g.,
`ATE_API_POSTGRES_CLOUDSQL_INSTANCE`, `ATE_API_POSTGRES_CLOUDSQL_GSA`)
and correctly synthesize the passwordless DSN and ConfigMaps for the
proxy.
* **Security & Documentation:**
* Added extensive documentation in `tools/setup-gcp/cloud-sql.md`
covering provisioning, schema privileges, deployment, and database
scaling.
* Updated `docs/threat-model.md` to reflect the new Cloud SQL egress
flows and Auth Proxy tunnel mechanics.
* **Dependencies:** Vendored required Google API clients (`sqladmin/v1`,
`servicenetworking/v1`, `iam/v1`) for the GCP setup tool.
- [X] Tests pass
- [X] Appropriate changes to documentation are included in the PR
Follow up on issue #1168 and the base PR #1215
In this PR:
**Actor read/lifecycle verbs — completes Actor end-to-end**
* `Get`/`Delete`/`Suspend`/`Pause`/`Resume` `ActorRequest`: dropped the
`+k8s:opaqueType` fence on the actor ref, added `+k8s:required` +
`+k8s:subfield(atespace)=+k8s:required`. Combined with the recursion
into `Validate_ObjectRef` this reproduces exactly what
`resources.ValidateObjectRef` enforced (ref required, atespace required,
name required, both DNS-1123 labels).
* `ListActorsRequest`: tagged to match `ListAtespacesRequest` (atespace
optional + short-name format, `page_size` minimum 1, `page_token` max
length 256).
* All six `validate*Request()` functions are now one-line calls to the
generated validators. With this, every RPC verb on `Actor` and
`Atespace` is fully declarative; `resources.ValidateObjectRef` has no
callers left in `actor.go`.
**Note:**`ListActorsRequest.page_token` now capped at 256 chars (parity
with `ListAtespaces`).
**Testing**
* Existing verb tests updated to the generated error shapes
(`field.ErrorMatcher` with origin matching).
* New coverage: `page_token` boundary cases for `ListActors`;
* `controlapi`, `functionaltest`, `actoridentity`, `store`, and `atepg`
suites all pass; `go vet` and `gofmt` clean.
- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
---------
Co-authored-by: Michelle Au <msau42@users.noreply.github.com>
Continues the declarative-validation (DV) migration from #1215, covering
the `ActorTemplate` resource, its full spec tree, the create path, and
read/delete verbs.
**Key Changes**
* **DV Tags & Hooks**: Added validation tags and custom hooks across the
`ActorTemplate` spec (`Metadata`, `SandboxConfig`, `SnapshotsConfig`,
`Container`, `Volume`, `Resources`).
* **Create Path**: Replaced hand-written validation with generated
validators in `CreateActorTemplate`. Custom rules (e.g., `on_commit ⊆
on_pause`) are now handled via custom hooks.
* **Read/Delete Verbs**: Converted `Get`, `List`, and `Delete` requests
to use DV.
* **State Updates**: Refactored `atepg` metadata setters to update
in-place.
* Regenerated apitool exemptions (documented 15 previously-exempt
fields).
**Deliberate Behavior Changes**
The gRPC path now strictly enforces CRD rules. Specific tightenings
include:
* Empty `worker_selector` is now rejected.
* `EnvVar.name` is strictly required.
* Negative values in optional enum fields are rejected (`minimum=1`).
* `page_token` is capped at 256 characters (matching other list
requests).
**Deliberately Not Done**
* **No update DV / status tags**: `ActorTemplates` are immutable to
clients. The reconciler updates the status directly against the store,
so tagging the status subtree or adding update validation isn't
necessary right now.
* **`config_name` matching `sandbox_class`**: Enforcing this requires
calling SandboxConfigLister (ServiceImpl), so it's left as a TODO.
**Testing**
* Added ~70 positive and negative test cases in the validator tables.
* Store-backed tests (`TestCreateActorTemplate*`,
`TestUpdateActorTemplateMetadata`, and the `atepg` suite) verified
against real Postgres via testcontainers.
- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
Follow up on issue #1168 and the base PR #1215
**Overview:**
Continues the declarative validation (DV) migration from #1215 by
converting the `Worker` resource. This moves immutable-field enforcement
from the storage layer up to the service layer. With `WorkerPoolSyncer`
now using RPCs, nearly all write paths share the exact same validation.
**Key Changes:**
* **Schema (`ateapi.proto`):** Added full DV tags (required, format,
immutable) to `Worker` pod-coordinates and metadata.
`WorkerStatus.state` is now strictly bounded to its enum range.
* **Service Layer:** Replaced ~60 lines of hand-written validation with
1-line generated calls. `CreateWorker` and `UpdateWorker` now scrub
server-owned fields and enforce immutability *before* hitting the store.
* **Storage Layer:** Removed `store.CheckWorkerMutation`.
`atepg.UpdateWorker` is optimized to only clone metadata instead of the
whole worker. Immutability contract tests were moved to the service
layer.
* **Behavior Tweaks (Tightenings):**
* `page_token` is now capped at 256 chars.
* `DeleteOptions.uid` must be a valid UUID.
* Global-ref atespace violations now correctly return `Forbidden`
instead of `Invalid`.
* Immutability checks now natively cover the `WorkerPoolSyncer` path.
* **Testing:** Added comprehensive positive/negative cases for all newly
tagged fields, boundary cases, and moved the immutability test suite.
All suites pass cleanly.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Follow up on issue #1168 and the base PR #1215
#### Identity RPCs (`MintJWT` / `MintCert`)
* **Requests:** Added full declarative validation tags (required,
formatting, enum bounds) to all fields in `MintJWTRequest` and
`MintCertRequest`.
* **Handlers:** Validation now occurs immediately after authentication
(the `purpose != ATUNNEL` policy check remains in the handler).
#### Notes
* **Stricter Validation:** Empty `atespace`/`actor_name` can no longer
mint malformed JWT subjects. UIDs must now be valid UUIDs.
* **Better Errors:** `MintJWT` empty-audience error is now an
`InvalidArgument` (was `Unknown`) checked *before* disk I/O.
`MintCert`'s atespace-on-global-ref error is now typed `Forbidden` (was
`Invalid`).
#### Testing
* Added 20 new table-driven test cases (`TestValidateMintJWTRequest` /
`TestValidateMintCertRequest`), as these requests previously lacked
validation tests.
* Updated existing verb tests to match the new generated error shapes.
* All test suites (`controlapi`, `functionaltest`, etc.), `go vet`, and
`gofmt` pass cleanly.
- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
Fix `TestDurableDirLifecycle` flaky test
https://github.com/agent-substrate/substrate/actions/runs/32416349094/job/96578260374
Cause:
- In the specific test (`_suspend_from_PAUSED`), the actor is suspended
(which strips its worker assignment, returning the worker to the pool)
and then immediately resumed.
- When the test calls `ResumeActor`, the control plane selects a new
worker, writes the assignment to the Kubernetes database, and
synchronously makes a direct gRPC Restore() call to the atelet on that
new worker.
- The `atelet` boots the actor and `atunnel`. `atunnel` then immediately
asks the `ateapi` (via the `ActorIdentity` service) to mint an actor
certificate so it can connect to the network.
Race: The `ActorIdentity` service verifies that the worker is actually
assigned to the actor by checking its internal `workercache.Cache`.
Because `ResumeActor` was so fast, the cache hasn't processed the new
worker assignment yet.
The `ActorIdentity` service looks at the stale cache, sees the worker is
unassigned, and hard-rejects the certificate minting with caller is not
permitted to mint credentials for this actor.
I faced this issue too when working on the postgresql updateWorker
optimization (#934)
Fix: Updated `authorizeActor` inside
`cmd/ateapi/internal/actoridentity/actoridentity.go`.
Now, if `authorizeActor` fails an authorization check (like discovering
a missing or mismatched worker assignment), it will automatically do a
one-time "live fetch" bypass (s.store.GetWorker()) directly against the
primary store to bypass the lagging informer cache.
- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
Fixes#920 partially.
Replaces per-write pg_notify with a worker_changes outbox table written
in the same transaction, and LISTEN with a 100ms polling watcher.
#### Motivation
* **Scalability Bottleneck**: `pg_notify` serializes the commits of all
notifying transactions through a global lock held across the commit
(including `fsync`). This artificially caps worker writes at ~600/s on
Cloud SQL regardless of instance size, whereas our target is O(10K)
worker updates/s.
* **Payload Limits**: Bypasses the 8KB `NOTIFY` payload limit that
previously caused writes to fail.
* **Reliability**: A cursor-based polling watcher survives reconnects
and failovers without missing events, which was a known flaw with the
ephemeral `LISTEN` approach.
*(Known Postgres pathology prior art:
[Recall.ai](https://www.recall.ai/blog/postgres-listen-notify-does-not-scale),
[DBOS](https://www.dbos.dev/blog/postgres-listen-notify-scalability)).*
### Performance Improvement
WorkerUpdate @ 1,000 QPS , 1M workers (preloaded) — before vs after the
change feed:
| | p50 | p90 | p95 | p99 |
|---|---|---|---|---|
| Before (per-update pg_notify) | 40.3s |55.6s | 61.2s | 63.8s |
| After (change-feed table) | 7.08 ms | 8.04 ms | 8.52 ms | 27.5 ms |
#### Changes Made
* **Schema**: Added transactional outbox table `worker_outbox`.
* **Write Path**: Worker writes now append to the `worker_outbox` feed
inside the same transaction instead of calling `pg_notify()`.
* **Watch Path**: Replaced `LISTEN` in `WatchWorkers` with a polling
watcher that queries the feed every 50ms.
* **Cleanup**: Implemented a janitor process during polling to
periodically delete old feed rows.
* **Tests**: Updated atomicity tests to verify feed inserts instead of
`pg_notify` payloads.
For full architecture:
https://docs.google.com/document/d/10K0wB6aTeFkJCL4HN3NbLJCdFGoLYdhIcqkFnqFHkKc/edit?tab=t.txqcjwvmhp3v
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
### Which issue(s) this PR is related to:
Fixes#721
Required for System Upgrade flow (#473)
### What this PR does / why we need it:
This PR implements graceful termination and zero-outage rolling updates
for `atenet-router` , building on the readiness probe and graceful drain
patterns established in #719 (`atelet`).
Draining a two-container networking pod (`atenet-router` Go
control-plane + `envoy` C++ proxy dataplane) introduces complex
lifecycle interdependencies. This PR resolves the dual-container SIGTERM
race, respects Envoy's `failClosed` `ext_proc` filter dependency,
preserves parked requests riding out worker pool saturation, and
accelerates idle deployments via an event-driven file handshake.
#### 1. Event-Driven Dataplane Synchronization (`emptyDir` Marker File)
Envoy fast-exits by default on SIGTERM. To keep Envoy alive while the Go
control-plane orchestrates the drain, we configured an IPC handshake
between containers via a pod-shared `emptyDir` volume mounted at
`/var/run/atenet`:
* **Go Router Container:** Removes any stale marker at startup, then
writes `/var/run/atenet/drain-complete` when its shutdown sequence
finishes.
* **Envoy Container `preStop` Hook:** Runs `while [ ! -f
/var/run/atenet/drain-complete ]; do sleep 0.5; done`.
**Outcome:** Envoy exits as soon as — the drain is done, rather than
wasting time in a fixed worst-case sleep. If the router crashes, Kubelet
terminates the `preStop` hook at `terminationGracePeriodSeconds:
60`—slower cleanup, never a wedge.
#### 2. The Multi-Container Shutdown Sequence (`drain.go`)
Because Kubernetes issues SIGTERM to both containers at once, a
coordination state machine ensures Envoy never drops a connection and
`ext_proc` is never stopped prematurely:
| Phase / (best-case ex.) Timeline | atenet-router | envoy | K8s /
Service Status |
| :--- | :--- | :--- | :--- |
| **SIGTERM Sent**<br>($t=0\text{s}$) | Catches SIGTERM<br>• Flips
`/readyz` $\rightarrow$ `503`.<br>• Starts 13s `drain-delay`. | Enters
`lifecycle.preStop` hook:<br>`while [ ! -f .../drain-complete ]; do
sleep 0.5; done`<br>• Kubelet holds SIGTERM back. | Deployment will
recreate the pod.<br>EndpointSlice controller begins dropping Old Pod
IP. |
| **Propagation**<br>($t=0\text{s}$ – $t=13\text{s}$) | Keeps serving
normally.<br>• `ext_proc` continues unparking/routing requests. | Runs
normally inside `preStop` loop.<br>• Serves active TCP connections. |
Service endpoint removal completes.<br>No new connections arrive at Old
Pod. |
| **Envoy drain**<br>($t=13\text{s}$) | Issues Envoy admin API
calls:<br>• `/healthcheck/fail`<br>•
`/drain_listeners?graceful&skip_exit`<br>• Polls `/stats` for active
downstream connections. | Begins graceful listener drain:<br>• GOAWAY /
`Connection: close` on established connections.<br>• In-flight requests
keep running. | In-flight HTTP requests finish executing through Envoy.
|
| **ext_proc drain**<br>($t\approx18\text{s}$) | Active Envoy
connections hit 0 (the poll exits early).<br>• Calls
`extproc.GracefulStop()`.<br>• Parked request streams finish. | All
downstream connections closed.<br>• Still waiting in the `preStop` loop.
| All client HTTP responses delivered. |
| **Handshake & Exit**<br>($t=18.5\text{s}$)| `ext_proc` drain
finishes.<br>• Writes `drain-complete` marker.<br>• Hard-stops xDS
(`Stop()`) & exits. | `preStop` loop detects marker file!<br>• `preStop`
exits 0.<br>• Kubelet sends SIGTERM to Envoy.<br>• Envoy exits
immediately. | Old Pod deleted cleanly at ~18.5s (well under the 60s
budget). |
Envoy ref:
https://www.envoyproxy.io/docs/envoy/latest/intro/arch_overview/operations/draining
#### Testing & Verification
1. Unit Tests (`drain_test.go`): Added comprehensive unit tests
covering:
- Graceful completion of in-flight `ext_proc` streams within deadline.
- Force-stopping straggler streams past `--drain-timeout`.
- Envoy admin API interaction and connection polling.
- Stale marker cleanup and marker file creation.
2. Local testing on kind cluster (Observations):
Test script: `hack/verify-atenet-drain.sh`
- Steady state: `/readyz=200`, `/healthz=200`; the instant the pod
turned Terminating: `/readyz=503` while `/healthz=200` — `NotReady` but
alive, for the whole drain.
- The parked request survived the shutdown: fired at a busy 1-worker
pool → parked → pod deleted while parked → worker freed → HTTP=200,
total=1.89s, body hello from: 169.254.17.2 | preserved memory count: 1 |
preserved file counter: 1 — park → resume → route, served by the
Terminating pod during its drain-delay window.
- Pod terminated 14s after deletion — the idle floor, confirmed across
three runs: >13s drain-delay (sequence ran), ≪60s grace (marker
handshake released Envoy's preStop; no `SIGKILL`).
- Log sequence captured verbatim, drain-delay honored to the millisecond
(17:34:41.706 → 17:34:54.707):
>Shutdown signal received; draining
>(+13.000s) Draining Envoy {window: 15s}
>Envoy drained
>Starting ext_proc drain
>ext_proc drain completed within deadline
>Drain-complete marker written {path: /var/run/atenet/drain-complete}
>Shutdown complete
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
#### What this PR does / why we need it:
`atelet` previously had no `SIGTERM` handling — `main()` ended in a bare
`svr.Serve(lis)`, so any pod eviction, `DaemonSet` rollout, or node
drain killed the process instantly, aborting in-flight
`Run`/`Checkpoint`/`Restore` RPCs mid-execution.
This PR adopts the same graceful-shutdown pattern `ate-api` uses: on
`SIGTERM`, mark not-ready, stop accepting new RPCs, let in-flight RPCs
finish, and force-stop after a deadline.
- `cmd/atelet/main.go`:
- `signal.NotifyContext` (kept separate from the work ctx so in-flight
RPCs aren't cancelled the moment `SIGTERM` arrives) + a local
`drainOnShutdown` mirroring `ateapi`'s: `MarkNotReady` → sleep
`--drain-delay` → `GracefulStop()` → force `Stop() `after
`--drain-timeout`.` /readyz` and `/healthz` are wired into the metrics
server (`serverboot.Readiness`), so readiness and liveness diverge
correctly during the drain.
- Flags: `--drain-delay` (default **`0s`** — `atelet` is dialed directly
by pod IP, no route-drain window is needed) and `--drain-timeout`
(default **`5m`** — `Checkpoint`/`Restore` stream multi-GiB snapshots to
object storage and
can take minutes; force-cancelling one mid-upload crashes the actor).
- `manifests/ate-install/atelet.yaml`: `terminationGracePeriodSeconds:
330 `(drain-delay + drain-timeout + slack — the sum must fit inside the
grace period or the kubelet `SIGKILLs` mid-drain), the drain flags, and
`/readyz` readiness + `/healthz` liveness probes.
**Test Scenarios Considered:**
| In-flight at `SIGTERM` | Force-stop (past timeout) |
|---|---|
| Idle | n/a |
| Checkpoint (suspend, external) | upload aborted → `CRASHED` (#362) |
| Pause (local checkpoint) | aborted → crash |
| Restore (resume) | RPC fails → ateapi retries |
| Run (cold boot) | RPC fails → retried |
| Multiple concurrent RPCs | one shared timeout; stragglers cut |
| New RPC during drain | rejected `Unavailable` → ateapi retries |
**Testing:**
- Unit (`cmd/atelet/main_test.go`): a loopback gRPC server with a
blocking handler holds an RPC in-flight across the drain.
- `TestDrainOnShutdownInFlightFinishes` — in-flight RPC completes during
`GracefulStop`; readiness flips to not-ready.
- `TestDrainOnShutdownForceStopsAfterTimeout` — an RPC running past
drain-timeout is force-cancelled by Stop().
- Live on a kind cluster (`hack/install-ate-kind.sh --deploy-atelet`,
then `kubectl delete pod` to deliver `SIGTERM`, with `--drain-delay=25s`
temporarily set to make the window observable):
- Steady state: `/readyz=200`, `/healthz=200` (old build had no /readyz
at all).
- During drain:`/readyz=503` while `/healthz=200` for the whole window —
`NotReady` but alive.
- Log sequence in order, with the 25s drain-delay honored exactly:
Shutdown signal received; draining → (+25s) Starting gRPC drain → Drain
completed within deadline → Shutdown complete.
#### Which issue(s) this PR is related to:
Fixes#719
Required for System Upgrade flow (#473)
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR