mirror of
https://github.com/agent-substrate/substrate.git
synced 2026-10-02 03:24:42 +08:00
Fixes #<issue_number_goes_here> NA This PR enhances the tests to validate volume lifecycle during state transition failures • **TestDeleteActor_VolumeDeletionFailure_RetrySuccess**- Verifies that when volume deletion fails during DeleteActor, the actor transitions to ACTOR_STATE_DELETING and its volumes to STATUS_DELETING. A subsequent retry successfully finalizes volume deletion and deletes the actor from the store. • **TestResumeActor_VolumeAttachFailureAndRetry** - Verifies that when volume attachment fails during ResumeActor, the actor remains in ACTOR_STATE_RESUMING with worker assignment and provisioned volumes intact. Retrying ResumeActor re-attempts volume attachment and successfully transitions the actor to ACTOR_STATE_RUNNING. • **TestResumeActor_VolumeAttachFailure_DeleteActor** - Verifies that an actor stuck in ACTOR_STATE_RESUMING after an attach failure rejects standard DeleteActor requests, but calling DeleteActor with AnyState=true successfully cleans up worker assignments and provisioned volumes. • **TestResumeActor_MultiVolumePartialAttachFailure_Retry** - Verifies that when attaching multiple volumes during ResumeActor and one volume fails, previously attached volumes remain attached and the actor stays in ACTOR_STATE_RESUMING. A subsequent retry attaches only the remaining unattached volumes and transitions the actor to ACTOR_STATE_RUNNING. • **TestSuspendActor_VolumeDetachFailure_RetrySuccess** - Verifies that volume detach failures during SuspendActor leave the actor in ACTOR_STATE_SUSPENDING with its worker assignment held. A subsequent retry succeeds in detaching the volume, releasing the worker, and transitioning the actor to ACTOR_STATE_SUSPENDED. • **TestSuspendActor_VolumeDetachFailure_DeleteActorAnyState** - Verifies that an actor stuck in ACTOR_STATE_SUSPENDING due to a volume detach failure rejects standard deletion, but cleanly detaches volumes and deletes the actor when called with AnyState=true. • **TestPauseActor_VolumeLifecycle_DetachAndResumeAttach** - Validates the volume lifecycle across pause and unpause, confirming external volumes are detached when transitioning from RUNNING to PAUSED and re- attached when resumed back to RUNNING. • **TestPauseActor_VolumeDetachFailure_RetrySuccess** - Verifies that volume detachment failure during PauseActor leaves the actor in ACTOR_STATE_PAUSING. A retry of PauseActor re-attempts volume detachment, succeeds, and advances the actor state to ACTOR_STATE_PAUSED. • **TestResumeActor_PausedLocalSnapshotMissing_Crashes** - Verifies that if local snapshot files on the worker are missing when resuming a PAUSED actor, the control plane marks the actor ACTOR_STATE_CRASHED and frees the assigned worker pod. • **TestDetachActorVolumes** - Tests control plane volume detachment logic across multiple scenarios, including container mount filtering, partial detach errors, unassigned workers, and idempotent not-found responses from plugins. • **TestCreateActorVolumes (Storage class parameters case)** - Verifies that parameters configured on a StorageClass (such as disk and filesystem types) are correctly propagated into the volume context of newly created external volumes. • **TestEnsureExternalSnapshotsReleased_DeletePrefixFailure** - Verifies that transient object store errors during actor deletion preserve un-deleted snapshot objects and return an error, which cleanly completes deletion on a retry. • **TestMountExternalVolumes** - Tests worker-side volume mounting, verifying host mount directory creation, handling pre-existing paths, plugin errors, and ensuring that partial failures during multi-volume mounts do not unmount already mounted volumes. • **TestVolumeHostDirectoryCleanup** - Verifies worker-side host directory management, ensuring volume unmounting preserves mount directories while directory reset removes empty directories and prevents data loss on non-empty ones. • **TestExternalVolume_NodeMigration** - Verifies cross-node actor migration with external volumes by suspending an actor, evicting its worker pod, and resuming it on a new worker node while ensuring application state is preserved - [ x] Tests pass - [ ] Appropriate changes to documentation are included in the PR