Files
Shiju 8983642e28 fix(gateway): give image preparation its own deadline (#4038)
* feat(server): separate image preparation and admission deadlines

Give sandbox image preparation its own deadline and start the admission
deadline after preparation finishes. Persist both phases across gateway
restarts and show recovery guidance for the phase that expired.

Preserve the current service authorization schema and regenerate the Go
bindings with the preparation timestamps.

Fixes #3952
Related to #3955

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(cli): simplify preparation timeout fallback selection

Use lazy Option fallbacks while preserving timeout messages and retained
sandbox behavior.

Signed-off-by: Shiju <shiju@nvidia.com>

* docs(server): clarify admission timer prerequisites

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(compute): enforce deadlines during initial sandbox create

Release stalled create operations after preparation expires and preserve
the timeout diagnosis across late driver results. Keep failed-create cleanup
bound to its original attempt so another replica can retry safely.

Signed-off-by: Shiju <shiju@nvidia.com>

* test(server): satisfy deadline regression lints

Drop the create-error mutex guard before matching its cloned value and use
idiomatic iteration and timeout matching in the deadline fixtures.

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(server): retain ownership of pending provisioning operations

Keep submitted create and start requests alive after caller cancellation,
monitor failure, or preparation timeout. Persist request ownership before
dispatch and retain staged uploads while the driver response is pending.

Require timeout cleanup to stop compute after driver settlement without
discarding an active cleanup claim. Fence late result handling against newer
operations, preserve failed-start recovery, and defer automatic restart while
another request owns the sandbox. Expose pending ownership in CLI JSON.

Add ordered multi-replica and cancellation regressions for the review findings.

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(openshell): preserve tracing and accept ready create responses

Carry the request span into the detached provisioning worker so compute
driver calls remain attached to their parent trace after task handoff.

Update the compensation regression to require no backend DELETE when the
durable cleanup claim fails, matching the operation ownership requirement.

Accept the gateway's current Ready snapshot when a sandbox becomes ready
before CREATE returns. Do not require the client to observe an earlier
provisioning phase. Cover the Ready-only watch and command attachment.

Signed-off-by: Shiju <shiju@nvidia.com>

---------

Signed-off-by: Shiju <shiju@nvidia.com>
2026-10-03 20:48:18 +00:00
..