mirror of
https://github.com/NVIDIA/OpenShell.git
synced 2026-10-02 07:34:45 +08:00
* fix(supervisor): wait for repair when the gateway refuses a startup policy write Startup writes the sandbox policy to the gateway in two cases: it uploads a discovered image policy when the gateway has none, and it writes the policy back after adding the proxy baseline filesystem paths. When the gateway refused either write with FAILED_PRECONDITION or INVALID_ARGUMENT, for example because the policy binds a provider that is not attached, startup treated the refusal as a permanent error and the supervisor exited. The sandbox never reached the ConfigurationInvalid repair state that other startup rejections use. Report such a refusal as a configuration rejection carrying the gateway's message, log it once per write and error code, and keep polling, so attaching the provider or replacing the policy completes startup. Other error codes keep their current handling: transient codes are retried, and permission, not-found and authentication failures still end startup. Skip the baseline-path write-back while a global policy is active. The gateway refuses every sandbox policy write in that state, so startup exited whenever a global policy lacked a baseline path. The supervisor now adds the paths to its own copy of the policy without saving a revision. Signed-off-by: Shiju <shiju@nvidia.com> * test(supervisor): stabilize startup refusal log capture Keep a second tracing dispatcher alive while capturing startup refusal logs. With only one dispatcher, a parallel test thread without a default subscriber can cache Interest::never for the shared OCSF callsite after the capture thread rebuilds the cache. Preserve the exact log-count, diagnostic, configuration-generation and repair assertions. Production startup behavior is unchanged. Signed-off-by: Shiju <shiju@nvidia.com> * fix(supervisor): reconcile stale startup rejection reports Refetch desired configuration immediately when a rejection report is aborted because its generation changed. Preserve acknowledged rejection pacing and all other report error handling. Signed-off-by: Shiju <shiju@nvidia.com> * test(supervisor): box startup repair race futures Keep the repair regressions below the large-future lint threshold without changing their inputs, scheduling, or assertions. Signed-off-by: Shiju <shiju@nvidia.com> --------- Signed-off-by: Shiju <shiju@nvidia.com>