feat(providers): report applied sandbox provider changes (#3391)

* feat(providers): report applied sandbox provider changes

Record exact provider mutation targets in shared configuration operations.
Require authenticated evidence that credentials, effective policy, and the
workload launch environment have been installed before reporting readiness.

Add bounded CLI and Rust SDK status and wait support, preserving ordinary
revision-scoped references for existing processes. Verify new-client rotation
and acknowledged detach revocation without external-stable resolver changes.

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(providers): align readiness times with protobuf contracts

Represent readiness receipts, status, and operation times with Timestamp
and report intervals with Duration. Reserve the scalar field tags, update
all consumers and generated bindings, and preserve timestamp presence
and nanosecond identity through storage and client validation.

Qualify both empty-map constructors in the Linux boundary test so its
module compiles while retaining the explicit default required by Clippy.

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(cli): preserve provider mutation storage uncertainty

Recognize the gateway's exact structured storage-uncertainty reason for
provider attach, detach, and update. Explain that the change may already
be saved and must be reconciled before retrying, without exposing server
messages or metadata. Preserve uncertainty ahead of generic retry hints.

Exercise saved mutations through the CLI and verify single submission,
redaction, missing receipt handling, and untrusted error-detail rejection.
Document the recovery guidance for users and the public CLI skill.

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(cli): explain denied provider profile lookups

Report exact and alias profile lookup denials with fixed permission and workspace guidance. Keep backend details redacted and stop before provider mutations.

Cover denied create and update calls through the CLI. Verify the complete provider list independently in the cross-workspace OIDC regression, extracting its JSON object from surrounding startup diagnostics.

Signed-off-by: Shiju <shiju@nvidia.com>

---------

Signed-off-by: Shiju <shiju@nvidia.com>
This commit is contained in:
Shiju
2026-09-17 19:10:43 +00:00
committed by GitHub
parent 769273f096
commit 9708ba9999
58 changed files with 14841 additions and 1620 deletions
+30 -1
View File
@@ -387,6 +387,8 @@ Public RPC contracts and durable protobuf formats have separate ownership. The `
`ReportEndpointStatus` is a sandbox-authenticated public gateway RPC. Its request, response, and `EndpointObservation` messages belong only to the public closure. `EndpointStatus` and `EndpointResult` also belong to the durable closure because `Sandbox.status.endpoint_statuses` persists them. The repeated status field uses a new wire tag; stored sandboxes without it decode with an empty endpoint list and retain their lifecycle fields. A fixed payload encoded with the earlier sandbox schema verifies that no database rewrite is required.
`GetSandboxProviderStatus` and `ReportProviderReadiness` are unary public gateway RPCs. The first lets authorized users inspect a provider change; the second accepts installation reports only from the sandbox's current authenticated supervisor session.
The removed `NetworkBinary.harness` field remains reserved by number and name,
so protobuf implementations cannot reuse its wire slot or source identifier.
The durable-policy compatibility decoder reads the former boolean before Prost
@@ -408,8 +410,10 @@ than extending the frozen message.
| SQL materializations | `StoredPolicyRevision`, `StoredDraftChunk` | Server-only typed results assembled from indexed columns and decoded payloads; not public RPC messages. |
| Public messages used directly as encoded storage roots | `Sandbox`, `SandboxWorkloadTemplate`, `Provider`, `Workspace`, `WorkspaceMember`, `SshSession`, `ServiceEndpoint` | The generated public type is also the persisted payload. `SshSession` is not in the current public RPC message closure. |
| Embedded encoded root | `SandboxPolicy` | Stored in policy rows and inside the JSON settings envelope. |
| Configuration operation storage root | `StoredConfigUpdateOperation` | One common operation resource stores the exact provider target, receipt projection, snapshot failure reason, and historical outcome. |
| Public dependencies of an operation | `ConfigUpdateOperation`, `ProviderMutationReceipt`, `ProviderReadinessReason` | Their complete message and enum closures are durable contracts. |
The descriptor-derived test inventories the complete message and enum closure of the encoded durable roots and its intersection with the public RPC closure. The tables here record the reviewed roots and classifications.
The descriptor-derived test owns the complete public, durable, and intersecting inventories and their reviewed fingerprints. The tables here record their roots and classifications. A synthetic sandbox-spec byte fixture verifies that an absent server-owned attachment epoch decodes to the valid empty initial identity; direct and template creation tests separately require the gateway to replace any caller-supplied epoch.
Public delete, membership-removal, and SSH-revocation responses use
`DeletionOutcome`, not a transport-success boolean. `COMPLETED` establishes
@@ -439,6 +443,9 @@ require a request UUID for the admission contract.
| `SshSession` | Defer a storage twin; govern its complete dependency closure as durable. |
| `ServiceEndpoint` | Defer a storage twin; govern its complete dependency closure as durable. |
| `SandboxPolicy` | Defer a storage twin; govern its complete dependency closure as durable. |
| `ConfigUpdateOperation` | Persist the common historical outcome within `StoredConfigUpdateOperation`; govern its complete dependency closure as durable. |
| `ProviderMutationReceipt` | Persist the immutable provider projection within `StoredConfigUpdateOperation`; govern its complete dependency closure as durable. |
| `ProviderReadinessReason` | Persist only the closed snapshot failure category within the operation; govern its enum values as durable. |
The public/storage overlap is deliberate for the current format. Storage twins
for the public roots are deferred: introducing them would require a broad
@@ -739,6 +746,28 @@ configuration, valid endpoint-bound static credentials from other attached
providers, and the dynamic credential snapshot. Provider environment revisions
include profile endpoint and binding changes.
The supervisor owns provider fetching, support negotiation, and credential resolution outside the workload. The authenticated sandbox boundary receives a revision and its prepared child environment from one snapshot; it preserves the issued placeholders without receiving the secret resolver or turning those placeholders into new references.
Provider mutations return immutable per-sandbox receipts that separate saved desired state from observed runtime installation. A receipt pins sandbox and provider identity, the attachment epoch, and the exact provider/configuration/policy fingerprints. Attachment-set changes replace the epoch in the same sandbox CAS write; credential updates stage distinct backend objects before publishing their handles and provider revision. This prevents a published revision from referring to an unfinished in-place credential write. Update receipts retain the sandbox target set selected before publication.
Each receipt projects a common configuration operation in the `config_update_operation` store; its receipt ID is the operation ID. The provider status path evaluates current installation evidence and records historical terminal outcomes through resource-version CAS. A previously applied operation does not bypass current session, freshness, or authority checks. The live provider projection can be pending after disconnection or superseded after another change even when historical operation state remains applied. Pending operations are evaluated through provider status queries; this path adds no background delivery engine or mutation replay contract.
When a provider mutation opts into admission with `request_id`, its replay record retains references to the original configuration operations and the original mutation ID. Replay returns those immutable receipts even if sandbox attachments have since changed; it does not select new targets or create replacement receipts. Missing operation evidence makes replay unavailable without executing the mutation again. Provider resources still require their original recorded version, and sandbox resources retain the replay contract's current-state projection.
Provider mutation and operation-result writes are separate. A result-storage failure can follow a saved mutation and returns structured uncertainty without a rollback or safe-retry claim. A failed initial snapshot remains failed rather than acquiring a different target during a later lookup. Operations contain only identities, revisions, timestamps, and closed reason categories.
The CLI recognizes the gateway's `CONFIG_OPERATION_STORAGE_UNCERTAIN` error reason and domain for attach, detach, and update. It reports fixed guidance to inspect and reconcile the saved change before retrying, while withholding arbitrary server messages and error metadata. An uncertain mutation never starts a readiness wait or automatic replay.
Provider receipts, installation status, and common operations represent absolute times with protobuf `Timestamp`; report intervals and evidence lifetimes use protobuf `Duration`. Receipt identity compares the full canonical timestamp without truncating nanoseconds. An absent observation or completion time represents missing evidence or an unfinished operation, independently of the Unix epoch.
Provider installation reports belong to the existing `ConnectSupervisor` session. Each report names that session, has an increasing sequence, and expires unless the supervisor reports again. Reconnection or disconnect invalidates prior observations; stored change records survive a gateway restart, but runtime evidence does not. Replaying an identical report cannot extend its lifetime.
Reports and status also compare the supervisor instance with the sandbox's persisted current instance. A different supervisor becoming current invalidates an older connection, including one retained by another gateway replica. Observations stay local to the gateway holding the supervisor session; a status request reaching a replica without that session returns pending. Multi-replica deployments therefore retain the existing supervisor-session routing requirement.
The supervisor reports success only after it installs the matching credentials, activates the effective policy, and receives an acknowledgment from the authenticated workload boundary that it installed the environment for future processes. Environment synchronization shares the process-launch lock, and its acknowledgment identifies the exact publication, including retries at the same provider revision. Failed policy installation cannot reuse evidence for a different installed policy. Ready and revoked statuses also recheck the requested sandbox, provider, attachment and configuration identities; revision fingerprints are compared only for equality. Revocation applies to future credential resolution and future processes. Requests already forwarded upstream can still finish.
Ordinary static credentials retain revision-scoped references. After update readiness, a newly launched process receives the updated reference; an existing process keeps its original revision. Installation completion does not retarget that reference or establish that an old upstream key can be retired.
## Provider Environment Resolution
The gateway resolves only the providers attached to a sandbox. It combines each
+1
View File
@@ -4,3 +4,4 @@
pub mod common;
pub mod gateway;
pub mod provider;
pub mod provider_readiness;
+134 -55
View File
@@ -1,6 +1,9 @@
// SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
// SPDX-License-Identifier: Apache-2.0
use super::provider_readiness::{
ProviderMutationExpectation, ProviderWaitOptions, finish_provider_mutation,
};
use crate::color::Colorize;
use crate::commands::common::{
format_epoch_ms, format_optional_epoch_ms, parse_credential_expiry_pairs,
@@ -19,10 +22,11 @@ use openshell_core::proto::{
LintProviderProfilesRequest, ListProviderProfilesRequest, ListProvidersRequest,
ListSandboxProvidersRequest, Provider, ProviderCredentialRefreshRecoveryAction,
ProviderCredentialRefreshStatus, ProviderCredentialRefreshStrategy,
ProviderCredentialTokenGrantType, ProviderProfile, ProviderProfileDiagnostic,
ProviderProfileImportItem, RotateProviderCredentialRequest, UpdateProviderProfilesRequest,
UpdateProviderRequest,
ProviderCredentialTokenGrantType, ProviderMutationKind, ProviderProfile,
ProviderProfileDiagnostic, ProviderProfileImportItem, RotateProviderCredentialRequest,
UpdateProviderProfilesRequest, UpdateProviderRequest,
};
use openshell_core::rpc_error::{ERROR_DOMAIN, decode_details};
use openshell_core::{ObjectId, ObjectName, ObjectWorkspace};
use openshell_providers::{
ProviderTypeProfile, RealDiscoveryContext, detect_provider_from_command, discover_from_profile,
@@ -34,6 +38,30 @@ use std::io::IsTerminal;
use std::path::{Path, PathBuf};
use tonic::{Code, Status};
fn provider_mutation_is_uncertain(status: &Status) -> bool {
// Only a validated gateway ErrorInfo identifies a possibly saved mutation.
// Message text, metadata, foreign domains, and malformed details are untrusted.
decode_details(status).is_some_and(|details| {
details.error_info().is_some_and(|info| {
info.domain == ERROR_DOMAIN && info.reason == "CONFIG_OPERATION_STORAGE_UNCERTAIN"
})
})
}
fn provider_mutation_error(status: &Status, operation: &str) -> miette::Report {
if provider_mutation_is_uncertain(status) {
// Emit fixed guidance without chaining the server's potentially sensitive
// message or metadata. An error does not establish that the write rolled back.
miette!(
"provider change may already be saved (CONFIG_OPERATION_STORAGE_UNCERTAIN); \
readiness receipt could not be recorded. Do not blindly retry the mutation; \
check provider and sandbox status and reconcile the saved change first."
)
} else {
miette!("provider {operation} failed ({})", status.code())
}
}
fn proto_timestamp_ms(timestamp: Option<&prost_types::Timestamp>) -> i64 {
timestamp
.and_then(|value| openshell_core::time::timestamp_to_millis(value).ok())
@@ -84,13 +112,16 @@ pub async fn sandbox_provider_list(
Ok(())
}
/// Save a provider attachment and optionally wait for its installed authority.
pub async fn sandbox_provider_attach(
server: &str,
name: &str,
provider: &str,
workspace: &str,
tls: &TlsOptions,
readiness: ProviderWaitOptions<'_>,
) -> Result<()> {
readiness.validate()?;
let mut client = grpc_client(server, tls).await?;
// Fetch current sandbox to get resource_version for CAS
@@ -100,7 +131,7 @@ pub async fn sandbox_provider_attach(
workspace_scope: Some(openshell_core::proto::workspace_selector(workspace)),
})
.await
.into_diagnostic()?
.map_err(|status| miette!("provider attachment lookup failed ({})", status.code()))?
.into_inner()
.sandbox
.ok_or_else(|| miette::miette!("sandbox not found"))?;
@@ -118,36 +149,50 @@ pub async fn sandbox_provider_attach(
.await
{
Ok(response) => response.into_inner(),
Err(status) if status.code() == Code::Aborted => {
// Explicit post-save uncertainty takes precedence over a generic retry hint.
Err(status)
if status.code() == Code::Aborted && !provider_mutation_is_uncertain(&status) =>
{
return Err(miette::miette!(
"Failed to attach provider: sandbox was modified by another operation.\n\
Please retry the command."
)
.with_source_code(status.message().to_string()));
));
}
Err(e) => return Err(e).into_diagnostic(),
Err(error) => return Err(provider_mutation_error(&error, "attachment")),
};
if response.attached {
println!(
"{} Attached provider {} to sandbox {}",
"✓".green().bold(),
provider,
name
);
} else {
println!("Provider {provider} is already attached to sandbox {name}.");
}
Ok(())
let receipt = response.receipt.ok_or_else(|| {
miette!(
"gateway did not return a provider receipt; saved attachment cannot establish readiness"
)
})?;
let mutation_id = receipt.mutation_id.clone();
finish_provider_mutation(
&client,
&mutation_id,
vec![receipt],
ProviderMutationExpectation {
workspace,
provider_name: provider,
kind: ProviderMutationKind::Attach,
sandbox: Some((name, sandbox.object_id())),
provider: None,
},
readiness,
)
.await
}
/// Save a provider detachment and optionally wait for future authority revocation.
pub async fn sandbox_provider_detach(
server: &str,
name: &str,
provider: &str,
workspace: &str,
tls: &TlsOptions,
readiness: ProviderWaitOptions<'_>,
) -> Result<()> {
readiness.validate()?;
let mut client = grpc_client(server, tls).await?;
// Fetch current sandbox to get resource_version for CAS
@@ -157,7 +202,7 @@ pub async fn sandbox_provider_detach(
workspace_scope: Some(openshell_core::proto::workspace_selector(workspace)),
})
.await
.into_diagnostic()?
.map_err(|status| miette!("provider detachment lookup failed ({})", status.code()))?
.into_inner()
.sandbox
.ok_or_else(|| miette::miette!("sandbox not found"))?;
@@ -175,27 +220,34 @@ pub async fn sandbox_provider_detach(
.await
{
Ok(response) => response.into_inner(),
Err(status) if status.code() == Code::Aborted => {
// Explicit post-save uncertainty takes precedence over a generic retry hint.
Err(status)
if status.code() == Code::Aborted && !provider_mutation_is_uncertain(&status) =>
{
return Err(miette::miette!(
"Failed to detach provider: sandbox was modified by another operation.\n\
Please retry the command."
)
.with_source_code(status.message().to_string()));
));
}
Err(e) => return Err(e).into_diagnostic(),
Err(error) => return Err(provider_mutation_error(&error, "detachment")),
};
if response.detached {
println!(
"{} Detached provider {} from sandbox {}",
"✓".green().bold(),
provider,
name
);
} else {
println!("Provider {provider} was not attached to sandbox {name}.");
}
Ok(())
let receipt = response.receipt.ok_or_else(|| miette!("gateway did not return a provider receipt; saved detachment cannot establish revocation"))?;
let mutation_id = receipt.mutation_id.clone();
finish_provider_mutation(
&client,
&mutation_id,
vec![receipt],
ProviderMutationExpectation {
workspace,
provider_name: provider,
kind: ProviderMutationKind::Detach,
sandbox: Some((name, sandbox.object_id())),
provider: None,
},
readiness,
)
.await
}
fn print_provider_attachment_table(providers: &[Provider]) {
@@ -710,6 +762,19 @@ async fn rollback_provider_create_after_gcloud_adc_failure(
}
}
fn provider_profile_lookup_error(status: &Status) -> miette::Report {
// A permission code supports recovery guidance, but cannot distinguish a
// missing membership from an insufficient role. Never expose backend text.
if status.code() == Code::PermissionDenied {
miette!(
"provider profile lookup denied (PERMISSION_DENIED): \
verify workspace membership and required permissions"
)
} else {
miette!("provider profile lookup failed ({})", status.code())
}
}
async fn fetch_provider_profile(
client: &mut crate::tls::GrpcClient,
provider_type: &str,
@@ -734,11 +799,11 @@ async fn fetch_provider_profile(
"provider profile '{requested}' not found; import a matching profile before using this provider type"
)
} else {
miette::miette!(fallback_status.to_string())
provider_profile_lookup_error(&fallback_status)
}
})?
}
Err(status) => return Err(miette::miette!(status.to_string())),
Err(status) => return Err(provider_profile_lookup_error(&status)),
};
Ok(response)
@@ -1025,11 +1090,10 @@ pub async fn provider_create_with_options(options: ProviderCreateOptions<'_>) ->
if profile_id.is_empty() {
return Err(miette::miette!("provider type is required"));
}
let provider_profile = fetch_provider_profile(&mut client, profile_id, profile_workspace)
.await
.map_err(|err| {
miette::miette!("unsupported provider type or profile: {profile_id} ({err})")
})?;
// Lookup already distinguishes absent profiles from permission and transport
// failures; those failures do not establish that the profile is unsupported.
let provider_profile =
fetch_provider_profile(&mut client, profile_id, profile_workspace).await?;
let provider_type = provider_profile.id.clone();
let adc_credential_key = if from_gcloud_adc {
@@ -2237,6 +2301,7 @@ fn print_provider_type_row(
);
}
/// Credential update inputs and observation choices for all attached sandboxes.
pub struct ProviderUpdateOptions<'a> {
pub server: &'a str,
pub name: &'a str,
@@ -2247,8 +2312,11 @@ pub struct ProviderUpdateOptions<'a> {
pub credential_expires_at: &'a [String],
pub workspace: &'a str,
pub tls: &'a TlsOptions,
/// Bound observation of the sandbox target set selected by the update.
pub readiness: ProviderWaitOptions<'a>,
}
/// Update provider credentials and report each selected sandbox independently.
pub async fn provider_update(options: ProviderUpdateOptions<'_>) -> Result<()> {
let ProviderUpdateOptions {
server,
@@ -2260,7 +2328,9 @@ pub async fn provider_update(options: ProviderUpdateOptions<'_>) -> Result<()> {
credential_expires_at,
workspace,
tls,
readiness,
} = options;
readiness.validate()?;
if from_existing && !credentials.is_empty() {
return Err(miette::miette!(
@@ -2298,7 +2368,7 @@ pub async fn provider_update(options: ProviderUpdateOptions<'_>) -> Result<()> {
{
None
}
Err(status) => return Err(status).into_diagnostic(),
Err(status) => return Err(miette!("provider update lookup failed ({})", status.code())),
};
if existing.is_none() && (from_existing || from_oidc_token) {
@@ -2388,19 +2458,28 @@ pub async fn provider_update(options: ProviderUpdateOptions<'_>) -> Result<()> {
clear_credential_expiration_keys,
})
.await
.into_diagnostic()?;
.map_err(|error| provider_mutation_error(&error, "update"))?;
let provider = response
.into_inner()
.provider
.ok_or_else(|| miette::miette!("provider missing from response"))?;
println!(
"{} Updated provider {}",
"✓".green().bold(),
provider.object_name()
);
Ok(())
let response = response.into_inner();
if response.mutation_id.is_empty() {
return Err(miette!(
"gateway did not return a provider mutation receipt; saved credentials cannot establish readiness"
));
}
finish_provider_mutation(
&client,
&response.mutation_id,
response.target_receipts,
ProviderMutationExpectation {
workspace,
provider_name: name,
kind: ProviderMutationKind::Update,
sandbox: None,
provider: response.provider.as_ref(),
},
readiness,
)
.await
}
pub async fn provider_delete(
@@ -0,0 +1,584 @@
// SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
// SPDX-License-Identifier: Apache-2.0
//! Secret-free provider receipts, status rendering, and bounded CLI waits.
use crate::tls::{GrpcClient, TlsOptions, grpc_client};
use futures::{StreamExt, stream};
use miette::{IntoDiagnostic, Result, miette};
use openshell_core::proto::{
GetSandboxProviderStatusRequest, Provider, ProviderMutationKind, ProviderMutationReceipt,
ProviderReadinessReason, ProviderReadinessState, ProviderReadinessStatus,
};
use openshell_sdk::provider_readiness::{
ProviderWaitOutcome, persisted_status, provider_status, provider_wait_deadline,
validate_provider_receipt, wait_for_provider_status_until,
};
use std::collections::HashSet;
use std::time::Duration;
use tokio::time::Instant;
const PROVIDER_WAIT_CONCURRENCY: usize = 16;
/// Output and deadline choices shared by provider mutation and status commands.
#[derive(Clone, Copy)]
pub struct ProviderWaitOptions<'a> {
/// Wait for current runtime installation after saving the desired mutation.
pub wait: bool,
/// One total deadline for all selected sandbox observations.
pub timeout: Duration,
/// CLI output format: table, JSON, or YAML.
pub output: &'a str,
}
impl Default for ProviderWaitOptions<'_> {
fn default() -> Self {
Self {
wait: false,
timeout: Duration::from_secs(30),
output: "table",
}
}
}
impl ProviderWaitOptions<'_> {
/// Reject invalid wait settings before performing a provider mutation.
pub fn validate(self) -> Result<()> {
provider_wait_deadline(self.timeout).into_diagnostic()?;
if !matches!(self.output, "table" | "json" | "yaml") {
return Err(miette!("unsupported provider status output format"));
}
Ok(())
}
}
struct DisplayStatus {
status: ProviderReadinessStatus,
outcome: &'static str,
complete: bool,
}
/// Caller-selected authority that every returned mutation receipt must retain.
#[derive(Clone, Copy)]
pub(super) struct ProviderMutationExpectation<'a> {
/// Workspace explicitly selected by the command.
pub workspace: &'a str,
/// Provider name explicitly selected by the command.
pub provider_name: &'a str,
/// Operation submitted to the gateway.
pub kind: ProviderMutationKind,
/// Requested sandbox name and previously fetched object ID for attach/detach.
pub sandbox: Option<(&'a str, &'a str)>,
/// Published provider object returned by an update, including its revision.
pub provider: Option<&'a Provider>,
}
fn validate_mutation_receipts(
mutation_id: &str,
receipts: &[ProviderMutationReceipt],
expected: &ProviderMutationExpectation<'_>,
) -> Result<()> {
let invalid = || miette!("gateway returned an invalid provider mutation receipt");
if mutation_id.is_empty() || expected.workspace.is_empty() || expected.provider_name.is_empty()
{
return Err(invalid());
}
let provider = match expected.kind {
ProviderMutationKind::Attach | ProviderMutationKind::Detach => {
if receipts.len() != 1
|| expected
.sandbox
.is_none_or(|(name, id)| name.is_empty() || id.is_empty())
{
return Err(invalid());
}
None
}
ProviderMutationKind::Update => {
let metadata = expected
.provider
.and_then(|provider| provider.metadata.as_ref())
.ok_or_else(invalid)?;
if metadata.id.is_empty()
|| metadata.name != expected.provider_name
|| metadata.workspace != expected.workspace
{
return Err(invalid());
}
Some(metadata)
}
ProviderMutationKind::Unspecified | ProviderMutationKind::Observe => return Err(invalid()),
};
let mut receipt_ids = HashSet::new();
let mut sandbox_ids = HashSet::new();
for receipt in receipts {
validate_provider_receipt(receipt).into_diagnostic()?;
let desired = receipt.desired.as_ref().ok_or_else(invalid)?;
// Validate the complete batch before polling or printing saved intent.
// Otherwise a substituted receipt can redirect a wait to another scope.
if receipt.mutation_id != mutation_id
|| receipt.workspace != expected.workspace
|| receipt.provider_name != expected.provider_name
|| receipt.kind != i32::from(expected.kind)
|| !receipt_ids.insert(receipt.receipt_id.as_str())
|| !sandbox_ids.insert(desired.sandbox_id.as_str())
|| expected
.sandbox
.is_some_and(|(name, id)| desired.sandbox_name != name || desired.sandbox_id != id)
|| provider.is_some_and(|provider| {
desired.provider_id != provider.id
|| desired.provider_resource_version != provider.resource_version
})
{
return Err(invalid());
}
}
Ok(())
}
/// Display persisted targets and optionally wait for each original receipt.
///
/// Every target is reported even when another target fails. Concurrency is
/// bounded and all requests share one deadline, including queued targets.
pub(super) async fn finish_provider_mutation(
client: &GrpcClient,
mutation_id: &str,
receipts: Vec<ProviderMutationReceipt>,
expected: ProviderMutationExpectation<'_>,
options: ProviderWaitOptions<'_>,
) -> Result<()> {
options.validate()?;
validate_mutation_receipts(mutation_id, &receipts, &expected)?;
let deadline = provider_wait_deadline(options.timeout).into_diagnostic()?;
let statuses = receipts.into_iter().map(persisted_status).collect();
let results = if options.wait {
wait_for_mutation_statuses(client, statuses, deadline).await
} else {
statuses
.into_iter()
.map(|status| DisplayStatus {
status,
outcome: "not_requested",
complete: false,
})
.collect()
};
print_statuses(mutation_id, &results, options.output)?;
if options.wait && results.iter().any(|result| !result.complete) {
return Err(miette!(
"provider readiness wait did not complete for every selected sandbox; inspect the reported outcomes"
));
}
Ok(())
}
async fn wait_for_mutation_statuses(
client: &GrpcClient,
statuses: Vec<ProviderReadinessStatus>,
deadline: Instant,
) -> Vec<DisplayStatus> {
let mut pending = statuses.into_iter().enumerate().collect::<Vec<_>>();
let mut results = Vec::with_capacity(pending.len());
while !pending.is_empty() && Instant::now() < deadline {
// Divide the remaining time among actual queued batches. When every
// target fits in one batch, healthy but slow RPCs get the full deadline;
// queued targets still get a turn if the preceding batch stalls.
let batches =
u32::try_from(pending.len().div_ceil(PROVIDER_WAIT_CONCURRENCY)).unwrap_or(u32::MAX);
let slice = deadline.saturating_duration_since(Instant::now()) / batches;
if slice.is_zero() {
break;
}
let round = stream::iter(pending.into_iter().map(|(index, status)| {
let mut client = client.clone();
async move {
let slice_deadline = (Instant::now() + slice).min(deadline);
let result =
wait_for_provider_status_until(&mut client, &status, slice_deadline).await;
(index, status, result)
}
}))
.buffer_unordered(PROVIDER_WAIT_CONCURRENCY)
.collect::<Vec<_>>()
.await;
pending = Vec::new();
for (index, previous, result) in round {
match result {
// The SDK pins the receipt in this status and retains its last
// observation. Only unfinished targets enter the next round.
Ok(result)
if result.outcome == ProviderWaitOutcome::TimedOut
&& Instant::now() < deadline =>
{
pending.push((index, result.status));
}
Ok(result) => {
results.push((
index,
DisplayStatus {
status: result.status,
outcome: match result.outcome {
ProviderWaitOutcome::Complete => "complete",
ProviderWaitOutcome::TimedOut => "timed_out",
ProviderWaitOutcome::Terminal => "terminal",
},
complete: result.outcome == ProviderWaitOutcome::Complete,
},
));
}
// A transport error cannot prove installation failure. Preserve
// the last observation and expose only the safe error category.
Err(_) => {
results.push((
index,
DisplayStatus {
status: previous,
outcome: "observation_error",
complete: false,
},
));
}
}
}
}
results.extend(pending.into_iter().map(|(index, status)| {
(
index,
DisplayStatus {
status,
outcome: "timed_out",
complete: false,
},
)
}));
// Fair polling must not change the command's stable target output order.
results.sort_unstable_by_key(|(index, _)| *index);
results.into_iter().map(|(_, result)| result).collect()
}
/// Show desired/observed provider state, optionally waiting for that exact state.
pub async fn sandbox_provider_status(
server: &str,
name: &str,
provider: &str,
receipt_id: &str,
workspace: &str,
tls: &TlsOptions,
options: ProviderWaitOptions<'_>,
) -> Result<()> {
options.validate()?;
let mut client = grpc_client(server, tls).await?;
let deadline = provider_wait_deadline(options.timeout).into_diagnostic()?;
let status = tokio::time::timeout_at(
deadline,
provider_status(
&mut client,
GetSandboxProviderStatusRequest {
sandbox_name: name.to_string(),
provider_name: provider.to_string(),
receipt_id: receipt_id.to_string(),
workspace_scope: Some(openshell_core::proto::workspace_selector(workspace)),
},
),
)
.await
.map_err(|_| miette!("provider status request timed out"))?
.into_diagnostic()?;
let receipt = status
.receipt
.as_ref()
.ok_or_else(|| miette!("gateway returned no provider receipt"))?;
let mutation_id = receipt.mutation_id.clone();
let result = if options.wait {
let result = wait_for_provider_status_until(&mut client, &status, deadline)
.await
.into_diagnostic()?;
DisplayStatus {
status: result.status,
outcome: match result.outcome {
ProviderWaitOutcome::Complete => "complete",
ProviderWaitOutcome::TimedOut => "timed_out",
ProviderWaitOutcome::Terminal => "terminal",
},
complete: result.outcome == ProviderWaitOutcome::Complete,
}
} else {
DisplayStatus {
status,
outcome: "not_requested",
complete: false,
}
};
let complete = result.complete;
print_statuses(&mutation_id, &[result], options.output)?;
if options.wait && !complete {
return Err(miette!(
"provider readiness wait did not complete; inspect the reported outcome"
));
}
Ok(())
}
fn state_label(state: i32) -> &'static str {
match ProviderReadinessState::try_from(state) {
Ok(ProviderReadinessState::Persisted) => "persisted",
Ok(ProviderReadinessState::Pending) => "pending",
Ok(ProviderReadinessState::Ready) => "ready",
Ok(ProviderReadinessState::Withheld) => "withheld",
Ok(ProviderReadinessState::Revoked) => "revoked",
Ok(ProviderReadinessState::Failed) => "failed",
Ok(ProviderReadinessState::Superseded) => "superseded",
_ => "unknown",
}
}
fn reason_label(reason: i32) -> String {
ProviderReadinessReason::try_from(reason).map_or_else(
|_| "unknown".to_string(),
|reason| {
reason
.as_str_name()
.trim_start_matches("PROVIDER_READINESS_REASON_")
.to_ascii_lowercase()
},
)
}
fn receipt_json(receipt: &ProviderMutationReceipt) -> serde_json::Value {
let desired = receipt.desired.as_ref().map(|desired| {
serde_json::json!({
"sandbox_id": desired.sandbox_id,
"sandbox_name": desired.sandbox_name,
"attachment_epoch": desired.attachment_epoch,
"provider_id": desired.provider_id,
"provider_resource_version": desired.provider_resource_version.to_string(),
"provider_env_revision": desired.provider_env_revision.to_string(),
"config_revision": desired.config_revision.to_string(),
"policy_hash": desired.policy_hash,
})
});
serde_json::json!({
"receipt_id": receipt.receipt_id,
"mutation_id": receipt.mutation_id,
"provider_name": receipt.provider_name,
"workspace": receipt.workspace,
"kind": receipt.kind().as_str_name().trim_start_matches("PROVIDER_MUTATION_KIND_").to_ascii_lowercase(),
"desired": desired,
"persisted_time": receipt.persisted_time.as_ref().map(ToString::to_string),
})
}
fn status_json(result: &DisplayStatus) -> serde_json::Value {
let observed = result.status.observed.as_ref().map(|observed| {
serde_json::json!({
"session_id": observed.session_id,
"sequence": observed.sequence.to_string(),
"attachment_epoch": observed.attachment_epoch,
"provider_env_revision": observed.provider_env_revision.to_string(),
"config_revision": observed.config_revision.to_string(),
"policy_hash": observed.policy_hash,
"credentials_installed": observed.credentials_installed,
"policy_active": observed.policy_active,
"launch_environment_installed": observed.launch_environment_installed,
"process_instance_id": observed.process_instance_id,
"reason": reason_label(observed.reason),
})
});
serde_json::json!({
"receipt": result.status.receipt.as_ref().map(receipt_json),
"state": state_label(result.status.state),
"reason": reason_label(result.status.reason),
"observed": observed,
"network_instance_id": result.status.network_instance_id,
"observed_time": result.status.observed_time.as_ref().map(ToString::to_string),
"evaluated_time": result.status.evaluated_time.as_ref().map(ToString::to_string),
"wait_outcome": result.outcome,
})
}
fn print_statuses(mutation_id: &str, results: &[DisplayStatus], output: &str) -> Result<()> {
let value = serde_json::json!({
"mutation_id": mutation_id,
"targets": results.iter().map(status_json).collect::<Vec<_>>(),
});
if crate::output::print_output_single(output, &value, Clone::clone)? {
return Ok(());
}
if results.is_empty() {
println!("Provider mutation {mutation_id} persisted; no attached sandboxes were selected.");
return Ok(());
}
for result in results {
let Some(receipt) = result.status.receipt.as_ref() else {
continue;
};
let sandbox = receipt
.desired
.as_ref()
.map_or("unknown", |desired| desired.sandbox_name.as_str());
println!(
"{} / {}: {} ({})",
sandbox,
receipt.provider_name,
state_label(result.status.state),
reason_label(result.status.reason)
);
println!(
" Receipt: {}",
if receipt.receipt_id.is_empty() {
"current state"
} else {
&receipt.receipt_id
}
);
println!(
" Persisted: {}",
receipt
.persisted_time
.as_ref()
.map_or_else(|| "-".to_string(), ToString::to_string)
);
if let Some(observed) = result.status.observed.as_ref() {
println!(
" Installed: credentials={}, policy={}, future process environment={}",
observed.credentials_installed,
observed.policy_active,
observed.launch_environment_installed
);
}
if result.outcome != "not_requested" {
println!(" Wait: {}", result.outcome);
}
}
Ok(())
}
#[cfg(test)]
mod tests {
use super::*;
use openshell_core::proto::ProviderDesiredIdentity;
#[test]
fn mutation_receipts_reject_duplicate_targets_and_allow_initial_fingerprints() {
let provider = Provider {
metadata: Some(openshell_core::proto::datamodel::v1::ObjectMeta {
id: "provider-id".to_string(),
name: "provider".to_string(),
workspace: "default".to_string(),
resource_version: 0,
..Default::default()
}),
..Default::default()
};
let expected = ProviderMutationExpectation {
workspace: "default",
provider_name: "provider",
kind: ProviderMutationKind::Update,
sandbox: None,
provider: Some(&provider),
};
let receipt = ProviderMutationReceipt {
receipt_id: "receipt".to_string(),
mutation_id: "mutation".to_string(),
provider_name: "provider".to_string(),
workspace: "default".to_string(),
kind: ProviderMutationKind::Update.into(),
persisted_time: Some(openshell_core::time::timestamp_from_millis(1).unwrap()),
desired: Some(ProviderDesiredIdentity {
sandbox_id: "sandbox-id".to_string(),
sandbox_name: "sandbox".to_string(),
provider_id: "provider-id".to_string(),
..Default::default()
}),
};
assert!(
validate_mutation_receipts("mutation", std::slice::from_ref(&receipt), &expected)
.is_ok()
);
assert!(validate_mutation_receipts("mutation", &[], &expected).is_ok());
let mut duplicate = receipt.clone();
duplicate.receipt_id = "other-receipt".to_string();
assert!(
validate_mutation_receipts("mutation", &[receipt.clone(), duplicate], &expected)
.is_err()
);
let mut duplicate = receipt.clone();
duplicate
.desired
.as_mut()
.expect("desired identity")
.sandbox_id = "other-sandbox".to_string();
assert!(validate_mutation_receipts("mutation", &[receipt, duplicate], &expected).is_err());
assert!(
validate_mutation_receipts(
"mutation",
&[],
&ProviderMutationExpectation {
provider: None,
..expected
}
)
.is_err()
);
}
#[test]
fn structured_status_renders_rfc3339_timestamps_without_losing_nanos() {
let timestamp = prost_types::Timestamp {
seconds: 1_000,
nanos: 123_456_789,
};
let receipt = ProviderMutationReceipt {
persisted_time: Some(timestamp),
..Default::default()
};
let mut status = persisted_status(receipt);
status.observed_time = Some(timestamp);
status.evaluated_time = Some(timestamp);
let value = status_json(&DisplayStatus {
status,
outcome: "not_requested",
complete: false,
});
let expected = "1970-01-01T00:16:40.123456789Z";
assert_eq!(value["receipt"]["persisted_time"], expected);
assert_eq!(value["observed_time"], expected);
assert_eq!(value["evaluated_time"], expected);
assert!(value["receipt"].get("persisted_at_ms").is_none());
assert!(value.get("observed_at_ms").is_none());
assert!(value.get("evaluated_at_ms").is_none());
let value = status_json(&DisplayStatus {
status: persisted_status(ProviderMutationReceipt::default()),
outcome: "not_requested",
complete: false,
});
assert!(value["receipt"]["persisted_time"].is_null());
assert!(value["observed_time"].is_null());
assert!(value["evaluated_time"].is_null());
}
#[test]
fn structured_status_preserves_opaque_revisions_and_separates_timeout() {
let receipt = ProviderMutationReceipt {
desired: Some(ProviderDesiredIdentity {
provider_env_revision: u64::MAX,
..Default::default()
}),
..Default::default()
};
let value = status_json(&DisplayStatus {
status: persisted_status(receipt),
outcome: "timed_out",
complete: false,
});
assert_eq!(value["state"], "persisted");
assert_eq!(value["wait_outcome"], "timed_out");
assert_eq!(
value["receipt"]["desired"]["provider_env_revision"],
u64::MAX.to_string()
);
assert!(value.get("environment").is_none());
assert!(value.get("load_error").is_none());
}
}
+150 -3
View File
@@ -944,6 +944,9 @@ enum ProviderCommands {
/// Credential expiry (`KEY=TIMESTAMP`). Accepts epoch milliseconds or RFC3339. A zero timestamp clears expiry.
#[arg(long = "credential-expires-at", value_name = "KEY=TIMESTAMP")]
credential_expires_at: Vec<String>,
#[command(flatten)]
readiness: ProviderReadinessArgs,
},
/// Delete providers by name.
@@ -1624,6 +1627,32 @@ enum SandboxCommands {
Template(SandboxTemplateCommands),
}
/// Common observation flags; the deadline starts after the mutation is saved.
#[derive(clap::Args, Debug)]
struct ProviderReadinessArgs {
/// Wait until the sandbox applies the credentials, policy, and environment for new processes.
#[arg(long)]
wait: bool,
/// Maximum wait in seconds after saving the change, shared by all selected sandboxes.
#[arg(long, default_value_t = 30, value_parser = clap::value_parser!(u64).range(1..=3600))]
timeout: u64,
/// Output format; JSON and YAML include change IDs and results for each sandbox.
#[arg(short = 'o', long = "output", value_enum, default_value_t = OutputFormat::Table)]
output: OutputFormat,
}
impl ProviderReadinessArgs {
fn options(&self) -> run::ProviderWaitOptions<'_> {
run::ProviderWaitOptions {
wait: self.wait,
timeout: std::time::Duration::from_secs(self.timeout),
output: self.output.as_str(),
}
}
}
#[derive(Subcommand, Debug)]
enum SandboxProviderCommands {
/// List providers attached to a sandbox.
@@ -1648,6 +1677,9 @@ enum SandboxProviderCommands {
/// Provider name to attach.
#[arg(add = ArgValueCompleter::new(completers::complete_provider_names))]
provider: String,
#[command(flatten)]
readiness: ProviderReadinessArgs,
},
/// Detach a provider from a sandbox.
@@ -1660,6 +1692,28 @@ enum SandboxProviderCommands {
/// Provider name to detach.
#[arg(add = ArgValueCompleter::new(completers::complete_provider_names))]
provider: String,
#[command(flatten)]
readiness: ProviderReadinessArgs,
},
/// Check whether a sandbox has applied a provider change.
#[command(help_template = LEAF_HELP_TEMPLATE, next_help_heading = "FLAGS")]
Status {
/// Sandbox name.
#[arg(add = ArgValueCompleter::new(completers::complete_sandbox_names))]
name: String,
/// Provider whose attachment or revocation is being inspected.
#[arg(add = ArgValueCompleter::new(completers::complete_provider_names))]
provider: String,
/// Change ID returned by attach, detach, or update; omitted means the latest saved state.
#[arg(long = "receipt")]
receipt: Option<String>,
#[command(flatten)]
readiness: ProviderReadinessArgs,
},
}
@@ -3356,23 +3410,50 @@ async fn run_async() -> Result<()> {
)
.await?;
}
SandboxProviderCommands::Attach { name, provider } => {
SandboxProviderCommands::Attach {
name,
provider,
readiness,
} => {
run::sandbox_provider_attach(
endpoint,
&name,
&provider,
&cli.workspace,
&tls,
readiness.options(),
)
.await?;
}
SandboxProviderCommands::Detach { name, provider } => {
SandboxProviderCommands::Detach {
name,
provider,
readiness,
} => {
run::sandbox_provider_detach(
endpoint,
&name,
&provider,
&cli.workspace,
&tls,
readiness.options(),
)
.await?;
}
SandboxProviderCommands::Status {
name,
provider,
receipt,
readiness,
} => {
run::sandbox_provider_status(
endpoint,
&name,
&provider,
receipt.as_deref().unwrap_or_default(),
&cli.workspace,
&tls,
readiness.options(),
)
.await?;
}
@@ -3733,6 +3814,7 @@ async fn run_async() -> Result<()> {
credentials,
config,
credential_expires_at,
readiness,
} => {
run::provider_update(run::ProviderUpdateOptions {
server: endpoint,
@@ -3744,6 +3826,7 @@ async fn run_async() -> Result<()> {
credential_expires_at: &credential_expires_at,
workspace: &cli.workspace,
tls: &tls,
readiness: readiness.options(),
})
.await?;
}
@@ -4030,7 +4113,9 @@ mod tests {
let Some(Commands::Sandbox {
command:
Some(SandboxCommands::Provider(SandboxProviderCommands::Attach { name, provider })),
Some(SandboxCommands::Provider(SandboxProviderCommands::Attach {
name, provider, ..
})),
}) = cli.command
else {
panic!("expected sandbox provider attach command");
@@ -4040,6 +4125,68 @@ mod tests {
assert_eq!(provider, "work-github");
}
#[test]
fn provider_readiness_commands_have_bounded_waits_and_structured_output() {
for action in ["attach", "detach", "status"] {
let cli = Cli::try_parse_from([
"openshell",
"sandbox",
"provider",
action,
"sandbox",
"provider",
"--wait",
"--timeout",
"45",
"--output",
"json",
])
.expect("readiness flags should parse");
let Some(Commands::Sandbox {
command: Some(SandboxCommands::Provider(command)),
}) = cli.command
else {
panic!("expected provider command")
};
let (SandboxProviderCommands::Attach { readiness, .. }
| SandboxProviderCommands::Detach { readiness, .. }
| SandboxProviderCommands::Status { readiness, .. }) = command
else {
panic!("expected readiness command");
};
assert!(readiness.wait);
assert_eq!(readiness.timeout, 45);
assert_eq!(readiness.output.as_str(), "json");
}
for timeout in ["0", "3601"] {
assert!(
Cli::try_parse_from([
"openshell",
"provider",
"update",
"provider",
"--wait",
"--timeout",
timeout,
])
.is_err()
);
}
assert!(
Cli::try_parse_from([
"openshell",
"sandbox",
"provider",
"status",
"sandbox",
"provider",
"--receipt",
"receipt-id",
])
.is_ok()
);
}
#[test]
fn completions_policy_flag_falls_back_to_file_paths() {
let temp = tempfile::tempdir().expect("failed to create tempdir");
+1
View File
@@ -32,6 +32,7 @@ pub use crate::commands::provider::{
provider_refresh_status, provider_rotate, provider_update, sandbox_provider_attach,
sandbox_provider_detach, sandbox_provider_list,
};
pub use crate::commands::provider_readiness::{ProviderWaitOptions, sandbox_provider_status};
use crate::color::Colorize;
use crate::policy_update::build_policy_update_plan;
@@ -215,6 +215,24 @@ impl OpenShell for TestOpenShell {
Ok(Response::new(GetGatewayConfigResponse::default()))
}
async fn get_sandbox_provider_status(
&self,
_request: tonic::Request<openshell_core::proto::GetSandboxProviderStatusRequest>,
) -> Result<Response<openshell_core::proto::GetSandboxProviderStatusResponse>, Status> {
Err(Status::unimplemented(
"provider readiness is not exercised by this mock",
))
}
async fn report_provider_readiness(
&self,
_request: tonic::Request<openshell_core::proto::ReportProviderReadinessRequest>,
) -> Result<Response<openshell_core::proto::ReportProviderReadinessResponse>, Status> {
Err(Status::unimplemented(
"provider installation reports are not exercised by this mock",
))
}
async fn get_sandbox_provider_environment(
&self,
_request: tonic::Request<GetSandboxProviderEnvironmentRequest>,
@@ -296,6 +314,7 @@ impl OpenShell for TestOpenShell {
providers.insert(provider_name, provider.clone());
Ok(Response::new(ProviderResponse {
provider: Some(provider),
..Default::default()
}))
}
@@ -311,6 +330,7 @@ impl OpenShell for TestOpenShell {
.ok_or_else(|| Status::not_found("provider not found"))?;
Ok(Response::new(ProviderResponse {
provider: Some(provider),
..Default::default()
}))
}
@@ -463,6 +483,7 @@ impl OpenShell for TestOpenShell {
providers.insert(updated_name, updated.clone());
Ok(Response::new(ProviderResponse {
provider: Some(updated),
..Default::default()
}))
}
async fn get_provider_refresh_status(
@@ -186,6 +186,24 @@ impl OpenShell for TestOpenShell {
))
}
async fn get_sandbox_provider_status(
&self,
_request: tonic::Request<openshell_core::proto::GetSandboxProviderStatusRequest>,
) -> Result<Response<openshell_core::proto::GetSandboxProviderStatusResponse>, Status> {
Err(Status::unimplemented(
"provider readiness is not exercised by this mock",
))
}
async fn report_provider_readiness(
&self,
_request: tonic::Request<openshell_core::proto::ReportProviderReadinessRequest>,
) -> Result<Response<openshell_core::proto::ReportProviderReadinessResponse>, Status> {
Err(Status::unimplemented(
"provider installation reports are not exercised by this mock",
))
}
async fn get_sandbox_provider_environment(
&self,
_request: tonic::Request<openshell_core::proto::GetSandboxProviderEnvironmentRequest>,
File diff suppressed because it is too large Load Diff
@@ -367,6 +367,24 @@ impl OpenShell for TestOpenShell {
}))
}
async fn get_sandbox_provider_status(
&self,
_request: tonic::Request<openshell_core::proto::GetSandboxProviderStatusRequest>,
) -> Result<Response<openshell_core::proto::GetSandboxProviderStatusResponse>, Status> {
Err(Status::unimplemented(
"provider readiness is not exercised by this mock",
))
}
async fn report_provider_readiness(
&self,
_request: tonic::Request<openshell_core::proto::ReportProviderReadinessRequest>,
) -> Result<Response<openshell_core::proto::ReportProviderReadinessResponse>, Status> {
Err(Status::unimplemented(
"provider installation reports are not exercised by this mock",
))
}
async fn get_sandbox_provider_environment(
&self,
_request: tonic::Request<GetSandboxProviderEnvironmentRequest>,
@@ -248,6 +248,24 @@ impl OpenShell for TestOpenShell {
Ok(Response::new(GetGatewayConfigResponse::default()))
}
async fn get_sandbox_provider_status(
&self,
_request: tonic::Request<openshell_core::proto::GetSandboxProviderStatusRequest>,
) -> Result<Response<openshell_core::proto::GetSandboxProviderStatusResponse>, Status> {
Err(Status::unimplemented(
"provider readiness is not exercised by this mock",
))
}
async fn report_provider_readiness(
&self,
_request: tonic::Request<openshell_core::proto::ReportProviderReadinessRequest>,
) -> Result<Response<openshell_core::proto::ReportProviderReadinessResponse>, Status> {
Err(Status::unimplemented(
"provider installation reports are not exercised by this mock",
))
}
async fn get_sandbox_provider_environment(
&self,
_request: tonic::Request<GetSandboxProviderEnvironmentRequest>,
+119 -8
View File
@@ -27,10 +27,11 @@ use crate::proto::{
DenialSummary, EndpointObservation as ProtoEndpointObservation,
EndpointResult as ProtoEndpointResult, ExchangeProviderSubjectTokenRequest,
GetDraftPolicyRequest, GetSandboxConfigRequest, GetSandboxProviderEnvironmentRequest,
IssueSandboxTokenRequest, NetworkActivitySummary, PolicyChunk, PolicySource, PolicyStatus,
RefreshSandboxTokenRequest, ReportEndpointStatusRequest, ReportPolicyStatusRequest,
SandboxPolicy as ProtoSandboxPolicy, SubmitPolicyAnalysisRequest, SubmitPolicyAnalysisResponse,
UpdateConfigRequest, open_shell_client::OpenShellClient, workspace_selector,
GetSandboxProviderEnvironmentResponse, IssueSandboxTokenRequest, NetworkActivitySummary,
PolicyChunk, PolicySource, PolicyStatus, RefreshSandboxTokenRequest,
ReportEndpointStatusRequest, ReportPolicyStatusRequest, SandboxPolicy as ProtoSandboxPolicy,
SubmitPolicyAnalysisRequest, SubmitPolicyAnalysisResponse, UpdateConfigRequest,
open_shell_client::OpenShellClient, workspace_selector,
};
use crate::sandbox_env;
use crate::time::{duration_to_std, timestamp_to_millis};
@@ -397,6 +398,26 @@ pub async fn connect_channel_pub(endpoint: &str) -> Result<AuthedChannel> {
connect_channel(endpoint).await
}
/// Report installed provider state for the current authenticated supervisor session.
///
/// The observation must carry the session ID returned by `ConnectSupervisor`.
/// Reconnects start a new report sequence; retries preserve the complete report.
pub async fn report_provider_readiness(
endpoint: &str,
sandbox_id: &str,
observation: crate::proto::ProviderReadinessObservation,
) -> Result<crate::proto::ReportProviderReadinessResponse> {
let mut client = connect(endpoint).await?;
client
.report_provider_readiness(crate::proto::ReportProviderReadinessRequest {
sandbox_id: sandbox_id.to_string(),
observation: Some(observation),
})
.await
.map(tonic::Response::into_inner)
.into_diagnostic()
}
/// Background task that renews the sandbox JWT at ~80% of its remaining
/// lifetime. The new token replaces the value in [`TOKEN_SLOT`], so all
/// in-flight and future clients pick it up on their next request. The
@@ -1003,9 +1024,9 @@ pub async fn sync_policy_and_fetch_snapshot(
/// Fetch provider environment variables for a sandbox from `OpenShell` server via gRPC.
///
/// Returns a map of environment variable names to values derived from provider
/// credentials configured on the sandbox. Returns an empty map if the sandbox
/// has no providers or the call fails.
/// Returns the credential snapshot and its exact readiness identity. An empty
/// environment represents a sandbox without provider credentials. Transport
/// failure returns an error so callers can revoke credentials and retry.
pub async fn fetch_provider_environment(
endpoint: &str,
sandbox_id: &str,
@@ -1022,7 +1043,14 @@ pub async fn fetch_provider_environment(
.await
.into_diagnostic()?;
let inner = response.into_inner();
provider_environment_result(response.into_inner())
}
/// Preserve snapshot authority and reject invalid credential expiration times.
/// Unknown delivery reasons withhold credentials rather than implying readiness.
fn provider_environment_result(
inner: GetSandboxProviderEnvironmentResponse,
) -> Result<ProviderEnvironmentResult> {
let credential_expires_at_ms = inner
.credential_expiration_times
.iter()
@@ -1035,6 +1063,10 @@ pub async fn fetch_provider_environment(
Ok(ProviderEnvironmentResult {
environment: inner.environment,
provider_env_revision: inner.provider_env_revision,
provider_attachment_epoch: inner.provider_attachment_epoch,
policy_hash: inner.policy_hash,
readiness_reason: crate::proto::ProviderReadinessReason::try_from(inner.readiness_reason)
.unwrap_or(crate::proto::ProviderReadinessReason::CredentialsWithheld),
credential_expires_at_ms,
dynamic_credentials: inner.dynamic_credentials,
static_credential_bindings: inner.static_credential_bindings,
@@ -1042,6 +1074,75 @@ pub async fn fetch_provider_environment(
})
}
#[cfg(test)]
mod provider_environment_tests {
use super::*;
#[test]
fn provider_environment_preserves_readiness_identity() {
let result = provider_environment_result(GetSandboxProviderEnvironmentResponse {
environment: HashMap::from([("TOKEN".to_string(), "synthetic".to_string())]),
provider_env_revision: 42,
provider_attachment_epoch: "attachment-epoch".to_string(),
policy_hash: "binding-policy".to_string(),
readiness_reason: crate::proto::ProviderReadinessReason::CredentialsWithheld.into(),
credential_expiration_times: HashMap::from([(
"TOKEN".to_string(),
prost_types::Timestamp {
seconds: 1_900_000_000,
nanos: 123_000_000,
},
)]),
..Default::default()
})
.expect("valid provider environment");
assert_eq!(result.provider_env_revision, 42);
assert_eq!(result.provider_attachment_epoch, "attachment-epoch");
assert_eq!(result.policy_hash, "binding-policy");
assert_eq!(
result.readiness_reason,
crate::proto::ProviderReadinessReason::CredentialsWithheld
);
assert_eq!(
result.environment.get("TOKEN").map(String::as_str),
Some("synthetic")
);
assert_eq!(
result.credential_expires_at_ms.get("TOKEN"),
Some(&1_900_000_000_123)
);
}
#[test]
fn provider_readiness_unknown_delivery_reason_is_withheld() {
let result = provider_environment_result(GetSandboxProviderEnvironmentResponse {
policy_hash: "binding-policy".to_string(),
readiness_reason: i32::MAX,
..Default::default()
})
.expect("valid provider environment");
assert_eq!(
result.readiness_reason,
crate::proto::ProviderReadinessReason::CredentialsWithheld
);
}
#[test]
fn provider_environment_rejects_invalid_credential_expiration() {
let result = provider_environment_result(GetSandboxProviderEnvironmentResponse {
credential_expiration_times: HashMap::from([(
"TOKEN".to_string(),
prost_types::Timestamp {
seconds: 1_900_000_000,
nanos: -1,
},
)]),
..Default::default()
});
assert!(result.is_err());
}
}
pub async fn exchange_provider_subject_token(
endpoint: &str,
sandbox_id: &str,
@@ -1127,6 +1228,8 @@ pub struct SettingsPollResult {
/// When `policy_source` is `Global`, the version of the global policy revision.
pub global_policy_version: u32,
pub provider_env_revision: u64,
/// Attachment identity captured with this effective configuration.
pub provider_attachment_epoch: String,
pub supervisor_middleware_services: Vec<crate::proto::SupervisorMiddlewareService>,
/// Workspace the sandbox belongs to.
pub workspace: String,
@@ -1147,6 +1250,7 @@ fn settings_poll_result(inner: crate::proto::GetSandboxConfigResponse) -> Settin
settings: inner.settings,
global_policy_version: inner.global_policy_version,
provider_env_revision: inner.provider_env_revision,
provider_attachment_epoch: inner.provider_attachment_epoch,
supervisor_middleware_services: inner.supervisor_middleware_services,
workspace: inner.workspace,
policy_validation_failure_mode: inner
@@ -1200,9 +1304,16 @@ mod settings_poll_tests {
}
}
/// Credential material and the authority snapshot that produced its bindings.
pub struct ProviderEnvironmentResult {
pub environment: HashMap<String, String>,
pub provider_env_revision: u64,
/// Attachment identity captured with the delivered credential records.
pub provider_attachment_epoch: String,
/// Effective policy used to derive the delivered endpoint bindings.
pub policy_hash: String,
/// Closed failure category; withheld material cannot establish readiness.
pub readiness_reason: crate::proto::ProviderReadinessReason,
pub credential_expires_at_ms: HashMap<String, i64>,
pub dynamic_credentials: HashMap<String, crate::proto::ProviderProfileCredential>,
pub static_credential_bindings: HashMap<String, crate::proto::StaticCredentialBinding>,
@@ -17,11 +17,27 @@ const MAX_RETAINED_CREDENTIAL_GENERATIONS: usize = 8;
#[derive(Debug, Clone, Default)]
pub struct ProviderCredentialSnapshot {
/// Identifies this local installation, including repairs at the same revision.
pub installation_id: String,
pub revision: u64,
pub child_env: HashMap<String, String>,
pub dynamic_credentials: HashMap<String, crate::proto::ProviderProfileCredential>,
}
/// One atomic workload-facing snapshot and its local installation identity.
///
/// Only the prepared environment crosses an isolation boundary; resolver
/// material remains in the supervisor. The identity distinguishes an empty
/// fail-closed snapshot from a later repair with the same provider revision.
pub struct ChildEnvironmentSnapshot {
/// Opaque identity of the installed supervisor snapshot.
pub installation_id: String,
/// Opaque provider content fingerprint, never an ordered counter.
pub revision: u64,
/// Prepared environment used by future workload processes.
pub environment: HashMap<String, String>,
}
#[derive(Debug)]
struct ProviderCredentialStateInner {
current: Arc<ProviderCredentialSnapshot>,
@@ -89,6 +105,7 @@ impl ProviderCredentialState {
revision,
);
let snapshot = Arc::new(ProviderCredentialSnapshot {
installation_id: uuid::Uuid::new_v4().to_string(),
revision,
child_env,
dynamic_credentials,
@@ -136,6 +153,7 @@ impl ProviderCredentialState {
&stable_handles,
);
let snapshot = Arc::new(ProviderCredentialSnapshot {
installation_id: uuid::Uuid::new_v4().to_string(),
revision,
child_env,
dynamic_credentials,
@@ -178,6 +196,7 @@ impl ProviderCredentialState {
/// or holding the gateway-side resolver material.
pub fn from_child_env_snapshot(revision: u64, child_env: HashMap<String, String>) -> Self {
let snapshot = Arc::new(ProviderCredentialSnapshot {
installation_id: uuid::Uuid::new_v4().to_string(),
revision,
child_env,
dynamic_credentials: HashMap::new(),
@@ -221,6 +240,7 @@ impl ProviderCredentialState {
}
inner.current = Arc::new(ProviderCredentialSnapshot {
installation_id: uuid::Uuid::new_v4().to_string(),
revision,
child_env,
dynamic_credentials: HashMap::new(),
@@ -386,6 +406,7 @@ impl ProviderCredentialState {
inner.suppressed_keys.insert(key.to_string());
let mut env = (*inner.current).clone();
env.child_env.remove(key);
env.installation_id = uuid::Uuid::new_v4().to_string();
inner.current = Arc::new(env);
}
@@ -431,6 +452,20 @@ impl ProviderCredentialState {
Ok(Self::resolve_child_env_snapshot(&inner))
}
/// Capture installation identity and prepared environment under one lock.
pub fn child_environment_snapshot(&self) -> std::io::Result<ChildEnvironmentSnapshot> {
let inner = self
.inner
.read()
.map_err(|_| std::io::Error::other("provider credential state poisoned"))?;
let (revision, environment) = Self::resolve_child_env_snapshot(&inner);
Ok(ChildEnvironmentSnapshot {
installation_id: inner.current.installation_id.clone(),
revision,
environment,
})
}
fn resolve_child_env_snapshot(
inner: &ProviderCredentialStateInner,
) -> (u64, HashMap<String, String>) {
@@ -491,8 +526,9 @@ impl ProviderCredentialState {
/// Compare and install a workload-facing environment snapshot.
///
/// Provider environment revisions are opaque content identities, not
/// ordered counters. The expected revision makes retries idempotent while
/// rejecting updates based on a stale view of the boundary state.
/// ordered counters. The expected revision rejects a different current
/// fingerprint. Callers must also fence publication order when a failed
/// refresh and its repair can share that fingerprint.
pub fn compare_and_install_child_env_snapshot(
&self,
expected_revision: u64,
@@ -503,14 +539,19 @@ impl ProviderCredentialState {
.inner
.write()
.map_err(|_| std::io::Error::other("provider credential state poisoned"))?;
if revision == inner.current.revision || expected_revision != inner.current.revision {
if expected_revision != inner.current.revision {
return Ok(inner.current.revision);
}
// A failed refresh can clear the map without changing its provider
// fingerprint. A matching expectation must therefore install even an
// equal revision. Boundary publication ordering fences delayed retries.
for key in &inner.suppressed_keys {
child_env.remove(key);
}
inner.current = Arc::new(ProviderCredentialSnapshot {
installation_id: uuid::Uuid::new_v4().to_string(),
revision,
child_env,
dynamic_credentials: HashMap::new(),
@@ -587,6 +628,7 @@ impl ProviderCredentialState {
}
inner.current = Arc::new(ProviderCredentialSnapshot {
installation_id: uuid::Uuid::new_v4().to_string(),
revision,
child_env,
dynamic_credentials,
@@ -652,6 +694,7 @@ impl ProviderCredentialState {
child_env.remove(key);
}
inner.current = Arc::new(ProviderCredentialSnapshot {
installation_id: uuid::Uuid::new_v4().to_string(),
revision,
child_env,
dynamic_credentials,
@@ -711,6 +754,7 @@ impl ProviderCredentialState {
let dynamic_credentials =
dynamic_credentials.unwrap_or_else(|| inner.current.dynamic_credentials.clone());
inner.current = Arc::new(ProviderCredentialSnapshot {
installation_id: uuid::Uuid::new_v4().to_string(),
revision,
child_env: HashMap::new(),
dynamic_credentials,
@@ -2427,6 +2471,26 @@ mod tests {
assert!(env.is_empty(), "an empty snapshot must revoke the old env");
}
#[test]
fn child_environment_repair_replaces_an_empty_map_at_the_same_revision() {
let state = ProviderCredentialState::from_child_env_snapshot(6, HashMap::new());
let failed = state.child_environment_snapshot().unwrap();
state
.compare_and_install_child_env_snapshot(
6,
6,
HashMap::from([("TOKEN".to_string(), "reference".to_string())]),
)
.unwrap();
let repaired = state.child_environment_snapshot().unwrap();
assert_ne!(failed.installation_id, repaired.installation_id);
assert_eq!(repaired.revision, 6);
assert_eq!(
repaired.environment.get("TOKEN").map(String::as_str),
Some("reference")
);
}
#[test]
fn poisoned_environment_update_returns_error_without_recovering_state() {
let state = ProviderCredentialState::from_child_env_snapshot(4, HashMap::new());
@@ -762,6 +762,30 @@ pub struct ExecSpec {
pub trait BoundaryExec: Send + Sync {
/// Spawn `spec` inside the boundary, returning an owned session.
async fn exec(&self, spec: ExecSpec) -> Result<ExecSession, BackendError>;
/// Install the current provider environment for future process launches.
///
/// Success requires an authenticated acknowledgment from the running
/// boundary. Implementations serialize this operation with exec so an
/// older publication cannot replace the acknowledged environment.
async fn synchronize_provider_environment(
&self,
) -> Result<ProviderEnvironmentInstallation, BackendError> {
Err(BackendError::Unsupported(
"provider environment installation acknowledgment is unavailable".to_string(),
))
}
}
/// Evidence that the running workload boundary installed one provider snapshot.
#[derive(Clone, Debug, PartialEq, Eq)]
pub struct ProviderEnvironmentInstallation {
/// Local supervisor snapshot that produced the acknowledged environment.
pub installation_id: String,
/// Opaque provider content fingerprint.
pub revision: u64,
/// Authenticated and confirmed workload boundary session.
pub session_id: SandboxSessionId,
}
// ============================================================================
@@ -536,7 +536,8 @@ pub enum Request {
provider_env: std::collections::HashMap<String, String>,
},
UpdateProviderEnvironment {
expected_revision: u64,
/// Ordered publication within this authenticated boundary session.
generation: u64,
revision: u64,
provider_env: std::collections::HashMap<String, String>,
},
@@ -634,12 +635,12 @@ impl fmt::Debug for Request {
)
.finish(),
Self::UpdateProviderEnvironment {
expected_revision,
generation,
revision,
provider_env,
} => formatter
.debug_struct("UpdateProviderEnvironment")
.field("expected_revision", expected_revision)
.field("generation", generation)
.field("revision", revision)
.field(
"provider_env_keys",
@@ -710,9 +711,13 @@ pub enum Response {
Started {
process_id: String,
provider_env_revision: u64,
provider_env_generation: u64,
},
ProviderEnvironmentUpdated {
revision: u64,
generation: u64,
/// True only for this installed request or its exact idempotent replay.
applied: bool,
},
ProcessAttached {
terminal: bool,
@@ -1298,7 +1303,7 @@ mod tests {
second.insert("A".to_string(), "1".to_string());
second.insert("B".to_string(), "2".to_string());
let build = |provider_env| Request::UpdateProviderEnvironment {
expected_revision: 1,
generation: 1,
revision: 2,
provider_env,
};
+174 -20
View File
@@ -25,8 +25,8 @@ use openshell_isolation_interface::contract::{
BoundaryInput, BoundaryLoopbackConnector, BoundaryOutput, BoundaryProcess, BoundarySignal,
BoundaryTerminal, ConfirmedBoundary, ExecSession, ExecSpec, IsolationBackend, LoopbackTarget,
MediationTiming, NetworkMediationSource, PendingDnsQuery, PendingTcpOpen, ProcessAttachment,
ReadyBoundary, RunningBoundary, SandboxContext, TcpOpenDecision, TcpOpenDenial,
VerifiedBackendDescriptor,
ProviderEnvironmentInstallation, ReadyBoundary, RunningBoundary, SandboxContext,
TcpOpenDecision, TcpOpenDenial, VerifiedBackendDescriptor,
};
use sha2::{Digest as _, Sha256};
use tokio::io::{AsyncReadExt, AsyncWriteExt};
@@ -373,7 +373,8 @@ impl ReadyBoundary for RemoteReady {
.await?;
let Response::Started {
process_id,
provider_env_revision,
provider_env_generation,
..
} = response
else {
return Err(unexpected_response("started", &response));
@@ -387,7 +388,7 @@ impl ReadyBoundary for RemoteReady {
exec: Arc::new(RemoteExec {
client: self.client.clone(),
provider_credentials: self.provider_credentials,
boundary_revision: tokio::sync::Mutex::new(provider_env_revision),
publication_generation: tokio::sync::Mutex::new(provider_env_generation),
}),
loopback_connector: Arc::new(RemoteLoopbackConnector {
client: self.client,
@@ -552,30 +553,39 @@ async fn pump_process_responses(
struct RemoteExec {
client: Arc<BoundaryClient>,
provider_credentials: openshell_core::provider_credentials::ProviderCredentialState,
boundary_revision: tokio::sync::Mutex<u64>,
// Shared by proactive synchronization and exec. Reserve generations before
// dispatch so a timed-out request cannot overwrite a newer publication.
publication_generation: tokio::sync::Mutex<u64>,
}
#[async_trait]
impl BoundaryExec for RemoteExec {
async fn exec(&self, spec: ExecSpec) -> Result<ExecSession, BackendError> {
let mut boundary_revision = self.boundary_revision.lock().await;
impl RemoteExec {
async fn synchronize(
&self,
generation: &mut u64,
) -> Result<ProviderEnvironmentInstallation, BackendError> {
for _ in 0..3 {
let (revision, provider_env) = self
let snapshot = self
.provider_credentials
.child_env_snapshot_with_gcp_resolved()
.child_environment_snapshot()
.map_err(|error| {
BackendError::Process(format!("snapshot provider environment: {error}"))
})?;
*generation = generation.checked_add(1).ok_or_else(|| {
BackendError::Process("provider environment publication exhausted".to_string())
})?;
let requested_generation = *generation;
let response = self
.client
.call_idempotent(Request::UpdateProviderEnvironment {
expected_revision: *boundary_revision,
revision,
provider_env,
generation: requested_generation,
revision: snapshot.revision,
provider_env: snapshot.environment,
})
.await?;
let Response::ProviderEnvironmentUpdated {
revision: effective_revision,
generation: effective_generation,
applied,
} = response
else {
return Err(unexpected_response(
@@ -583,9 +593,18 @@ impl BoundaryExec for RemoteExec {
&response,
));
};
*boundary_revision = effective_revision;
if effective_revision == revision {
return open_exec_session(self.client.clone(), spec).await;
*generation = (*generation).max(effective_generation);
if applied
&& effective_generation == requested_generation
&& effective_revision == snapshot.revision
{
// The acknowledgment is for the exact request, including its
// map, even when a repaired map reuses a provider fingerprint.
return Ok(ProviderEnvironmentInstallation {
installation_id: snapshot.installation_id,
revision: snapshot.revision,
session_id: self.client.runtime_descriptor.session_id,
});
}
}
Err(BackendError::Process(
@@ -594,6 +613,22 @@ impl BoundaryExec for RemoteExec {
}
}
#[async_trait]
impl BoundaryExec for RemoteExec {
async fn exec(&self, spec: ExecSpec) -> Result<ExecSession, BackendError> {
let mut generation = self.publication_generation.lock().await;
self.synchronize(&mut generation).await?;
open_exec_session(self.client.clone(), spec).await
}
async fn synchronize_provider_environment(
&self,
) -> Result<ProviderEnvironmentInstallation, BackendError> {
let mut generation = self.publication_generation.lock().await;
self.synchronize(&mut generation).await
}
}
struct RemoteLoopbackConnector {
client: Arc<BoundaryClient>,
}
@@ -1872,6 +1907,7 @@ mod tests {
requests: Arc<std::sync::atomic::AtomicUsize>,
mediation_failures: Arc<std::sync::atomic::AtomicUsize>,
mediation_ready: bool,
provider_environment_generation: u64,
}
type TestGrpcStream = Pin<
@@ -1900,6 +1936,7 @@ mod tests {
let wait_for_half_close = self.wait_for_half_close;
let requests = self.requests.clone();
let mediation_ready = self.mediation_ready;
let provider_environment_generation = self.provider_environment_generation;
let (outbound, outbound_rx) = tokio::sync::mpsc::channel(1);
tokio::spawn(async move {
let mut frame = Vec::new();
@@ -1961,9 +1998,15 @@ mod tests {
}
Request::Terminate { .. } => Response::Terminated,
Request::TerminateBoundary => Response::BoundaryTerminated,
Request::UpdateProviderEnvironment { revision, .. } => {
Response::ProviderEnvironmentUpdated { revision }
}
Request::UpdateProviderEnvironment {
revision,
generation,
..
} => Response::ProviderEnvironmentUpdated {
revision,
generation: generation.max(provider_environment_generation),
applied: generation > provider_environment_generation,
},
Request::Resize { .. } => Response::Resized,
Request::LoopbackConnect { .. } => Response::PortConnected,
Request::StartAgent {
@@ -1972,6 +2015,7 @@ mod tests {
} => Response::Started {
process_id: "test-generation:main:0".to_string(),
provider_env_revision,
provider_env_generation: provider_environment_generation,
},
Request::AcceptNetwork => Response::Error {
kind: crate::boundary_protocol::BoundaryErrorKind::Unavailable,
@@ -2036,6 +2080,7 @@ mod tests {
requests: requests.clone(),
mediation_failures: Arc::new(std::sync::atomic::AtomicUsize::new(0)),
mediation_ready: false,
provider_environment_generation: 0,
};
let server = tokio::spawn(async move {
let (stream, _) = listener.accept().await.unwrap();
@@ -2099,6 +2144,7 @@ mod tests {
requests: server_requests.clone(),
mediation_failures: Arc::new(std::sync::atomic::AtomicUsize::new(0)),
mediation_ready: false,
provider_environment_generation: 0,
};
tokio::spawn(async move {
tonic::transport::Server::builder()
@@ -2173,6 +2219,7 @@ mod tests {
requests: server_requests.clone(),
mediation_failures: server_failures.clone(),
mediation_ready: true,
provider_environment_generation: 0,
};
tokio::spawn(async move {
tonic::transport::Server::builder()
@@ -2226,6 +2273,7 @@ mod tests {
requests,
mediation_failures: Arc::new(std::sync::atomic::AtomicUsize::new(0)),
mediation_ready: false,
provider_environment_generation: 0,
};
let server = tokio::spawn(async move {
let (stream, _) = listener.accept().await.unwrap();
@@ -2398,6 +2446,7 @@ mod tests {
requests: Arc::new(std::sync::atomic::AtomicUsize::new(0)),
mediation_failures: Arc::new(std::sync::atomic::AtomicUsize::new(0)),
mediation_ready: false,
provider_environment_generation: 0,
};
tonic::transport::Server::builder()
.add_service(IsolationBoundaryServer::new(service))
@@ -2717,6 +2766,7 @@ mod tests {
requests: handled.clone(),
mediation_failures: Arc::new(std::sync::atomic::AtomicUsize::new(0)),
mediation_ready: false,
provider_environment_generation: 0,
};
let server = tokio::spawn(async move {
loop {
@@ -2775,6 +2825,110 @@ mod tests {
server.abort();
}
#[tokio::test]
async fn reconstructed_backend_resumes_the_running_boundary_publication_generation() {
let certificate = test_certificate();
let listener = tokio::net::TcpListener::bind("127.0.0.1:0").await.unwrap();
let address = listener.local_addr().unwrap();
let requests = Arc::new(std::sync::atomic::AtomicUsize::new(0));
let service = TestGrpcBoundary {
wait_for_half_close: false,
expected_token: "a".repeat(32),
requests: requests.clone(),
mediation_failures: Arc::new(std::sync::atomic::AtomicUsize::new(0)),
mediation_ready: false,
provider_environment_generation: 50,
};
let server = tokio::spawn(async move {
let (stream, _) = listener.accept().await.unwrap();
let stream = tokio_rustls::TlsAcceptor::from(certificate.server_config)
.accept(stream)
.await
.unwrap();
tonic::transport::Server::builder()
.add_service(IsolationBoundaryServer::new(service))
.serve_with_incoming(tokio_stream::iter([Ok::<_, std::io::Error>(TestTlsIo(
Box::new(stream),
))]))
.await
.unwrap();
});
let credentials =
openshell_core::provider_credentials::ProviderCredentialState::from_child_env_snapshot(
6,
HashMap::new(),
);
let context = sandbox();
let ready = Box::new(RemoteReady {
client: Arc::new(BoundaryClient::new(
tls_runtime_descriptor(address, certificate.client_tls),
test_bearer(&"a".repeat(32)),
)),
agent: context.agent,
policy: context.policy,
sandbox_id: context.sandbox_id,
ca_file_paths: Arc::new(std::sync::Mutex::new(None)),
provider_credentials: credentials.clone(),
});
let running = ready.start_agent().await.unwrap();
assert_eq!(requests.load(Ordering::Acquire), 1);
let installed = running
.exec()
.synchronize_provider_environment()
.await
.unwrap();
assert_eq!(
requests.load(Ordering::Acquire),
2,
"the first update must advance the returned generation, without rejected catch-up calls"
);
assert_eq!(
installed.installation_id,
credentials.snapshot().installation_id
);
assert_eq!(installed.session_id, test_session_id());
server.abort();
let _ = server.await;
}
#[tokio::test]
async fn provider_environment_acknowledges_the_exact_local_snapshot_and_checked_generation() {
let certificate = test_certificate();
let (address, server) = spawn_tls_boundary(certificate.server_config, "a".repeat(32)).await;
let credentials =
openshell_core::provider_credentials::ProviderCredentialState::from_child_env_snapshot(
6,
HashMap::new(),
);
let client = Arc::new(BoundaryClient::new(
tls_runtime_descriptor(address, certificate.client_tls),
test_bearer(&"a".repeat(32)),
));
let exec = RemoteExec {
client,
provider_credentials: credentials.clone(),
publication_generation: tokio::sync::Mutex::new(0),
};
let empty = exec.synchronize_provider_environment().await.unwrap();
credentials
.install_child_env_snapshot(6, HashMap::from([("TOKEN".into(), "restored".into())]));
let repaired = exec.synchronize_provider_environment().await.unwrap();
assert_eq!(empty.revision, repaired.revision);
assert_ne!(empty.installation_id, repaired.installation_id);
assert_eq!(
repaired.installation_id,
credentials.snapshot().installation_id
);
assert_eq!(*exec.publication_generation.lock().await, 2);
*exec.publication_generation.lock().await = u64::MAX;
assert!(
exec.synchronize_provider_environment().await.is_err(),
"publication exhaustion must fail before sending another request"
);
server.abort();
let _ = server.await;
}
#[tokio::test]
async fn tls_tcp_flushes_large_control_requests_before_reading_response() {
let certificate = test_certificate();
@@ -65,6 +65,7 @@ async fn renewing_boundary_client(
requests: Arc::new(std::sync::atomic::AtomicUsize::new(0)),
mediation_failures: Arc::new(std::sync::atomic::AtomicUsize::new(0)),
mediation_ready: false,
provider_environment_generation: 0,
},
expected_token: Arc::new(std::sync::RwLock::new("a".repeat(32))),
failures: Arc::new(std::sync::atomic::AtomicUsize::new(0)),
+107 -16
View File
@@ -1863,10 +1863,10 @@ mod linux {
provider_env,
),
Request::UpdateProviderEnvironment {
expected_revision,
generation,
revision,
provider_env,
} => self.update_provider_environment(expected_revision, revision, provider_env),
} => self.update_provider_environment(generation, revision, provider_env),
Request::Wait { process_id } => self.wait(&process_id),
Request::Signal { process_id, signal } => self.signal(&process_id, signal),
Request::Terminate { process_id } => self.terminate(&process_id),
@@ -2336,6 +2336,7 @@ mod linux {
Response::Started {
process_id: process.process_id(),
provider_env_revision: process.provider_credentials.snapshot().revision,
provider_env_generation: *lock(&process.provider_environment_generation),
}
} else {
guest_error(
@@ -2400,12 +2401,13 @@ mod linux {
Response::Started {
process_id,
provider_env_revision,
provider_env_generation: 0,
}
}
fn update_provider_environment(
&self,
expected_revision: u64,
generation: u64,
revision: u64,
provider_env: std::collections::HashMap<String, String>,
) -> Response {
@@ -2419,14 +2421,34 @@ mod linux {
};
process.clone()
};
// The session-scoped publication order is separate from opaque
// provider fingerprints. Holding it through installation prevents
// a delayed request from replacing a newer same-revision repair.
let mut installed_generation = lock(&process.provider_environment_generation);
let current = match process.provider_credentials.child_environment_snapshot() {
Ok(snapshot) => snapshot,
Err(error) => return guest_error(BoundaryErrorKind::Process, error.to_string()),
};
if generation <= *installed_generation {
return Response::ProviderEnvironmentUpdated {
revision: current.revision,
generation: *installed_generation,
applied: false,
};
}
let revision = match process
.provider_credentials
.compare_and_install_child_env_snapshot(expected_revision, revision, provider_env)
.compare_and_install_child_env_snapshot(current.revision, revision, provider_env)
{
Ok(revision) => revision,
Err(error) => return guest_error(BoundaryErrorKind::Process, error.to_string()),
};
Response::ProviderEnvironmentUpdated { revision }
*installed_generation = generation;
Response::ProviderEnvironmentUpdated {
revision,
generation,
applied: true,
}
}
fn wait(&self, process_id: &str) -> Response {
@@ -2699,6 +2721,7 @@ mod linux {
attached: Arc<AtomicBool>,
boundary_runtime: Arc<BoundaryRuntimeState>,
provider_credentials: ProviderCredentialState,
provider_environment_generation: Mutex<u64>,
}
struct ManagedProcessLaunch {
@@ -2792,6 +2815,7 @@ mod linux {
attached: Arc::new(AtomicBool::new(false)),
boundary_runtime,
provider_credentials,
provider_environment_generation: Mutex::new(0),
})
}
@@ -4598,13 +4622,14 @@ mod linux {
let Response::Started {
process_id,
provider_env_revision: 0,
provider_env_generation: 0,
} = start()
else {
panic!("initial start did not succeed");
};
let update = RequestEnvelope::new(Request::UpdateProviderEnvironment {
expected_revision: 0,
generation: 1,
revision: 7,
provider_env: std::collections::HashMap::from([(
"REPLAY_TEST".to_string(),
@@ -4614,11 +4639,19 @@ mod linux {
.expect("build replayed update");
assert_eq!(
boundary.dispatch(update.clone()),
Response::ProviderEnvironmentUpdated { revision: 7 }
Response::ProviderEnvironmentUpdated {
revision: 7,
generation: 1,
applied: true,
}
);
assert_eq!(
boundary.dispatch(update.clone()),
Response::ProviderEnvironmentUpdated { revision: 7 },
Response::ProviderEnvironmentUpdated {
revision: 7,
generation: 1,
applied: true,
},
"the same request ID and payload must replay its recorded response"
);
let mut changed = RequestEnvelope::new(Request::Terminate {
@@ -4662,6 +4695,7 @@ mod linux {
Response::Started {
process_id: process_id.clone(),
provider_env_revision: 7,
provider_env_generation: 1,
}
);
@@ -4684,9 +4718,44 @@ mod linux {
Response::Error { kind, .. } if kind == BoundaryErrorKind::Denied
));
// Failed refresh and recovery can retain the same provider
// fingerprint. Distinct publications must still replace the map,
// while a delayed older clear must never undo the repair.
assert_eq!(
boundary.update_provider_environment(2, 7, std::collections::HashMap::default()),
Response::ProviderEnvironmentUpdated {
revision: 7,
generation: 2,
applied: true
}
);
assert_eq!(
boundary.update_provider_environment(
3,
7,
std::collections::HashMap::from([(
"REPLAY_TEST".to_string(),
"reconnected".to_string()
),])
),
Response::ProviderEnvironmentUpdated {
revision: 7,
generation: 3,
applied: true
}
);
assert_eq!(
boundary.update_provider_environment(2, 7, std::collections::HashMap::default()),
Response::ProviderEnvironmentUpdated {
revision: 7,
generation: 3,
applied: false
}
);
let exec_spec = ExecSpecWire {
program: "/bin/sh".to_string(),
args: vec!["-c".to_string(), "printf reconnected".to_string()],
args: vec!["-c".to_string(), "printf '%s' \"$REPLAY_TEST\"".to_string()],
env: Vec::new(),
workdir: None,
pty: false,
@@ -4869,19 +4938,24 @@ mod linux {
Response::Started {
process_id: process.process_id(),
provider_env_revision: 0,
provider_env_generation: 0,
}
);
assert_eq!(
boundary.update_provider_environment(
0,
1,
2,
std::collections::HashMap::from([(
"ROTATED_TOKEN".to_string(),
"refreshed".to_string(),
)]),
),
Response::ProviderEnvironmentUpdated { revision: 2 }
Response::ProviderEnvironmentUpdated {
revision: 2,
generation: 1,
applied: true,
}
);
assert_eq!(
boundary.update_provider_environment(
@@ -4892,11 +4966,19 @@ mod linux {
"stale".to_string(),
)]),
),
Response::ProviderEnvironmentUpdated { revision: 2 }
Response::ProviderEnvironmentUpdated {
revision: 2,
generation: 1,
applied: false,
}
);
assert_eq!(
boundary.update_provider_environment(2, 1, std::collections::HashMap::new()),
Response::ProviderEnvironmentUpdated { revision: 1 },
Response::ProviderEnvironmentUpdated {
revision: 1,
generation: 2,
applied: true,
},
"a numerically smaller opaque revision must revoke the environment"
);
assert_eq!(
@@ -4908,12 +4990,20 @@ mod linux {
"out-of-order".to_string(),
)]),
),
Response::ProviderEnvironmentUpdated { revision: 1 },
"a stale expected revision must not overwrite current state"
Response::ProviderEnvironmentUpdated {
revision: 1,
generation: 2,
applied: false,
},
"a stale publication must not overwrite current state"
);
assert_eq!(
boundary.update_provider_environment(1, 1, std::collections::HashMap::new()),
Response::ProviderEnvironmentUpdated { revision: 1 },
Response::ProviderEnvironmentUpdated {
revision: 1,
generation: 2,
applied: false,
},
"a duplicate update must be idempotent"
);
@@ -4938,6 +5028,7 @@ mod linux {
Response::Started {
process_id: process.process_id(),
provider_env_revision: 1,
provider_env_generation: 2,
},
"a replacement control must resume from the boundary's current revision"
);
+35
View File
@@ -130,6 +130,40 @@ let _sandbox = client
# }
```
## Wait for a provider change
Provider attach, detach, and update responses include a `ProviderMutationReceipt`: a saved record identifying the exact change requested for one sandbox. Pass that record to `provider_readiness::wait_for_provider` to wait until the current sandbox runtime confirms it applied the change. Detach completes with `Revoked`; attach and update complete with `Ready`.
```rust
use std::time::Duration;
use openshell_sdk::{OpenShellClient, raw::ProviderMutationReceipt};
use openshell_sdk::provider_readiness::{
ProviderWaitOutcome, wait_for_provider,
};
async fn wait_for_change(
client: &OpenShellClient,
change: &ProviderMutationReceipt,
) -> Result<(), Box<dyn std::error::Error>> {
let mut grpc = client.raw_grpc_fresh().await?;
let result = wait_for_provider(&mut grpc, change, Duration::from_secs(30)).await?;
match result.outcome {
ProviderWaitOutcome::Complete => println!("The sandbox applied the change."),
ProviderWaitOutcome::TimedOut => println!("Still waiting; check the same change again."),
ProviderWaitOutcome::Terminal => println!("The change failed, was withheld, or was replaced."),
}
Ok(())
}
```
The result preserves the last known status when the deadline expires. A later change cannot satisfy a wait for the original request. `provider_status` queries once, and `wait_for_provider_until` accepts a shared deadline for waiting on the sandboxes selected by one provider update. Status responses contain configuration identities and safe reason categories, without credentials or raw installation errors.
For ordinary static credentials, launch a new client after update readiness to receive the updated reference. Existing processes keep their revision-scoped references; a successful wait does not retarget them or prove that the old upstream key can be retired. After detach completes, retained references cannot resolve and new processes do not receive them.
The status's `operation` field is the common operation's historical outcome, keyed by the receipt ID. These helpers complete from the live provider state and its matching evidence; a historical applied operation cannot override a disconnected, expired, or superseded live result.
These helpers use the raw client's authentication slot. They do not perform OIDC refresh themselves. Follow the raw-client refresh guidance above if a request returns `Unauthenticated`, then resume waiting for the same change ID.
## Modules
| Module | Purpose |
@@ -145,6 +179,7 @@ let _sandbox = client
| `pagination` | Lazy `Pager<T>` and response `Page<T>`. |
| `types` | Curated request/response types and proto conversions. |
| `raw` | Escape hatch re-exporting the generated tonic clients. |
| `provider_readiness` | Check and wait for an exact provider change to take effect. |
## Notes
+1
View File
@@ -37,6 +37,7 @@ pub mod edge_tunnel;
pub mod error;
pub mod oidc;
pub mod pagination;
pub mod provider_readiness;
pub mod raw;
pub mod refresh;
pub mod transport;
@@ -0,0 +1,841 @@
// SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
// SPDX-License-Identifier: Apache-2.0
//! Receipt-bound provider status and bounded waits over the raw provider API.
//!
//! A wait pins the original desired identity. It cannot succeed by following a
//! later mutation, and its deadline covers RPC execution as well as polling.
use crate::raw::GrpcClient;
use openshell_core::proto::{
GetSandboxProviderStatusRequest, ProviderMutationKind, ProviderMutationReceipt,
ProviderReadinessReason, ProviderReadinessState, ProviderReadinessStatus,
};
use std::future::Future;
use std::time::Duration;
use thiserror::Error;
use tokio::time::Instant;
use tonic::service::Interceptor;
use tonic::service::interceptor::InterceptedService;
use tonic::transport::Channel;
/// Maximum duration accepted by provider waits.
pub const MAX_PROVIDER_WAIT: Duration = Duration::from_hours(1);
/// Completion of a bounded wait, separate from the gateway's installed state.
#[derive(Clone, Copy, Debug, PartialEq, Eq)]
pub enum ProviderWaitOutcome {
/// The original receipt's authority was installed or revoked as requested.
Complete,
/// The deadline elapsed; the last observed state remains authoritative.
TimedOut,
/// Installation failed, was withheld, or the desired authority was superseded.
Terminal,
}
/// Last safe status and the reason a provider wait returned.
#[derive(Clone, Debug)]
pub struct ProviderWaitResult {
/// Status for the original receipt, including on timeout.
pub status: ProviderReadinessStatus,
/// Whether installation completed, failed, or exhausted the deadline.
pub outcome: ProviderWaitOutcome,
}
/// Errors safe to display without including raw secret-bearing RPC messages.
#[derive(Clone, Debug, Error)]
pub enum ProviderReadinessError {
/// Provider waits require a positive timeout of at most one hour.
#[error("provider wait timeout must be greater than zero and at most 3600 seconds")]
InvalidTimeout,
/// A receipt must identify a sandbox, workspace, provider, and desired state.
#[error("provider readiness receipt is incomplete")]
InvalidReceipt,
/// The gateway returned no usable status or an unknown protocol state.
#[error("gateway returned an invalid provider readiness status")]
InvalidStatus,
/// The RPC failed; only its protocol code is exposed.
#[error("provider readiness request failed ({code})")]
Rpc {
/// gRPC status code without the server's diagnostic message or metadata.
code: tonic::Code,
},
}
/// Query a provider receipt or reconstruct current state when its ID is empty.
/// A receipt ID selects its provider when the request omits the provider name.
///
/// # Errors
/// Returns a safe protocol error if the gateway is unavailable or omits status.
pub async fn provider_status<I>(
client: &mut GrpcClient<InterceptedService<Channel, I>>,
request: GetSandboxProviderStatusRequest,
) -> Result<ProviderReadinessStatus, ProviderReadinessError>
where
I: Interceptor + Clone + Send + Sync,
{
let expected = request.clone();
let status = client
.get_sandbox_provider_status(request)
.await
.map_err(|error| ProviderReadinessError::Rpc { code: error.code() })?
.into_inner()
.status
.ok_or(ProviderReadinessError::InvalidStatus)?;
validate_status_for_request(&status, &expected)?;
Ok(status)
}
fn validate_status_for_request(
status: &ProviderReadinessStatus,
expected: &GetSandboxProviderStatusRequest,
) -> Result<(), ProviderReadinessError> {
validate_status(status)?;
let receipt = status
.receipt
.as_ref()
.ok_or(ProviderReadinessError::InvalidStatus)?;
// Either a receipt ID or a provider name must select the authority. Every
// supplied selector must still bind the response to that authority.
if (expected.receipt_id.is_empty() && expected.provider_name.is_empty())
|| (!expected.receipt_id.is_empty() && receipt.receipt_id != expected.receipt_id)
|| (!expected.provider_name.is_empty() && receipt.provider_name != expected.provider_name)
|| expected.workspace_scope.as_ref().is_some_and(|scope| {
!matches!(
scope.selection.as_ref(),
Some(openshell_core::proto::workspace_selector::Selection::Workspace(workspace))
if workspace == &receipt.workspace
)
})
|| receipt
.desired
.as_ref()
.is_none_or(|desired| desired.sandbox_name != expected.sandbox_name)
{
return Err(ProviderReadinessError::InvalidStatus);
}
Ok(())
}
/// Wait for the exact authority named by a provider receipt.
///
/// Detach requires `REVOKED`; attach and update require `READY`. A reconstructed
/// observation waits for ready or revoked according to its original provider ID.
/// This raw helper uses the authentication slot on the supplied client.
///
/// # Errors
/// Rejects an invalid timeout/receipt, malformed responses, or failed RPCs.
pub async fn wait_for_provider<I>(
client: &mut GrpcClient<InterceptedService<Channel, I>>,
receipt: &ProviderMutationReceipt,
timeout: Duration,
) -> Result<ProviderWaitResult, ProviderReadinessError>
where
I: Interceptor + Clone + Send + Sync,
{
let deadline = provider_wait_deadline(timeout)?;
wait_for_provider_until(client, receipt, deadline).await
}
/// Wait using a shared deadline, so a multi-sandbox update has one total bound.
///
/// # Errors
/// Returns safe protocol errors without retaining the underlying RPC message.
pub async fn wait_for_provider_until<I>(
client: &mut GrpcClient<InterceptedService<Channel, I>>,
receipt: &ProviderMutationReceipt,
deadline: Instant,
) -> Result<ProviderWaitResult, ProviderReadinessError>
where
I: Interceptor + Clone + Send + Sync,
{
wait_with(receipt, deadline, Duration::from_millis(250), |request| {
let mut client = client.clone();
async move { provider_status(&mut client, request).await }
})
.await
}
/// Continue a bounded wait from a status that has already been observed.
///
/// Completed observations return immediately. A later failed or timed-out RPC
/// cannot discard the last observed pending state and replace it with saved intent.
///
/// # Errors
/// Rejects malformed receipts/statuses and reports only safe protocol errors.
pub async fn wait_for_provider_status_until<I>(
client: &mut GrpcClient<InterceptedService<Channel, I>>,
status: &ProviderReadinessStatus,
deadline: Instant,
) -> Result<ProviderWaitResult, ProviderReadinessError>
where
I: Interceptor + Clone + Send + Sync,
{
let receipt = status
.receipt
.as_ref()
.ok_or(ProviderReadinessError::InvalidStatus)?;
wait_with_initial_status(
receipt,
status.clone(),
deadline,
Duration::from_millis(250),
|request| {
let mut client = client.clone();
async move { provider_status(&mut client, request).await }
},
)
.await
}
/// Validate a timeout and create a deadline shared across all selected sandboxes.
///
/// # Errors
/// A zero timeout or a timeout over one hour is rejected before mutation.
pub fn provider_wait_deadline(timeout: Duration) -> Result<Instant, ProviderReadinessError> {
if timeout.is_zero() || timeout > MAX_PROVIDER_WAIT {
return Err(ProviderReadinessError::InvalidTimeout);
}
Ok(Instant::now() + timeout)
}
/// Represent saved intent before any runtime observation has been fetched.
#[must_use]
pub fn persisted_status(receipt: ProviderMutationReceipt) -> ProviderReadinessStatus {
ProviderReadinessStatus {
receipt: Some(receipt),
state: ProviderReadinessState::Persisted.into(),
reason: ProviderReadinessReason::WaitingForSupervisor.into(),
..Default::default()
}
}
fn validate_status(status: &ProviderReadinessStatus) -> Result<(), ProviderReadinessError> {
if status.receipt.is_none()
|| matches!(
ProviderReadinessState::try_from(status.state),
Err(_) | Ok(ProviderReadinessState::Unspecified)
)
|| ProviderReadinessReason::try_from(status.reason).is_err()
{
return Err(ProviderReadinessError::InvalidStatus);
}
// Absent observation times are expected before a report arrives. Present
// times must be canonical so output cannot normalize malformed timestamps.
for timestamp in [&status.observed_time, &status.evaluated_time]
.into_iter()
.flatten()
{
openshell_core::time::validate_timestamp(timestamp)
.map_err(|_| ProviderReadinessError::InvalidStatus)?;
}
let receipt = status
.receipt
.as_ref()
.ok_or(ProviderReadinessError::InvalidStatus)?;
validate_provider_receipt(receipt).map_err(|_| ProviderReadinessError::InvalidStatus)?;
let state = ProviderReadinessState::try_from(status.state)
.map_err(|_| ProviderReadinessError::InvalidStatus)?;
if matches!(
state,
ProviderReadinessState::Ready | ProviderReadinessState::Revoked
) {
if state != completion_state(receipt).map_err(|_| ProviderReadinessError::InvalidStatus)? {
return Err(ProviderReadinessError::InvalidStatus);
}
// A status read has the same installation contract as a wait. Enum
// labels alone must never report ready or revoked to a direct caller.
validate_completion(status)?;
}
Ok(())
}
fn completion_state(
receipt: &ProviderMutationReceipt,
) -> Result<ProviderReadinessState, ProviderReadinessError> {
let desired = receipt
.desired
.as_ref()
.ok_or(ProviderReadinessError::InvalidReceipt)?;
let attached = !desired.provider_id.is_empty();
match ProviderMutationKind::try_from(receipt.kind) {
Ok(ProviderMutationKind::Attach | ProviderMutationKind::Update) if attached => {
Ok(ProviderReadinessState::Ready)
}
Ok(ProviderMutationKind::Detach) if !attached => Ok(ProviderReadinessState::Revoked),
Ok(ProviderMutationKind::Observe) => Ok(if attached {
ProviderReadinessState::Ready
} else {
ProviderReadinessState::Revoked
}),
_ => Err(ProviderReadinessError::InvalidReceipt),
}
}
/// Validate saved receipt identity without fetching a runtime observation.
///
/// Initial attachment epochs may be empty and revision fingerprints may be zero;
/// the mutation kind must still agree with the presence of a provider identity.
///
/// # Errors
/// Rejects incomplete receipts, invalid persistence timestamps, and inconsistent
/// mutation/provider identities.
pub fn validate_provider_receipt(
receipt: &ProviderMutationReceipt,
) -> Result<(), ProviderReadinessError> {
let desired = receipt
.desired
.as_ref()
.ok_or(ProviderReadinessError::InvalidReceipt)?;
let persisted_time = receipt
.persisted_time
.as_ref()
.ok_or(ProviderReadinessError::InvalidReceipt)?;
openshell_core::time::validate_timestamp(persisted_time)
.map_err(|_| ProviderReadinessError::InvalidReceipt)?;
if receipt.receipt_id.is_empty()
|| receipt.mutation_id.is_empty()
|| desired.sandbox_id.is_empty()
|| desired.sandbox_name.is_empty()
|| receipt.provider_name.is_empty()
|| receipt.workspace.is_empty()
|| matches!(
ProviderMutationKind::try_from(receipt.kind),
Err(_) | Ok(ProviderMutationKind::Unspecified)
)
{
return Err(ProviderReadinessError::InvalidReceipt);
}
completion_state(receipt)?;
Ok(())
}
fn request_for_receipt(
receipt: &ProviderMutationReceipt,
) -> Result<GetSandboxProviderStatusRequest, ProviderReadinessError> {
validate_provider_receipt(receipt)?;
let desired = receipt
.desired
.as_ref()
.ok_or(ProviderReadinessError::InvalidReceipt)?;
Ok(GetSandboxProviderStatusRequest {
sandbox_name: desired.sandbox_name.clone(),
provider_name: receipt.provider_name.clone(),
receipt_id: receipt.receipt_id.clone(),
workspace_scope: Some(openshell_core::proto::workspace_selector(
&receipt.workspace,
)),
})
}
fn disposition(
receipt: &ProviderMutationReceipt,
status: &mut ProviderReadinessStatus,
) -> Result<Option<ProviderWaitOutcome>, ProviderReadinessError> {
validate_status(status)?;
let observed_receipt = status
.receipt
.as_ref()
.ok_or(ProviderReadinessError::InvalidStatus)?;
// Status reconstruction creates an immutable observation receipt too. The
// full receipt must remain fixed; a later snapshot cannot move this target.
if observed_receipt.desired != receipt.desired
|| observed_receipt.provider_name != receipt.provider_name
|| observed_receipt.workspace != receipt.workspace
|| (!receipt.receipt_id.is_empty()
&& (observed_receipt.receipt_id != receipt.receipt_id
|| observed_receipt.mutation_id != receipt.mutation_id
|| observed_receipt.kind != receipt.kind
|| observed_receipt.persisted_time != receipt.persisted_time))
{
status.receipt = Some(receipt.clone());
status.state = ProviderReadinessState::Superseded.into();
status.reason = ProviderReadinessReason::DesiredStateChanged.into();
return Ok(Some(ProviderWaitOutcome::Terminal));
}
let state = ProviderReadinessState::try_from(status.state)
.map_err(|_| ProviderReadinessError::InvalidStatus)?;
let revoked = completion_state(receipt)? == ProviderReadinessState::Revoked;
match state {
ProviderReadinessState::Ready if !revoked => {
validate_completion(status)?;
Ok(Some(ProviderWaitOutcome::Complete))
}
ProviderReadinessState::Revoked if revoked => {
validate_completion(status)?;
Ok(Some(ProviderWaitOutcome::Complete))
}
ProviderReadinessState::Withheld
| ProviderReadinessState::Failed
| ProviderReadinessState::Superseded => Ok(Some(ProviderWaitOutcome::Terminal)),
ProviderReadinessState::Ready | ProviderReadinessState::Revoked => {
Err(ProviderReadinessError::InvalidStatus)
}
_ => Ok(None),
}
}
fn validate_completion(status: &ProviderReadinessStatus) -> Result<(), ProviderReadinessError> {
let desired = status
.receipt
.as_ref()
.and_then(|receipt| receipt.desired.as_ref())
.ok_or(ProviderReadinessError::InvalidStatus)?;
let observed = status
.observed
.as_ref()
.ok_or(ProviderReadinessError::InvalidStatus)?;
if desired.policy_hash.is_empty()
|| status.network_instance_id.is_empty()
|| observed.session_id.is_empty()
|| observed.process_instance_id.is_empty()
|| !observed.credentials_installed
|| !observed.policy_active
|| !observed.launch_environment_installed
|| observed.reason != i32::from(ProviderReadinessReason::Unspecified)
|| status.reason != i32::from(ProviderReadinessReason::Unspecified)
|| observed.attachment_epoch != desired.attachment_epoch
|| observed.provider_env_revision != desired.provider_env_revision
|| observed.config_revision != desired.config_revision
|| observed.policy_hash != desired.policy_hash
{
return Err(ProviderReadinessError::InvalidStatus);
}
Ok(())
}
async fn wait_with<F, Fut>(
receipt: &ProviderMutationReceipt,
deadline: Instant,
interval: Duration,
fetch: F,
) -> Result<ProviderWaitResult, ProviderReadinessError>
where
F: FnMut(GetSandboxProviderStatusRequest) -> Fut,
Fut: Future<Output = Result<ProviderReadinessStatus, ProviderReadinessError>>,
{
wait_with_initial_status(
receipt,
persisted_status(receipt.clone()),
deadline,
interval,
fetch,
)
.await
}
async fn wait_with_initial_status<F, Fut>(
receipt: &ProviderMutationReceipt,
mut last: ProviderReadinessStatus,
deadline: Instant,
interval: Duration,
mut fetch: F,
) -> Result<ProviderWaitResult, ProviderReadinessError>
where
F: FnMut(GetSandboxProviderStatusRequest) -> Fut,
Fut: Future<Output = Result<ProviderReadinessStatus, ProviderReadinessError>>,
{
let request = request_for_receipt(receipt)?;
if let Some(outcome) = disposition(receipt, &mut last)? {
return Ok(ProviderWaitResult {
status: last,
outcome,
});
}
loop {
if Instant::now() >= deadline {
return Ok(ProviderWaitResult {
status: last,
outcome: ProviderWaitOutcome::TimedOut,
});
}
// The timeout also cancels an RPC that never responds; sleeping alone
// between polls cannot enforce a bounded wait against an unavailable peer.
match tokio::time::timeout_at(deadline, fetch(request.clone())).await {
Ok(result) => last = result?,
Err(_) => {
return Ok(ProviderWaitResult {
status: last,
outcome: ProviderWaitOutcome::TimedOut,
});
}
}
if let Some(outcome) = disposition(receipt, &mut last)? {
return Ok(ProviderWaitResult {
status: last,
outcome,
});
}
tokio::time::sleep_until((Instant::now() + interval).min(deadline)).await;
}
}
#[cfg(test)]
mod tests {
use super::*;
use openshell_core::proto::ProviderDesiredIdentity;
fn receipt(kind: ProviderMutationKind) -> ProviderMutationReceipt {
ProviderMutationReceipt {
receipt_id: "receipt".into(),
mutation_id: "mutation".into(),
provider_name: "provider".into(),
workspace: "default".into(),
persisted_time: Some(openshell_core::time::timestamp_from_millis(1).unwrap()),
kind: kind.into(),
desired: Some(ProviderDesiredIdentity {
sandbox_id: "sandbox-id".into(),
sandbox_name: "sandbox".into(),
provider_id: if kind == ProviderMutationKind::Detach {
String::new()
} else {
"provider-id".into()
},
attachment_epoch: "attachment-epoch".into(),
policy_hash: "policy-hash".into(),
provider_env_revision: u64::MAX,
..Default::default()
}),
}
}
fn completed_status(
receipt: &ProviderMutationReceipt,
state: ProviderReadinessState,
) -> ProviderReadinessStatus {
let desired = receipt.desired.as_ref().unwrap();
ProviderReadinessStatus {
receipt: Some(receipt.clone()),
state: state.into(),
network_instance_id: "network-instance".into(),
observed: Some(openshell_core::proto::ProviderReadinessObservation {
session_id: "session".into(),
process_instance_id: "process-instance".into(),
attachment_epoch: desired.attachment_epoch.clone(),
provider_env_revision: desired.provider_env_revision,
config_revision: desired.config_revision,
policy_hash: desired.policy_hash.clone(),
credentials_installed: true,
policy_active: true,
launch_environment_installed: true,
..Default::default()
}),
..Default::default()
}
}
#[test]
fn receipt_requires_a_present_canonical_persistence_time() {
let mut receipt = receipt(ProviderMutationKind::Attach);
for (seconds, nanos, valid) in [
(0, 1, true),
(0, 0, true),
(-1, 999_999_999, true),
(1, -1, false),
(1, 1_000_000_000, false),
(253_402_300_800, 0, false),
(-62_135_596_801, 0, false),
] {
let timestamp = receipt.persisted_time.as_mut().unwrap();
timestamp.seconds = seconds;
timestamp.nanos = nanos;
assert_eq!(
validate_provider_receipt(&receipt).is_ok(),
valid,
"seconds={seconds}, nanos={nanos}"
);
}
receipt.persisted_time = None;
assert!(matches!(
validate_provider_receipt(&receipt),
Err(ProviderReadinessError::InvalidReceipt)
));
}
#[test]
fn changed_persistence_nanoseconds_supersede_the_original_receipt() {
let receipt = receipt(ProviderMutationKind::Attach);
let mut changed = receipt.clone();
changed.persisted_time.as_mut().unwrap().nanos += 1;
let mut status = completed_status(&changed, ProviderReadinessState::Ready);
assert_eq!(
disposition(&receipt, &mut status).unwrap(),
Some(ProviderWaitOutcome::Terminal)
);
assert_eq!(status.state, i32::from(ProviderReadinessState::Superseded));
assert_eq!(status.receipt.as_ref(), Some(&receipt));
}
#[test]
fn status_rejects_malformed_observation_and_evaluation_times() {
let receipt = receipt(ProviderMutationKind::Attach);
let mut malformed = openshell_core::time::timestamp_from_millis(1).unwrap();
malformed.nanos = -1;
for observation in [true, false] {
let mut status = completed_status(&receipt, ProviderReadinessState::Ready);
if observation {
status.observed_time = Some(malformed);
} else {
status.evaluated_time = Some(malformed);
}
assert!(matches!(
validate_status(&status),
Err(ProviderReadinessError::InvalidStatus)
));
}
}
#[test]
fn direct_status_accepts_receipt_only_and_explicit_provider_selectors() {
let receipt = receipt(ProviderMutationKind::Attach);
let status = completed_status(&receipt, ProviderReadinessState::Ready);
let mut request = GetSandboxProviderStatusRequest {
sandbox_name: "sandbox".into(),
receipt_id: receipt.receipt_id.clone(),
workspace_scope: Some(openshell_core::proto::workspace_selector("default")),
..Default::default()
};
assert!(validate_status_for_request(&status, &request).is_ok());
request.provider_name = receipt.provider_name;
assert!(validate_status_for_request(&status, &request).is_ok());
request.receipt_id.clear();
assert!(validate_status_for_request(&status, &request).is_ok());
}
#[test]
fn direct_status_rejects_mismatched_explicit_selectors() {
let receipt = receipt(ProviderMutationKind::Attach);
let status = completed_status(&receipt, ProviderReadinessState::Ready);
let request = GetSandboxProviderStatusRequest {
sandbox_name: "sandbox".into(),
receipt_id: receipt.receipt_id,
workspace_scope: Some(openshell_core::proto::workspace_selector("default")),
..Default::default()
};
for invalid in [
GetSandboxProviderStatusRequest {
receipt_id: String::new(),
..request.clone()
},
GetSandboxProviderStatusRequest {
provider_name: "other-provider".into(),
..request.clone()
},
GetSandboxProviderStatusRequest {
receipt_id: "other-receipt".into(),
..request.clone()
},
GetSandboxProviderStatusRequest {
sandbox_name: "other-sandbox".into(),
..request.clone()
},
GetSandboxProviderStatusRequest {
workspace_scope: Some(openshell_core::proto::workspace_selector("other-workspace")),
..request
},
] {
assert!(matches!(
validate_status_for_request(&status, &invalid),
Err(ProviderReadinessError::InvalidStatus)
));
}
}
#[tokio::test]
async fn pending_process_waits_until_exact_receipt_is_ready() {
let receipt = receipt(ProviderMutationKind::Attach);
let mut calls = 0;
let result = wait_with(
&receipt,
Instant::now() + Duration::from_secs(1),
Duration::ZERO,
|_| {
calls += 1;
let mut status = completed_status(&receipt, ProviderReadinessState::Ready);
status.state = if calls == 1 {
ProviderReadinessState::Pending
} else {
ProviderReadinessState::Ready
}
.into();
status.reason = if calls == 1 {
ProviderReadinessReason::WaitingForProcess
} else {
ProviderReadinessReason::Unspecified
}
.into();
async { Ok(status) }
},
)
.await
.unwrap();
assert_eq!(calls, 2);
assert_eq!(result.outcome, ProviderWaitOutcome::Complete);
}
#[tokio::test]
async fn deadline_bounds_a_hung_status_rpc() {
let receipt = receipt(ProviderMutationKind::Update);
let result = wait_with(
&receipt,
Instant::now() + Duration::from_millis(10),
Duration::ZERO,
|_| std::future::pending(),
)
.await
.unwrap();
assert_eq!(result.outcome, ProviderWaitOutcome::TimedOut);
assert_eq!(
result.status.state,
i32::from(ProviderReadinessState::Persisted)
);
}
#[test]
fn different_authority_is_superseded_even_with_a_smaller_fingerprint() {
let receipt = receipt(ProviderMutationKind::Update);
let mut newer_receipt = receipt.clone();
newer_receipt
.desired
.as_mut()
.unwrap()
.provider_env_revision = 1;
let mut status = completed_status(&newer_receipt, ProviderReadinessState::Ready);
assert_eq!(
disposition(&receipt, &mut status).unwrap(),
Some(ProviderWaitOutcome::Terminal)
);
assert_eq!(status.state, i32::from(ProviderReadinessState::Superseded));
}
#[test]
fn detach_requires_revocation_and_installation_failure_is_terminal() {
let receipt = receipt(ProviderMutationKind::Detach);
let mut status = completed_status(&receipt, ProviderReadinessState::Ready);
assert!(disposition(&receipt, &mut status).is_err());
status.state = ProviderReadinessState::Revoked.into();
assert_eq!(
disposition(&receipt, &mut status).unwrap(),
Some(ProviderWaitOutcome::Complete)
);
status.state = ProviderReadinessState::Failed.into();
status.reason = ProviderReadinessReason::CredentialInstallFailed.into();
assert_eq!(
disposition(&receipt, &mut status).unwrap(),
Some(ProviderWaitOutcome::Terminal)
);
}
#[test]
fn initial_empty_epoch_requires_a_complete_policy_identity() {
let mut receipt = receipt(ProviderMutationKind::Observe);
receipt.desired.as_mut().unwrap().attachment_epoch.clear();
let mut status = completed_status(&receipt, ProviderReadinessState::Ready);
assert_eq!(
disposition(&receipt, &mut status).unwrap(),
Some(ProviderWaitOutcome::Complete)
);
receipt.desired.as_mut().unwrap().policy_hash.clear();
let mut status = completed_status(&receipt, ProviderReadinessState::Ready);
assert!(disposition(&receipt, &mut status).is_err());
}
#[test]
fn direct_status_rejects_missing_installation_and_inconsistent_authority() {
let attached = receipt(ProviderMutationKind::Attach);
let mut status = completed_status(&attached, ProviderReadinessState::Ready);
status
.observed
.as_mut()
.unwrap()
.launch_environment_installed = false;
assert!(validate_status(&status).is_err());
let mut missing_provider = attached;
missing_provider
.desired
.as_mut()
.unwrap()
.provider_id
.clear();
assert!(
validate_status(&completed_status(
&missing_provider,
ProviderReadinessState::Ready
))
.is_err()
);
let mut detach = receipt(ProviderMutationKind::Detach);
detach.desired.as_mut().unwrap().provider_id = "still-attached".into();
assert!(
validate_status(&completed_status(&detach, ProviderReadinessState::Revoked)).is_err()
);
let mut observe = receipt(ProviderMutationKind::Observe);
observe.desired.as_mut().unwrap().provider_id.clear();
assert!(
validate_status(&completed_status(&observe, ProviderReadinessState::Revoked)).is_ok()
);
}
#[tokio::test]
async fn initial_completed_status_returns_without_another_rpc() {
let receipt = receipt(ProviderMutationKind::Attach);
let status = completed_status(&receipt, ProviderReadinessState::Ready);
let result = wait_with_initial_status(
&receipt,
status,
Instant::now() + Duration::from_millis(10),
Duration::ZERO,
|_| std::future::pending(),
)
.await
.unwrap();
assert_eq!(result.outcome, ProviderWaitOutcome::Complete);
}
#[tokio::test]
async fn initial_pending_status_survives_a_hung_followup_rpc() {
let receipt = receipt(ProviderMutationKind::Attach);
let mut status = persisted_status(receipt.clone());
status.state = ProviderReadinessState::Pending.into();
status.reason = ProviderReadinessReason::WaitingForProcess.into();
let result = wait_with_initial_status(
&receipt,
status.clone(),
Instant::now() + Duration::from_millis(10),
Duration::ZERO,
|_| std::future::pending(),
)
.await
.unwrap();
assert_eq!(result.outcome, ProviderWaitOutcome::TimedOut);
assert_eq!(result.status, status);
}
#[test]
fn rejects_unbounded_waits_and_unknown_states() {
assert!(provider_wait_deadline(Duration::ZERO).is_err());
assert!(provider_wait_deadline(MAX_PROVIDER_WAIT + Duration::from_secs(1)).is_err());
let receipt = receipt(ProviderMutationKind::Attach);
let mut status = persisted_status(receipt.clone());
status.state = 99;
assert!(disposition(&receipt, &mut status).is_err());
}
#[test]
fn claimed_ready_without_process_acknowledgment_is_invalid() {
let receipt = receipt(ProviderMutationKind::Attach);
let mut status = completed_status(&receipt, ProviderReadinessState::Ready);
status
.observed
.as_mut()
.unwrap()
.launch_environment_installed = false;
assert!(disposition(&receipt, &mut status).is_err());
}
}
+10 -7
View File
@@ -23,13 +23,16 @@ pub use openshell_core::proto::open_shell_client::OpenShellClient as GrpcClient;
pub use openshell_core::proto::{
CreateSandboxRequest, CreateSandboxTemplateRequest, CreateWorkspaceRequest,
DeleteSandboxRequest, DeleteSandboxTemplateRequest, DeleteWorkspaceRequest, ExecSandboxRequest,
GetSandboxRequest, GetSandboxTemplateRequest, GetWorkspaceRequest, HealthRequest,
ListProvidersRequest, ListSandboxTemplatesRequest, ListSandboxesRequest, ListWorkspacesRequest,
Sandbox, SandboxPhase as ProtoSandboxPhase, SandboxResources, SandboxServiceLevel,
SandboxSpec as ProtoSandboxSpec, SandboxStartup, SandboxTemplate, SandboxTemplateResponse,
SandboxWorkloadConfig, SandboxWorkloadTemplate, SandboxWorkloadTemplateProvenance,
SandboxWorkloadTemplateSpec, ServiceStatus as ProtoServiceStatus, StartSandboxRequest,
StopSandboxRequest, Workspace,
GetSandboxProviderStatusRequest, GetSandboxProviderStatusResponse, GetSandboxRequest,
GetSandboxTemplateRequest, GetWorkspaceRequest, HealthRequest, ListProvidersRequest,
ListSandboxTemplatesRequest, ListSandboxesRequest, ListWorkspacesRequest,
ProviderDesiredIdentity, ProviderMutationKind, ProviderMutationReceipt,
ProviderReadinessObservation, ProviderReadinessReason, ProviderReadinessState,
ProviderReadinessStatus, Sandbox, SandboxPhase as ProtoSandboxPhase, SandboxResources,
SandboxServiceLevel, SandboxSpec as ProtoSandboxSpec, SandboxStartup, SandboxTemplate,
SandboxTemplateResponse, SandboxWorkloadConfig, SandboxWorkloadTemplate,
SandboxWorkloadTemplateProvenance, SandboxWorkloadTemplateSpec,
ServiceStatus as ProtoServiceStatus, StartSandboxRequest, StopSandboxRequest, Workspace,
};
/// Type alias for the gRPC client wrapped in the SDK's auth interceptor.
+18
View File
@@ -708,6 +708,24 @@ impl OpenShell for TestOpenShell {
Err(Status::unimplemented("unused"))
}
async fn get_sandbox_provider_status(
&self,
_request: tonic::Request<proto::GetSandboxProviderStatusRequest>,
) -> Result<Response<proto::GetSandboxProviderStatusResponse>, Status> {
Err(Status::unimplemented(
"provider readiness is not exercised by this mock",
))
}
async fn report_provider_readiness(
&self,
_request: tonic::Request<proto::ReportProviderReadinessRequest>,
) -> Result<Response<proto::ReportProviderReadinessResponse>, Status> {
Err(Status::unimplemented(
"provider installation reports are not exercised by this mock",
))
}
async fn get_sandbox_provider_environment(
&self,
_: tonic::Request<proto::GetSandboxProviderEnvironmentRequest>,
@@ -9,6 +9,7 @@ package openshell.storage.v1;
import "datamodel.proto";
import "google/protobuf/duration.proto";
import "google/protobuf/timestamp.proto";
import "openshell.proto";
import "options.proto";
import "sandbox.proto";
@@ -108,6 +109,31 @@ message StoredProviderProfile {
openshell.v1.ProviderProfile profile = 2;
}
// Durable, non-secret progress for a sandbox-scoped desired-state mutation.
// The target_* fields let a recovery worker bind the exact effective snapshot
// revision after the desired-state transaction commits.
message StoredConfigUpdateOperation {
openshell.datamodel.v1.ObjectMeta metadata = 1;
openshell.v1.ConfigUpdateOperation operation = 2;
uint32 target_policy_version = 3;
uint64 target_settings_revision = 4;
openshell.v1.SandboxPhase initial_phase = 5;
string idempotency_key = 6;
uint32 attempt_count = 7;
reserved 8;
reserved "next_attempt_at_ms";
google.protobuf.Timestamp next_attempt_time = 108;
uint32 response_policy_version = 9;
string response_policy_hash = 10;
uint64 response_settings_revision = 11;
bool response_deleted = 12;
map<string, string> response_annotations = 13;
// Immutable provider mutation intent associated with this operation.
openshell.v1.ProviderMutationReceipt provider_receipt = 14;
// Closed reason captured when resolving the operation's provider snapshot.
openshell.v1.ProviderReadinessReason provider_snapshot_reason = 15;
}
// Stored payload for a policy revision row in the generic objects table.
message PolicyRevisionPayload {
// Serialized policy contents.
@@ -0,0 +1,763 @@
// SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
// SPDX-License-Identifier: Apache-2.0
//! Durable, exact-target configuration operations, independent of delivery transport.
//!
//! Provider receipts project the common operation resource. Live installation
//! evidence remains session-bound; an applied operation records historical
//! completion and never substitutes for a fresh provider readiness evaluation.
#![allow(clippy::result_large_err)] // Internal operation helpers preserve gRPC error details.
use std::collections::HashMap;
use openshell_core::proto::{
ConfigApplyOutcome, ConfigComponent, ConfigSnapshotRevision, ConfigUpdateOperation,
ConfigUpdateOperationState, ObjectMeta, ProviderMutationKind, ProviderMutationReceipt,
ProviderReadinessReason, ProviderReadinessState, ProviderReadinessStatus,
config_snapshot_revision,
};
use openshell_core::rpc_error::{self, ErrorDetails, StatusExt};
use prost::Message;
use sha2::{Digest, Sha256};
use tonic::{Code, Status};
use crate::persistence::{ObjectRecord, ObjectType, PersistenceError, Store, WriteCondition};
use crate::storage_proto::StoredConfigUpdateOperation;
/// Object-store namespace shared by configuration completion resources.
pub const CONFIG_UPDATE_OPERATION_OBJECT_TYPE: &str = "config_update_operation";
const MAX_TRANSITION_RETRIES: usize = 8;
impl ObjectType for StoredConfigUpdateOperation {
fn object_type() -> &'static str {
CONFIG_UPDATE_OPERATION_OBJECT_TYPE
}
}
/// Provider-specific view of one common configuration operation.
#[derive(Clone, Debug)]
pub struct ProviderOperation {
/// Immutable desired authority captured for this operation.
pub(crate) receipt: ProviderMutationReceipt,
/// A failed initial snapshot is never reconstructed into another target.
pub(crate) snapshot_reason: ProviderReadinessReason,
/// Durable historical outcome; callers must separately evaluate live readiness.
pub(crate) operation: ConfigUpdateOperation,
}
fn storage_unavailable() -> Status {
// A provider mutation can already have committed when operation persistence
// fails. No retry hint is attached: repeating the mutation is not proven safe.
Status::with_error_details(
Code::Unavailable,
"configuration operation storage unavailable; the saved mutation may remain in effect",
ErrorDetails::with_error_info(
"CONFIG_OPERATION_STORAGE_UNCERTAIN",
rpc_error::ERROR_DOMAIN,
HashMap::new(),
),
)
}
fn invalid_record() -> Status {
Status::with_error_details(
Code::Internal,
"configuration operation identity is inconsistent",
ErrorDetails::with_error_info(
"CONFIG_OPERATION_INVALID",
rpc_error::ERROR_DOMAIN,
HashMap::new(),
),
)
}
fn persisted_time(receipt: &ProviderMutationReceipt) -> Result<prost_types::Timestamp, Status> {
// Receipt identity includes the full canonical timestamp. Absence is not
// the Unix epoch, and reducing nanos to milliseconds would merge identities.
let timestamp = receipt.persisted_time.ok_or_else(invalid_record)?;
openshell_core::time::validate_timestamp(&timestamp).map_err(|_| invalid_record())?;
Ok(timestamp)
}
fn observation_id(
receipt: &ProviderMutationReceipt,
snapshot_reason: ProviderReadinessReason,
) -> Result<String, Status> {
let desired = receipt.desired.as_ref().ok_or_else(invalid_record)?;
let mut digest = Sha256::new();
digest.update(b"openshell/provider-observation/v1\0");
let target_bytes = desired.encode_to_vec();
// The target contains only public identity and opaque revision fields. Its
// protobuf has no maps, so encoding is canonical. Length framing prevents
// component concatenation ambiguities across workspaces and provider names.
for component in [
receipt.workspace.as_bytes(),
receipt.provider_name.as_bytes(),
target_bytes.as_slice(),
] {
let length = u64::try_from(component.len()).map_err(|_| invalid_record())?;
digest.update(length.to_be_bytes());
digest.update(component);
}
// A failed capture and its later successful repair are different targets;
// the original failed operation must remain immutable.
digest.update((snapshot_reason as i32).to_be_bytes());
let mut bytes = [0_u8; 16];
for (byte, hashed) in bytes.iter_mut().zip(digest.finalize()) {
*byte = hashed;
}
Ok(uuid::Builder::from_custom_bytes(bytes)
.into_uuid()
.to_string())
}
/// Record an exact provider target in the shared configuration-operation store.
///
/// The caller captures the snapshot only after its provider mutation finishes.
/// This write does not roll back a preceding mutation on failure. An incomplete
/// snapshot is persisted as failed, retaining its original non-secret reason.
/// Observation-only requests reuse the original receipt for the same complete
/// target; source mutations retain their distinct caller-created receipt IDs.
pub async fn record_provider_operation(
store: &Store,
mut receipt: ProviderMutationReceipt,
snapshot_reason: ProviderReadinessReason,
) -> Result<ProviderMutationReceipt, Status> {
let observation = receipt.kind == ProviderMutationKind::Observe as i32;
if observation {
receipt.receipt_id = observation_id(&receipt, snapshot_reason)?;
}
let desired = receipt.desired.as_ref().ok_or_else(invalid_record)?;
if receipt.receipt_id.is_empty()
|| receipt.workspace.is_empty()
|| desired.sandbox_id.is_empty()
{
return Err(invalid_record());
}
let persisted_time = persisted_time(&receipt)?;
let failed = snapshot_reason != ProviderReadinessReason::Unspecified;
let operation = ConfigUpdateOperation {
operation_id: receipt.receipt_id.clone(),
sandbox_id: desired.sandbox_id.clone(),
component: ConfigComponent::ProviderEnvironment.into(),
target_revision: Some(ConfigSnapshotRevision {
component: Some(config_snapshot_revision::Component::ProviderTarget(
desired.clone(),
)),
}),
state: if failed {
ConfigUpdateOperationState::Failed.into()
} else {
ConfigUpdateOperationState::Pending.into()
},
outcome: if failed {
ConfigApplyOutcome::FailedClosed.into()
} else {
ConfigApplyOutcome::Unspecified.into()
},
sanitized_error: if failed {
snapshot_reason.as_str_name().to_string()
} else {
String::new()
},
created_time: Some(persisted_time),
updated_time: Some(persisted_time),
completed_time: failed.then_some(persisted_time),
};
let stored = StoredConfigUpdateOperation {
metadata: Some(ObjectMeta {
id: receipt.receipt_id.clone(),
name: receipt.receipt_id.clone(),
workspace: receipt.workspace.clone(),
created_time: Some(persisted_time),
..Default::default()
}),
operation: Some(operation),
provider_receipt: Some(receipt.clone()),
provider_snapshot_reason: snapshot_reason.into(),
..Default::default()
};
let result = store
.put_if(
CONFIG_UPDATE_OPERATION_OBJECT_TYPE,
&receipt.receipt_id,
&receipt.receipt_id,
&receipt.workspace,
&stored.encode_to_vec(),
None,
WriteCondition::MustCreate,
)
.await;
match result {
Ok(_) => Ok(receipt),
Err(PersistenceError::UniqueViolation { .. }) if observation => {
// Concurrent observers race only on the insert. A conflict never
// updates the winner's timestamp, mutation identity, or outcome.
// Exact comparison also fails closed on an ID collision/corruption.
let existing =
get_provider_operation(store, &receipt.receipt_id, &receipt.workspace).await?;
if existing.receipt.kind != receipt.kind
|| existing.receipt.provider_name != receipt.provider_name
|| existing.receipt.desired != receipt.desired
|| existing.snapshot_reason != snapshot_reason
{
return Err(invalid_record());
}
Ok(existing.receipt)
}
Err(_) => Err(storage_unavailable()),
}
}
// The record ID, receipt ID, and operation ID deliberately name the same
// durable identity; their distinct schema field names must compare equal.
#[allow(clippy::suspicious_operation_groupings)]
fn decode_provider_operation(
record: &ObjectRecord,
) -> Result<(StoredConfigUpdateOperation, ProviderOperation), Status> {
let stored = StoredConfigUpdateOperation::decode(record.payload.as_slice())
.map_err(|_| invalid_record())?;
let receipt = stored
.provider_receipt
.as_ref()
.ok_or_else(invalid_record)?;
let operation = stored.operation.as_ref().ok_or_else(invalid_record)?;
let metadata = stored.metadata.as_ref().ok_or_else(invalid_record)?;
let desired = receipt.desired.as_ref().ok_or_else(invalid_record)?;
let persisted_time = persisted_time(receipt)?;
let snapshot_reason = ProviderReadinessReason::try_from(stored.provider_snapshot_reason)
.map_err(|_| invalid_record())?;
if record.id != receipt.receipt_id
|| record.workspace != receipt.workspace
|| metadata.id != record.id
|| metadata.workspace != record.workspace
|| operation.operation_id != record.id
|| operation.sandbox_id != desired.sandbox_id
|| operation.created_time != Some(persisted_time)
|| metadata.created_time != Some(persisted_time)
|| operation.component != ConfigComponent::ProviderEnvironment as i32
|| operation
.target_revision
.as_ref()
.and_then(|revision| revision.component.as_ref())
!= Some(&config_snapshot_revision::Component::ProviderTarget(
desired.clone(),
))
|| ConfigUpdateOperationState::try_from(operation.state).is_err()
{
return Err(invalid_record());
}
let provider = ProviderOperation {
receipt: receipt.clone(),
snapshot_reason,
operation: operation.clone(),
};
Ok((stored, provider))
}
async fn load_provider_operation(
store: &Store,
operation_id: &str,
workspace: &str,
) -> Result<(ObjectRecord, StoredConfigUpdateOperation, ProviderOperation), Status> {
let record = store
.get(CONFIG_UPDATE_OPERATION_OBJECT_TYPE, operation_id)
.await
.map_err(|_| storage_unavailable())?
.filter(|record| record.workspace == workspace)
.ok_or_else(|| Status::not_found("provider operation not found"))?;
let (stored, provider) = decode_provider_operation(&record)?;
Ok((record, stored, provider))
}
/// Read a provider projection after the RPC has authorized the workspace.
///
/// Looking up an operation from a different workspace returns the same result
/// as a missing operation and never reveals its receipt or desired target.
pub async fn get_provider_operation(
store: &Store,
operation_id: &str,
workspace: &str,
) -> Result<ProviderOperation, Status> {
let (_, _, provider) = load_provider_operation(store, operation_id, workspace).await?;
Ok(provider)
}
fn terminal(state: ConfigUpdateOperationState) -> bool {
matches!(
state,
ConfigUpdateOperationState::Applied
| ConfigUpdateOperationState::Inactive
| ConfigUpdateOperationState::Failed
| ConfigUpdateOperationState::Superseded
| ConfigUpdateOperationState::Cancelled
)
}
fn completion(
status: &ProviderReadinessStatus,
) -> Result<Option<(ConfigUpdateOperationState, ConfigApplyOutcome)>, Status> {
match ProviderReadinessState::try_from(status.state).map_err(|_| invalid_record())? {
ProviderReadinessState::Ready | ProviderReadinessState::Revoked => {
let desired = status
.receipt
.as_ref()
.and_then(|receipt| receipt.desired.as_ref())
.ok_or_else(invalid_record)?;
let observed = status.observed.as_ref().ok_or_else(invalid_record)?;
// The RPC validates current session ownership and freshness. Check
// the complete target again before converting its result into a
// durable terminal transition; a partial install is never applied.
if status.reason != ProviderReadinessReason::Unspecified as i32
|| observed.reason != ProviderReadinessReason::Unspecified as i32
|| observed.attachment_epoch != desired.attachment_epoch
|| observed.provider_env_revision != desired.provider_env_revision
|| observed.config_revision != desired.config_revision
|| observed.policy_hash != desired.policy_hash
|| !observed.credentials_installed
|| !observed.policy_active
|| !observed.launch_environment_installed
|| observed.process_instance_id.is_empty()
|| observed.session_id.is_empty()
|| status.network_instance_id.is_empty()
|| (status.state == ProviderReadinessState::Revoked as i32)
!= desired.provider_id.is_empty()
{
return Err(invalid_record());
}
Ok(Some((
ConfigUpdateOperationState::Applied,
ConfigApplyOutcome::Applied,
)))
}
ProviderReadinessState::Superseded => Ok(Some((
ConfigUpdateOperationState::Superseded,
ConfigApplyOutcome::IgnoredStale,
))),
// Live installation failures can recover without changing the desired
// revision. They remain visible in the provider projection but do not
// terminate the operation or authorize a retry of the source mutation.
_ => Ok(None),
}
}
/// Persist exact completion by CAS and attach its historical resource to a view.
///
/// Call only after evaluating authenticated current-session evidence. The live
/// status is never changed by this function: expired or superseded evidence
/// cannot become ready because an earlier observation was durably applied.
pub async fn observe_provider_status(
store: &Store,
status: &mut ProviderReadinessStatus,
) -> Result<(), Status> {
let receipt = status.receipt.as_ref().ok_or_else(invalid_record)?.clone();
let completion = completion(status)?;
for _ in 0..MAX_TRANSITION_RETRIES {
let (record, mut stored, provider) =
load_provider_operation(store, &receipt.receipt_id, &receipt.workspace).await?;
if provider.receipt != receipt {
return Err(invalid_record());
}
let state = ConfigUpdateOperationState::try_from(provider.operation.state)
.map_err(|_| invalid_record())?;
let Some((terminal_state, outcome)) = completion.filter(|_| !terminal(state)) else {
status.operation = Some(provider.operation);
return Ok(());
};
// Capturing the desired snapshot failed permanently for this operation.
// A later read must not fabricate a different, successful target.
if provider.snapshot_reason != ProviderReadinessReason::Unspecified {
return Err(invalid_record());
}
let operation = stored.operation.as_mut().ok_or_else(invalid_record)?;
operation.state = terminal_state.into();
operation.outcome = outcome.into();
operation.sanitized_error = if terminal_state == ConfigUpdateOperationState::Superseded {
ProviderReadinessReason::DesiredStateChanged
.as_str_name()
.to_string()
} else {
String::new()
};
let completed_time =
openshell_core::time::timestamp_from_system_time(std::time::SystemTime::now())
.map_err(|error| {
Status::internal(format!("create operation completion timestamp: {error}"))
})?;
operation.updated_time = Some(completed_time);
operation.completed_time = Some(completed_time);
let result = store
.put_if(
CONFIG_UPDATE_OPERATION_OBJECT_TYPE,
&record.id,
&record.name,
&record.workspace,
&stored.encode_to_vec(),
record.labels.as_deref(),
WriteCondition::MatchResourceVersion(record.resource_version),
)
.await;
match result {
Ok(_) => {
status.operation = stored.operation;
return Ok(());
}
// A competing observer may have completed this operation. Reload
// the authoritative row; a terminal outcome is immutable.
Err(PersistenceError::Conflict { .. }) => {}
Err(_) => return Err(storage_unavailable()),
}
}
Err(rpc_error::resource_version_conflict(
"configuration operation changed concurrently; query its status again",
None,
))
}
#[cfg(test)]
mod tests {
use super::*;
use openshell_core::proto::{
ProviderDesiredIdentity, ProviderMutationKind, ProviderReadinessObservation,
};
use uuid::Uuid;
fn receipt() -> ProviderMutationReceipt {
ProviderMutationReceipt {
receipt_id: Uuid::new_v4().to_string(),
mutation_id: Uuid::new_v4().to_string(),
provider_name: "provider".to_string(),
workspace: "default".to_string(),
kind: ProviderMutationKind::Update.into(),
desired: Some(ProviderDesiredIdentity {
sandbox_id: Uuid::new_v4().to_string(),
sandbox_name: "sandbox".to_string(),
attachment_epoch: Uuid::new_v4().to_string(),
provider_id: Uuid::new_v4().to_string(),
provider_resource_version: 3,
provider_env_revision: 5,
config_revision: 7,
policy_hash: "policy".to_string(),
}),
persisted_time: Some(prost_types::Timestamp {
seconds: 1_700_000_000,
nanos: 123_456_789,
}),
}
}
fn ready(receipt: &ProviderMutationReceipt) -> ProviderReadinessStatus {
let desired = receipt.desired.as_ref().unwrap();
ProviderReadinessStatus {
receipt: Some(receipt.clone()),
state: ProviderReadinessState::Ready.into(),
network_instance_id: Uuid::new_v4().to_string(),
observed: Some(ProviderReadinessObservation {
session_id: Uuid::new_v4().to_string(),
sequence: 1,
attachment_epoch: desired.attachment_epoch.clone(),
provider_env_revision: desired.provider_env_revision,
config_revision: desired.config_revision,
policy_hash: desired.policy_hash.clone(),
credentials_installed: true,
policy_active: true,
launch_environment_installed: true,
process_instance_id: Uuid::new_v4().to_string(),
reason: ProviderReadinessReason::Unspecified.into(),
}),
..Default::default()
}
}
#[tokio::test]
async fn provider_receipt_uses_the_common_operation_namespace_and_exact_target() {
let store = crate::persistence::test_store().await;
let receipt = receipt();
record_provider_operation(
&store,
receipt.clone(),
ProviderReadinessReason::Unspecified,
)
.await
.unwrap();
let operation = get_provider_operation(&store, &receipt.receipt_id, "default")
.await
.unwrap();
assert_eq!(operation.receipt, receipt);
assert_eq!(operation.operation.operation_id, receipt.receipt_id);
assert_eq!(operation.operation.created_time, receipt.persisted_time);
assert_eq!(operation.operation.updated_time, receipt.persisted_time);
assert!(operation.operation.completed_time.is_none());
assert_eq!(
operation.operation.state,
ConfigUpdateOperationState::Pending as i32
);
assert!(
store
.get("provider_mutation_receipt", &receipt.receipt_id)
.await
.unwrap()
.is_none()
);
assert_eq!(
get_provider_operation(&store, &receipt.receipt_id, "other")
.await
.unwrap_err()
.code(),
Code::NotFound
);
}
#[tokio::test]
async fn incomplete_provider_target_stays_failed_after_later_installation() {
let store = crate::persistence::test_store().await;
let receipt = receipt();
record_provider_operation(
&store,
receipt.clone(),
ProviderReadinessReason::SnapshotMismatch,
)
.await
.unwrap();
let mut status = ready(&receipt);
observe_provider_status(&store, &mut status).await.unwrap();
let operation = status.operation.unwrap();
assert_eq!(operation.state, ConfigUpdateOperationState::Failed as i32);
assert_eq!(operation.completed_time, receipt.persisted_time);
}
#[tokio::test]
async fn receipt_timestamp_requires_presence_and_canonical_nanos() {
let store = crate::persistence::test_store().await;
for invalid_time in [
None,
Some(prost_types::Timestamp {
seconds: 0,
nanos: -1,
}),
Some(prost_types::Timestamp {
seconds: openshell_core::time::MAX_TIMESTAMP_SECONDS + 1,
nanos: 0,
}),
] {
let mut receipt = receipt();
receipt.persisted_time = invalid_time;
let error =
record_provider_operation(&store, receipt, ProviderReadinessReason::Unspecified)
.await
.unwrap_err();
assert_eq!(error.code(), Code::Internal);
}
assert_eq!(
store
.count_in_workspace(CONFIG_UPDATE_OPERATION_OBJECT_TYPE, "default")
.await
.unwrap(),
0
);
// The Unix epoch is a valid explicit timestamp, not missing data.
let mut epoch = receipt();
epoch.persisted_time = Some(prost_types::Timestamp::default());
let recorded =
record_provider_operation(&store, epoch.clone(), ProviderReadinessReason::Unspecified)
.await
.unwrap();
assert_eq!(recorded, epoch);
assert_eq!(
get_provider_operation(&store, &epoch.receipt_id, "default")
.await
.unwrap()
.receipt,
epoch
);
}
#[tokio::test]
async fn receipt_timestamp_nanos_are_part_of_exact_completion_identity() {
let store = crate::persistence::test_store().await;
let receipt = receipt();
record_provider_operation(
&store,
receipt.clone(),
ProviderReadinessReason::Unspecified,
)
.await
.unwrap();
let mut changed = ready(&receipt);
changed
.receipt
.as_mut()
.unwrap()
.persisted_time
.as_mut()
.unwrap()
.nanos += 1;
assert_eq!(
observe_provider_status(&store, &mut changed)
.await
.unwrap_err()
.code(),
Code::Internal
);
let stored = get_provider_operation(&store, &receipt.receipt_id, "default")
.await
.unwrap();
assert_eq!(stored.receipt, receipt);
assert_eq!(
stored.operation.state,
ConfigUpdateOperationState::Pending as i32
);
assert!(stored.operation.completed_time.is_none());
let mut exact = ready(&receipt);
observe_provider_status(&store, &mut exact).await.unwrap();
let completed = exact.operation.unwrap();
assert_eq!(completed.created_time, receipt.persisted_time);
assert!(completed.completed_time.is_some());
assert_eq!(completed.updated_time, completed.completed_time);
openshell_core::time::validate_timestamp(completed.completed_time.as_ref().unwrap())
.unwrap();
}
#[tokio::test]
async fn stale_or_partial_provider_evidence_cannot_complete_an_operation() {
let store = crate::persistence::test_store().await;
let receipt = receipt();
record_provider_operation(
&store,
receipt.clone(),
ProviderReadinessReason::Unspecified,
)
.await
.unwrap();
let mut stale = ready(&receipt);
stale.observed.as_mut().unwrap().config_revision += 1;
assert!(observe_provider_status(&store, &mut stale).await.is_err());
let mut partial = ready(&receipt);
partial
.observed
.as_mut()
.unwrap()
.launch_environment_installed = false;
assert!(observe_provider_status(&store, &mut partial).await.is_err());
assert_eq!(
get_provider_operation(&store, &receipt.receipt_id, "default")
.await
.unwrap()
.operation
.state,
ConfigUpdateOperationState::Pending as i32
);
}
#[tokio::test]
async fn applied_history_never_upgrades_an_expired_live_view() {
let store = crate::persistence::test_store().await;
let receipt = receipt();
record_provider_operation(
&store,
receipt.clone(),
ProviderReadinessReason::Unspecified,
)
.await
.unwrap();
let mut status = ready(&receipt);
observe_provider_status(&store, &mut status).await.unwrap();
let before = store
.get(CONFIG_UPDATE_OPERATION_OBJECT_TYPE, &receipt.receipt_id)
.await
.unwrap()
.unwrap();
status.state = ProviderReadinessState::Pending.into();
status.reason = ProviderReadinessReason::SupervisorLeaseExpired.into();
observe_provider_status(&store, &mut status).await.unwrap();
assert_eq!(status.state, ProviderReadinessState::Pending as i32);
assert_eq!(
status.operation.unwrap().state,
ConfigUpdateOperationState::Applied as i32
);
let after = store
.get(CONFIG_UPDATE_OPERATION_OBJECT_TYPE, &receipt.receipt_id)
.await
.unwrap()
.unwrap();
assert_eq!(before.resource_version, after.resource_version);
}
#[tokio::test]
async fn competing_terminal_observers_preserve_the_first_committed_outcome() {
let store = crate::persistence::test_store().await;
let receipt = receipt();
record_provider_operation(
&store,
receipt.clone(),
ProviderReadinessReason::Unspecified,
)
.await
.unwrap();
let mut applied = ready(&receipt);
let mut superseded = applied.clone();
superseded.state = ProviderReadinessState::Superseded.into();
superseded.reason = ProviderReadinessReason::DesiredStateChanged.into();
let (first, second) = tokio::join!(
observe_provider_status(&store, &mut applied),
observe_provider_status(&store, &mut superseded),
);
first.unwrap();
second.unwrap();
assert_eq!(applied.operation, superseded.operation);
let stored = get_provider_operation(&store, &receipt.receipt_id, "default")
.await
.unwrap();
assert!(matches!(
ConfigUpdateOperationState::try_from(stored.operation.state).unwrap(),
ConfigUpdateOperationState::Applied | ConfigUpdateOperationState::Superseded
));
}
#[tokio::test]
async fn failed_observation_and_recovered_snapshot_have_distinct_immutable_receipts() {
let store = crate::persistence::test_store().await;
let mut observed = receipt();
observed.kind = ProviderMutationKind::Observe.into();
let failed = record_provider_operation(
&store,
observed.clone(),
ProviderReadinessReason::CredentialsWithheld,
)
.await
.unwrap();
observed.mutation_id = Uuid::new_v4().to_string();
observed.persisted_time.as_mut().unwrap().nanos += 1;
let repeated = record_provider_operation(
&store,
observed.clone(),
ProviderReadinessReason::CredentialsWithheld,
)
.await
.unwrap();
assert_eq!(repeated, failed);
let repaired =
record_provider_operation(&store, observed, ProviderReadinessReason::Unspecified)
.await
.unwrap();
assert_ne!(repaired.receipt_id, failed.receipt_id);
assert_eq!(
get_provider_operation(&store, &failed.receipt_id, "default")
.await
.unwrap()
.operation
.state,
ConfigUpdateOperationState::Failed as i32
);
assert_eq!(
store
.count_in_workspace(CONFIG_UPDATE_OPERATION_OBJECT_TYPE, "default")
.await
.unwrap(),
2
);
}
}
+24 -7
View File
@@ -7,6 +7,7 @@ mod auth_rpc;
pub mod mutation_replay;
pub mod policy;
pub mod provider;
pub mod provider_readiness;
mod sandbox;
pub use sandbox::mint_persisted_authentication;
mod service;
@@ -37,7 +38,8 @@ use openshell_core::proto::{
GetProviderRefreshStatusResponse, GetProviderRequest, GetSandboxConfigRequest,
GetSandboxConfigResponse, GetSandboxLogsRequest, GetSandboxLogsResponse,
GetSandboxPolicyStatusRequest, GetSandboxPolicyStatusResponse,
GetSandboxProviderEnvironmentRequest, GetSandboxProviderEnvironmentResponse, GetSandboxRequest,
GetSandboxProviderEnvironmentRequest, GetSandboxProviderEnvironmentResponse,
GetSandboxProviderStatusRequest, GetSandboxProviderStatusResponse, GetSandboxRequest,
GetSandboxTemplateRequest, GetServiceRequest, GetWorkspaceRequest, GetWorkspaceResponse,
GpuResourceCapabilities, HealthRequest, HealthResponse, ImportProviderProfilesRequest,
ImportProviderProfilesResponse, IssueSandboxTokenRequest, IssueSandboxTokenResponse,
@@ -52,12 +54,13 @@ use openshell_core::proto::{
RefreshSandboxTokenResponse, RejectDraftChunkRequest, RejectDraftChunkResponse, RelayFrame,
RemoveWorkspaceMemberRequest, RemoveWorkspaceMemberResponse, ReportEndpointStatusRequest,
ReportEndpointStatusResponse, ReportMainProcessExitRequest, ReportMainProcessExitResponse,
ReportPolicyStatusRequest, ReportPolicyStatusResponse, ResourceCapabilities,
RevokeSshSessionRequest, RevokeSshSessionResponse, RotateProviderCredentialRequest,
RotateProviderCredentialResponse, SandboxResponse, SandboxTemplateResponse,
ServiceEndpointResponse, ServiceStatus, StartSandboxRequest, StopSandboxRequest,
SubmitPolicyAnalysisRequest, SubmitPolicyAnalysisResponse, SupervisorMessage, TcpForwardFrame,
UndoDraftChunkRequest, UndoDraftChunkResponse, UpdateConfigRequest, UpdateConfigResponse,
ReportPolicyStatusRequest, ReportPolicyStatusResponse, ReportProviderReadinessRequest,
ReportProviderReadinessResponse, ResourceCapabilities, RevokeSshSessionRequest,
RevokeSshSessionResponse, RotateProviderCredentialRequest, RotateProviderCredentialResponse,
SandboxResponse, SandboxTemplateResponse, ServiceEndpointResponse, ServiceStatus,
StartSandboxRequest, StopSandboxRequest, SubmitPolicyAnalysisRequest,
SubmitPolicyAnalysisResponse, SupervisorMessage, TcpForwardFrame, UndoDraftChunkRequest,
UndoDraftChunkResponse, UpdateConfigRequest, UpdateConfigResponse,
UpdateProviderProfilesRequest, UpdateProviderProfilesResponse, UpdateProviderRequest,
WatchSandboxRequest, open_shell_server::OpenShell,
};
@@ -349,6 +352,20 @@ impl OpenShell for OpenShellService {
mutation_replay::run(&self.state, request).await
}
async fn get_sandbox_provider_status(
&self,
request: Request<GetSandboxProviderStatusRequest>,
) -> Result<Response<GetSandboxProviderStatusResponse>, Status> {
provider_readiness::handle_get_sandbox_provider_status(&self.state, request).await
}
async fn report_provider_readiness(
&self,
request: Request<ReportProviderReadinessRequest>,
) -> Result<Response<ReportProviderReadinessResponse>, Status> {
provider_readiness::handle_report_provider_readiness(&self.state, request).await
}
async fn delete_sandbox(
&self,
request: Request<DeleteSandboxRequest>,
@@ -20,8 +20,8 @@ use openshell_core::proto::{
DeleteSandboxRequest, DeleteSandboxResponse, DeleteServiceRequest, DeleteServiceResponse,
DetachSandboxProviderRequest, DetachSandboxProviderResponse, EditDraftChunkRequest,
EditDraftChunkResponse, ExposeServiceRequest, ImportProviderProfilesRequest,
ImportProviderProfilesResponse, Provider, ProviderProfile, ProviderProfileDiagnostic,
ProviderResponse, RejectDraftChunkRequest, RejectDraftChunkResponse,
ImportProviderProfilesResponse, Provider, ProviderMutationReceipt, ProviderProfile,
ProviderProfileDiagnostic, ProviderResponse, RejectDraftChunkRequest, RejectDraftChunkResponse,
RotateProviderCredentialRequest, RotateProviderCredentialResponse, Sandbox, SandboxResponse,
ServiceEndpointResponse, StartSandboxRequest, StopSandboxRequest, UndoDraftChunkRequest,
UndoDraftChunkResponse, UpdateConfigRequest, UpdateConfigResponse,
@@ -37,7 +37,6 @@ use super::{
Mutation, Scope, Success, global_scope, named_scope, replay_unavailable, restore_resource,
storage_error, uncertain,
};
use crate::ServerState;
use crate::auth::principal::Principal;
use crate::auth::workspace_authz::{
MinWorkspaceRole, authorize_workspace, authorize_workspace_selector,
@@ -48,6 +47,7 @@ use crate::storage_proto::{
StoredProviderCredentialRefreshStateV2 as StoredProviderCredentialRefreshState,
StoredProviderProfile,
};
use crate::{ServerState, config_update_operation};
#[derive(Clone, Serialize, Deserialize)]
pub(in crate::grpc) struct Reference {
@@ -159,7 +159,16 @@ pub(in crate::grpc) enum Outcome {
id: String,
outcome: i32,
},
Provider(Reference),
Attachment {
id: String,
changed: bool,
receipt: ProviderReceiptReference,
},
Provider {
reference: Reference,
mutation_id: String,
target_receipts: Vec<ProviderReceiptReference>,
},
Service {
reference: Reference,
sandbox_id: String,
@@ -189,6 +198,32 @@ pub(in crate::grpc) enum Outcome {
Refresh(Refresh),
}
/// Replay retains immutable operation identities, never a fresh target snapshot
/// or a serialized provider response that could carry credential material.
#[derive(Serialize, Deserialize)]
pub(in crate::grpc) struct ProviderReceiptReference {
id: String,
workspace: String,
}
impl ProviderReceiptReference {
fn new(receipt: &ProviderMutationReceipt) -> Self {
Self {
id: receipt.receipt_id.clone(),
workspace: receipt.workspace.clone(),
}
}
async fn restore(self, store: &Store) -> Result<ProviderMutationReceipt, Status> {
// The original operation is authoritative even after readiness changes.
// A missing operation cannot authorize replaying the mutation itself.
config_update_operation::get_provider_operation(store, &self.id, &self.workspace)
.await
.map(|operation| operation.receipt)
.map_err(|_| replay_unavailable())
}
}
/// Public profile declarations and diagnostics have a nonsecret contract. These
/// fields may echo declaration text; they are not arbitrary sanitized payloads.
#[derive(Serialize, Deserialize)]
@@ -388,17 +423,34 @@ macro_rules! attachment_mutation {
$method,
$handler,
User,
|response: &Response<$resp>| sandbox_receipt(
response.get_ref().sandbox.as_ref(),
response.get_ref().$field
),
|response: &Response<$resp>| {
let response = response.get_ref();
Ok(Outcome::Attachment {
id: response
.sandbox
.as_ref()
.ok_or_else(uncertain)?
.object_id()
.into(),
changed: response.$field,
receipt: ProviderReceiptReference::new(
response.receipt.as_ref().ok_or_else(uncertain)?,
),
})
},
async |store: &Store, outcome: Outcome| {
let Outcome::Sandbox { id, changed } = outcome else {
let Outcome::Attachment {
id,
changed,
receipt,
} = outcome
else {
return Err(replay_unavailable());
};
Ok($resp {
sandbox: Some(live(store, &id).await?),
$field: changed,
receipt: Some(receipt.restore(store).await?),
})
}
);
@@ -480,17 +532,37 @@ macro_rules! provider_mutation {
$method,
$handler,
Admin,
|response: &Response<ProviderResponse>| Ok(Outcome::Provider(Reference::new(
response.get_ref().provider.as_ref().ok_or_else(uncertain)?
))),
|response: &Response<ProviderResponse>| {
let response = response.get_ref();
Ok(Outcome::Provider {
reference: Reference::new(response.provider.as_ref().ok_or_else(uncertain)?),
mutation_id: response.mutation_id.clone(),
target_receipts: response
.target_receipts
.iter()
.map(ProviderReceiptReference::new)
.collect(),
})
},
async |store: &Store, outcome: Outcome| {
let Outcome::Provider(reference) = outcome else {
let Outcome::Provider {
reference,
mutation_id,
target_receipts,
} = outcome
else {
return Err(replay_unavailable());
};
let mut receipts = Vec::with_capacity(target_receipts.len());
for receipt in target_receipts {
receipts.push(receipt.restore(store).await?);
}
Ok(ProviderResponse {
provider: Some(provider::redact_provider_credentials(
reference.restore(store).await?,
)),
mutation_id,
target_receipts: receipts,
})
}
);
@@ -455,6 +455,121 @@ async fn service_deletion_replays_without_parent_and_does_not_delete_replacement
);
}
#[tokio::test]
async fn provider_replay_preserves_receipts_after_attachment_changes() {
let (_directory, state) = protected_state().await;
let created = run(
&state,
authed_request(CreateProviderRequest {
provider: Some(Provider {
metadata: Some(meta("receipt-provider")),
r#type: "openai".into(),
credentials: HashMap::from([("OPENAI_API_KEY".into(), "private-value".into())]),
..Default::default()
}),
workspace_scope: Some(scope()),
request_id: id(),
}),
)
.await
.unwrap()
.into_inner();
let sandbox = Sandbox {
metadata: Some(meta("receipt-sandbox")),
spec: Some(SandboxSpec::default()),
..Default::default()
};
state.store.put_message(&sandbox).await.unwrap();
let attach = AttachSandboxProviderRequest {
sandbox_name: "receipt-sandbox".into(),
provider_name: "receipt-provider".into(),
workspace_scope: Some(scope()),
request_id: id(),
..Default::default()
};
let attached = run(&state, authed_request(attach.clone()))
.await
.unwrap()
.into_inner();
assert!(attached.attached);
assert!(attached.receipt.is_some());
assert_eq!(replay(&state, attach.clone()).await, attached);
let mut provider = created.provider.unwrap();
provider.credentials = HashMap::from([("OPENAI_API_KEY".into(), "replacement-value".into())]);
let update = UpdateProviderRequest {
provider: Some(provider),
workspace_scope: Some(scope()),
request_id: id(),
..Default::default()
};
let updated = run(&state, authed_request(update.clone()))
.await
.unwrap()
.into_inner();
assert!(!updated.mutation_id.is_empty());
assert_eq!(updated.target_receipts.len(), 1);
assert_eq!(updated.target_receipts[0].mutation_id, updated.mutation_id);
let detach = DetachSandboxProviderRequest {
sandbox_name: "receipt-sandbox".into(),
provider_name: "receipt-provider".into(),
workspace_scope: Some(scope()),
request_id: id(),
..Default::default()
};
let detached = run(&state, authed_request(detach.clone()))
.await
.unwrap()
.into_inner();
assert!(detached.detached);
assert!(detached.receipt.is_some());
assert_eq!(replay(&state, detach.clone()).await, detached);
// Replay returns the original target and epoch even though it is detached
// now. It neither selects new targets nor executes the stale attach again.
assert_eq!(replay(&state, update.clone()).await, updated);
let replayed_attach = replay(&state, attach).await;
assert_eq!(replayed_attach.receipt, attached.receipt);
assert!(replayed_attach.attached);
assert_eq!(replayed_attach.sandbox, detached.sandbox);
let operations = state
.store
.count_in_workspace(
config_update_operation::CONFIG_UPDATE_OPERATION_OBJECT_TYPE,
"default",
)
.await
.unwrap();
assert_eq!(operations, 3);
// Lost operation evidence makes replay unavailable. It must not rotate
// credentials or produce a replacement receipt to repair that evidence.
state
.store
.delete(
config_update_operation::CONFIG_UPDATE_OPERATION_OBJECT_TYPE,
&updated.target_receipts[0].receipt_id,
)
.await
.unwrap();
assert_eq!(
reason(&run(&state, authed_request(update)).await.unwrap_err()),
"REQUEST_REPLAY_UNAVAILABLE"
);
assert_eq!(replay(&state, detach).await, detached);
let current: Provider = state
.store
.get_message(updated.provider.as_ref().unwrap().object_id())
.await
.unwrap()
.unwrap();
assert_eq!(
current.get_resource_version(),
updated.provider.unwrap().get_resource_version()
);
}
#[tokio::test]
async fn provider_replay_is_redacted_and_update_does_not_recheck_stale_version() {
let (_directory, state) = protected_state().await;
+115 -10
View File
@@ -51,11 +51,11 @@ use openshell_core::proto::{
GetSandboxPolicyStatusRequest, GetSandboxPolicyStatusResponse,
GetSandboxProviderEnvironmentRequest, GetSandboxProviderEnvironmentResponse,
ListSandboxPoliciesRequest, ListSandboxPoliciesResponse, PolicyChunk, PolicyMergeOperation,
PolicySource, PolicyStatus, PushSandboxLogsRequest, PushSandboxLogsResponse,
RejectDraftChunkRequest, RejectDraftChunkResponse, ReportPolicyStatusRequest,
ReportPolicyStatusResponse, SandboxLogLine, SandboxPolicyRevision, SettingScope, SettingValue,
SubmitPolicyAnalysisRequest, SubmitPolicyAnalysisResponse, UndoDraftChunkRequest,
UndoDraftChunkResponse, UpdateConfigRequest, UpdateConfigResponse,
PolicySource, PolicyStatus, ProviderReadinessReason, PushSandboxLogsRequest,
PushSandboxLogsResponse, RejectDraftChunkRequest, RejectDraftChunkResponse,
ReportPolicyStatusRequest, ReportPolicyStatusResponse, SandboxLogLine, SandboxPolicyRevision,
SettingScope, SettingValue, SubmitPolicyAnalysisRequest, SubmitPolicyAnalysisResponse,
UndoDraftChunkRequest, UndoDraftChunkResponse, UpdateConfigRequest, UpdateConfigResponse,
};
use openshell_core::proto::{
L7DenyRule, L7Rule, NetworkBinary, NetworkEndpoint, NetworkPolicyRule, Provider, Sandbox,
@@ -2499,6 +2499,16 @@ pub(super) async fn handle_get_sandbox_config(
let sandbox =
super::sandbox::fetch_and_authorize_sandbox(state, &principal, &sandbox_id).await?;
Ok(Response::new(load_sandbox_config(state, &sandbox).await?))
}
/// Resolve the same effective configuration for authenticated RPCs and trusted
/// readiness snapshots. Callers must authorize external access before entering.
pub(super) async fn load_sandbox_config(
state: &Arc<ServerState>,
sandbox: &Sandbox,
) -> Result<GetSandboxConfigResponse, Status> {
let sandbox_id = sandbox.object_id().to_string();
let workspace = sandbox.object_workspace().to_string();
let sandbox_provider_names = sandbox
.spec
@@ -2710,7 +2720,7 @@ pub(super) async fn handle_get_sandbox_config(
)
.await?;
Ok(Response::new(GetSandboxConfigResponse {
Ok(GetSandboxConfigResponse {
policy,
version,
policy_hash,
@@ -2727,7 +2737,12 @@ pub(super) async fn handle_get_sandbox_config(
.as_str()
.to_string(),
extension_authentication_enabled: state.sandbox_jwt_issuer.is_some(),
}))
provider_attachment_epoch: sandbox
.spec
.as_ref()
.map(|spec| spec.provider_attachment_epoch.clone())
.unwrap_or_default(),
})
}
#[cfg(test)]
@@ -3245,6 +3260,20 @@ pub(super) async fn handle_get_sandbox_provider_environment(
.await
.map_err(|e| Status::internal(format!("fetch sandbox failed: {e}")))?
.ok_or_else(|| Status::not_found("sandbox not found"))?;
Ok(Response::new(
load_sandbox_provider_environment(state, &sandbox, supports_static_credential_bindings)
.await?,
))
}
/// Materialize a privileged provider snapshot after the caller has authorized
/// access. Omission reasons remain separate from a legitimately empty snapshot.
pub(super) async fn load_sandbox_provider_environment(
state: &Arc<ServerState>,
sandbox: &Sandbox,
supports_static_credential_bindings: bool,
) -> Result<GetSandboxProviderEnvironmentResponse, Status> {
let sandbox_id = sandbox.object_id().to_string();
let workspace = sandbox.object_workspace().to_string();
let spec = sandbox
@@ -3267,7 +3296,7 @@ pub(super) async fn handle_get_sandbox_provider_environment(
state.as_ref(),
&provider_profile_catalog,
&workspace,
&sandbox,
sandbox,
&sandbox_id,
)
.await?;
@@ -3295,6 +3324,8 @@ pub(super) async fn handle_get_sandbox_provider_environment(
)
.await?;
let mut readiness_reason = provider_environment.readiness_reason;
if supports_static_credential_bindings {
let unbound_static_keys = provider_environment
.static_credential_keys
@@ -3306,6 +3337,9 @@ pub(super) async fn handle_get_sandbox_provider_environment(
})
.cloned()
.collect::<Vec<_>>();
if !unbound_static_keys.is_empty() {
readiness_reason = ProviderReadinessReason::CredentialsWithheld;
}
for key in unbound_static_keys {
warn!(
sandbox_id = %sandbox_id,
@@ -3319,6 +3353,9 @@ pub(super) async fn handle_get_sandbox_provider_environment(
provider_environment.static_credential_keys.remove(&key);
}
} else {
if !provider_environment.static_credential_keys.is_empty() {
readiness_reason = ProviderReadinessReason::UnsupportedSupervisor;
}
for key in &provider_environment.static_credential_keys {
provider_environment.environment.remove(key);
provider_environment.credential_expiration_times.remove(key);
@@ -3351,14 +3388,17 @@ pub(super) async fn handle_get_sandbox_provider_environment(
.map(|timestamp| (key, timestamp))
})
.collect();
Ok(Response::new(GetSandboxProviderEnvironmentResponse {
Ok(GetSandboxProviderEnvironmentResponse {
environment: provider_environment.environment,
provider_env_revision,
credential_expiration_times,
dynamic_credentials: provider_environment.dynamic_credentials,
static_credential_bindings: provider_environment.static_credential_bindings,
non_secret_environment_keys,
}))
provider_attachment_epoch: spec.provider_attachment_epoch.clone(),
policy_hash: deterministic_policy_hash(&effective_policy),
readiness_reason: readiness_reason.into(),
})
}
// ---------------------------------------------------------------------------
@@ -11667,6 +11707,63 @@ mod tests {
assert_eq!(v2_env.get("GITHUB_TOKEN"), Some(&"ghp-test".to_string()));
}
#[tokio::test]
async fn provider_readiness_snapshot_uses_baseline_static_binding_contract() {
let state = test_server_state().await;
state
.store
.put_message(&test_provider("work-github", "github"))
.await
.unwrap();
let mut sandbox = test_sandbox(
"sb-static-ready",
"static-ready",
test_policy_with_rule("sandbox_only", "sandbox.example.com"),
vec!["work-github".to_string()],
);
let attachment_epoch = uuid::Uuid::new_v4().to_string();
sandbox.spec.as_mut().unwrap().provider_attachment_epoch = attachment_epoch.clone();
state.store.put_message(&sandbox).await.unwrap();
// Static binding support is sufficient to install this snapshot. The
// environment and policy must describe the same attachment authority.
let response = handle_get_sandbox_provider_environment(
&state,
with_user(Request::new(GetSandboxProviderEnvironmentRequest {
sandbox_id: "sb-static-ready".to_string(),
supports_static_credential_bindings: true,
})),
)
.await
.unwrap()
.into_inner();
let config = load_sandbox_config(&state, &sandbox).await.unwrap();
assert_eq!(
response.readiness_reason,
ProviderReadinessReason::Unspecified as i32
);
assert_eq!(response.provider_attachment_epoch, attachment_epoch);
assert_eq!(
response.provider_attachment_epoch,
config.provider_attachment_epoch
);
assert_eq!(response.provider_env_revision, config.provider_env_revision);
assert_eq!(response.policy_hash, config.policy_hash);
assert!(response.environment.contains_key("GITHUB_TOKEN"));
assert!(
response
.static_credential_bindings
.contains_key("GITHUB_TOKEN")
);
assert!(
!response
.non_secret_environment_keys
.iter()
.any(|key| key == "GITHUB_TOKEN")
);
}
#[tokio::test]
async fn provider_environment_withholds_static_credentials_from_legacy_supervisors() {
use openshell_core::proto::GetSandboxProviderEnvironmentRequest;
@@ -11701,6 +11798,10 @@ mod tests {
assert!(!response.environment.contains_key("GITHUB_TOKEN"));
assert!(response.static_credential_bindings.is_empty());
assert_eq!(
response.readiness_reason,
ProviderReadinessReason::UnsupportedSupervisor as i32
);
}
#[tokio::test]
@@ -11740,6 +11841,10 @@ mod tests {
.into_inner();
assert!(!response.environment.contains_key("OPENAI_API_KEY"));
assert_eq!(
response.readiness_reason,
ProviderReadinessReason::CredentialsWithheld as i32
);
assert!(
!response
.static_credential_bindings
+519 -67
View File
@@ -66,6 +66,9 @@ pub(super) fn redact_provider_credentials(mut provider: Provider) -> Provider {
#[derive(Debug, Clone, Default, PartialEq)]
pub(super) struct ProviderEnvironment {
/// A closed reason for omitted injectable material. An empty environment
/// alone cannot distinguish successful revocation from withheld authority.
pub readiness_reason: openshell_core::proto::ProviderReadinessReason,
pub environment: HashMap<String, String>,
pub credential_expiration_times: HashMap<String, i64>,
pub dynamic_credentials: HashMap<String, ProviderProfileCredential>,
@@ -455,7 +458,6 @@ async fn update_provider_record_validating(
candidate.object_name(),
candidate.object_workspace(),
candidate.object_id(),
&removed_credential_handles,
&updated_credential_values,
&existing_handles,
)
@@ -464,9 +466,6 @@ async fn update_provider_record_validating(
candidate.credential_handles.remove(key);
candidate.credentials.remove(key);
}
for key in credential_update.deferred_store_values.keys() {
candidate.credentials.remove(key);
}
if credentials.is_some_and(crate::credentials::CredentialRuntime::stores_provider_credentials) {
for key in updated_credential_values.keys() {
candidate.credentials.remove(key);
@@ -533,16 +532,25 @@ async fn update_provider_record_validating(
}
};
finish_provider_credential_update(
// The provider CAS already excludes these handles. Keep the committed
// result and its receipts available if retirement fails; the unused
// backend objects still require cleanup.
if let Err(err) = finish_provider_credential_update(
credentials,
candidate.object_name(),
candidate.object_workspace(),
candidate.object_id(),
credential_update,
&removed_credential_handles,
&existing_handles,
)
.await?;
.await
{
warn!(
provider_name = %candidate.object_name(),
code = ?err.code(),
"failed to retire unused provider credentials after publication"
);
}
// Update resource_version from successful write
if let Some(metadata) = candidate.metadata.as_mut() {
@@ -784,16 +792,17 @@ fn credential_handles_removed_by_update(
#[derive(Debug, Clone, Default)]
struct ProviderCredentialUpdate {
pre_stored_handles: HashMap<String, CredentialHandle>,
deferred_store_values: HashMap<String, String>,
replaced_handles: HashMap<String, CredentialHandle>,
}
// Each candidate owns distinct backend objects before its provider CAS. A
// published resource version therefore identifies fully stored credentials,
// and a concurrent loser cannot overwrite the winner's credential values.
async fn prepare_provider_credential_update(
credentials: Option<&crate::credentials::CredentialRuntime>,
provider_name: &str,
workspace: &str,
provider_id: &str,
_removed_handles: &HashMap<String, CredentialHandle>,
updated_values: &HashMap<String, String>,
existing_handles: &HashMap<String, CredentialHandle>,
) -> Result<ProviderCredentialUpdate, Status> {
@@ -804,42 +813,41 @@ async fn prepare_provider_credential_update(
return Ok(ProviderCredentialUpdate::default());
}
let mut update = ProviderCredentialUpdate::default();
let mut values_requiring_new_handles = HashMap::new();
for (credential_key, value) in updated_values {
match existing_handles.get(credential_key) {
Some(existing_handle) if credentials.storage_owns_handle(existing_handle) => {
update
.deferred_store_values
.insert(credential_key.clone(), value.clone());
}
Some(replaced_handle) => {
values_requiring_new_handles.insert(credential_key.clone(), value.clone());
update
.replaced_handles
.insert(credential_key.clone(), replaced_handle.clone());
}
None => {
values_requiring_new_handles.insert(credential_key.clone(), value.clone());
}
}
}
if !values_requiring_new_handles.is_empty() {
update.pre_stored_handles = credentials
.store_provider_credentials(
provider_name,
workspace,
provider_id,
&values_requiring_new_handles,
&HashMap::new(),
let object_id = uuid::Uuid::new_v4().to_string();
let pre_stored_handles = credentials
.store_provider_credentials_with_object_id(
provider_name,
workspace,
provider_id,
&object_id,
updated_values,
&HashMap::new(),
)
.await
.map_err(|err| {
Status::new(
err.code(),
"credential storage failed before provider publication",
)
.await?;
}
})?;
let replaced_handles = updated_values
.keys()
.filter_map(|key| {
existing_handles
.get(key)
.map(|handle| (key.clone(), handle.clone()))
})
.collect();
Ok(update)
Ok(ProviderCredentialUpdate {
pre_stored_handles,
replaced_handles,
})
}
// Retire only handles replaced by the successful provider CAS. Readers of an
// older record can fail resolution after retirement, but cannot resolve its
// handles to credential values from a different provider resource version.
async fn finish_provider_credential_update(
credentials: Option<&crate::credentials::CredentialRuntime>,
provider_name: &str,
@@ -847,7 +855,6 @@ async fn finish_provider_credential_update(
provider_id: &str,
update: ProviderCredentialUpdate,
removed_handles: &HashMap<String, CredentialHandle>,
existing_handles: &HashMap<String, CredentialHandle>,
) -> Result<(), Status> {
let Some(credentials) = credentials else {
return Ok(());
@@ -856,18 +863,6 @@ async fn finish_provider_credential_update(
return Ok(());
}
if !update.deferred_store_values.is_empty() {
credentials
.store_provider_credentials(
provider_name,
workspace,
provider_id,
&update.deferred_store_values,
existing_handles,
)
.await?;
}
let mut handles_to_delete = removed_handles.clone();
handles_to_delete.extend(update.replaced_handles);
if !handles_to_delete.is_empty() {
@@ -878,17 +873,21 @@ async fn finish_provider_credential_update(
provider_id,
&handles_to_delete,
)
.await?;
.await
.map_err(|err| {
Status::new(
err.code(),
"credential retirement failed after provider publication",
)
})?;
}
Ok(())
}
// TODO(credential-drivers): A gateway crash between CAS success and
// finish_provider_credential_update leaves replaced/removed credential handles
// orphaned in the backing store. This best-effort cleanup only covers pre-CAS
// failures. A background reconciliation loop should be added to detect and
// reclaim orphaned handles.
// A failed CAS never owns the published handles. Best-effort cleanup removes
// only this candidate's staged objects; a crash may leave unused backend objects
// but cannot change the credential data named by the committed provider record.
async fn cleanup_pre_stored_provider_credentials(
credentials: Option<&crate::credentials::CredentialRuntime>,
provider_name: &str,
@@ -908,7 +907,7 @@ async fn cleanup_pre_stored_provider_credentials(
{
warn!(
provider_name = %provider_name,
error = %err,
code = ?err.code(),
"failed to clean up staged provider credentials after provider update failure"
);
}
@@ -1128,6 +1127,7 @@ pub(super) async fn resolve_provider_environment_from_records_with_policy_bindin
let mut expires = HashMap::new();
let mut static_credential_bindings = HashMap::new();
let mut static_credential_keys = HashSet::new();
let mut readiness_reason = openshell_core::proto::ProviderReadinessReason::Unspecified;
let now_ms = crate::persistence::current_time_ms();
validate_provider_environment_records_unique_at(store, catalog, records, now_ms).await?;
let registry = openshell_providers::ProviderRegistry::new();
@@ -1209,6 +1209,8 @@ pub(super) async fn resolve_provider_environment_from_records_with_policy_bindin
key = %key,
"withholding provider credential not declared by resolved profile"
);
readiness_reason =
openshell_core::proto::ProviderReadinessReason::CredentialsWithheld;
continue;
}
if is_non_injectable_provider_credential(provider, key)
@@ -1232,6 +1234,8 @@ pub(super) async fn resolve_provider_environment_from_records_with_policy_bindin
key = %key,
"withholding static provider credential from endpointless profile"
);
readiness_reason =
openshell_core::proto::ProviderReadinessReason::CredentialsWithheld;
continue;
}
let expires_at_ms = provider
@@ -1248,6 +1252,8 @@ pub(super) async fn resolve_provider_environment_from_records_with_policy_bindin
expires_at_ms,
"skipping expired provider credential"
);
readiness_reason =
openshell_core::proto::ProviderReadinessReason::CredentialExpired;
continue;
}
expires.entry(key.clone()).or_insert(expires_at_ms);
@@ -1283,6 +1289,15 @@ pub(super) async fn resolve_provider_environment_from_records_with_policy_bindin
let resolved_refs = credentials
.resolve_provider_handles(provider, now_ms)
.await?;
// Expired handles are removed by the credential runtime before values
// reach this loop. Preserve omission evidence without exposing handles.
if provider.credential_handles.keys().any(|key| {
!is_non_injectable_provider_credential(provider, key)
&& !broker_only_credential_keys.contains(key)
&& !resolved_refs.values.contains_key(key)
}) {
readiness_reason = openshell_core::proto::ProviderReadinessReason::CredentialExpired;
}
for (key, value) in resolved_refs.values {
if accepted_stored_credential_keys
.as_ref()
@@ -1293,6 +1308,8 @@ pub(super) async fn resolve_provider_environment_from_records_with_policy_bindin
key = %key,
"withholding provider credential handle not declared by resolved profile"
);
readiness_reason =
openshell_core::proto::ProviderReadinessReason::CredentialsWithheld;
continue;
}
if is_non_injectable_provider_credential(provider, &key)
@@ -1312,6 +1329,8 @@ pub(super) async fn resolve_provider_environment_from_records_with_policy_bindin
key = %key,
"withholding static provider credential handle from endpointless profile"
);
readiness_reason =
openshell_core::proto::ProviderReadinessReason::CredentialsWithheld;
continue;
}
if let Some(expires_at_ms) = resolved_refs.expires_at_ms.get(&key).copied() {
@@ -1356,6 +1375,7 @@ pub(super) async fn resolve_provider_environment_from_records_with_policy_bindin
}
Ok(ProviderEnvironment {
readiness_reason,
environment: env,
credential_expiration_times: expires,
dynamic_credentials: resolve_dynamic_credentials_from_records(catalog, records),
@@ -2503,6 +2523,7 @@ pub(super) async fn handle_create_provider(
);
Ok(Response::new(ProviderResponse {
provider: Some(provider),
..Default::default()
}))
}
Err(err) => {
@@ -2537,6 +2558,7 @@ pub(super) async fn handle_get_provider(
Ok(Response::new(ProviderResponse {
provider: Some(provider),
..Default::default()
}))
}
@@ -3755,6 +3777,13 @@ pub(super) async fn handle_update_provider(
if state.credentials.stores_provider_credentials() && !provider.credentials.is_empty() {
state.compute.ensure_workspace(&workspace).await?;
}
// Freeze this operation's target identities before updating authority.
// Attachments made later are separate operations; frozen attachment epochs
// prevent an intervening detach/reattach from satisfying an older receipt.
let targets =
sandboxes_using_provider_records(state.store.as_ref(), &workspace, provider.object_name())
.await?;
let mutation_id = uuid::Uuid::new_v4().to_string();
let catalog = state
.provider_profile_sources
.snapshot_catalog(state.store.as_ref(), &workspace)
@@ -3770,6 +3799,24 @@ pub(super) async fn handle_update_provider(
.await;
match result {
Ok(provider) => {
let provider_version = provider
.metadata
.as_ref()
.map_or(0, |metadata| metadata.resource_version);
let mut target_receipts = Vec::with_capacity(targets.len());
for sandbox in &targets {
target_receipts.push(
super::provider_readiness::record_provider_mutation(
state,
sandbox,
provider.object_name(),
openshell_core::proto::ProviderMutationKind::Update,
Some((provider.object_id(), provider_version)),
&mutation_id,
)
.await?,
);
}
emit_provider_lifecycle(
&provider.r#type,
LifecycleOperation::Update,
@@ -3777,6 +3824,8 @@ pub(super) async fn handle_update_provider(
);
Ok(Response::new(ProviderResponse {
provider: Some(provider),
target_receipts,
mutation_id,
}))
}
Err(err) => {
@@ -8636,7 +8685,7 @@ mod tests {
}
#[tokio::test]
async fn update_provider_record_overwrites_credentials_with_runtime() {
async fn update_provider_credential_publication_waits_for_staged_storage() {
let store = test_store().await;
let config = openshell_core::Config::new(None).with_credential_drivers(["test-static"]);
let credentials = crate::credentials::CredentialRuntime::from_config(&config).unwrap();
@@ -8666,16 +8715,40 @@ mod tests {
.handle
.clone();
let updated = update_provider_record_validating(
let (store_hit, release_store) = credentials.gate_next_store();
let update = update_provider_record_validating(
&store,
"default",
&catalog,
provider_with_credential_value("openai-local", "openai", "OPENAI_API_KEY", "sk-second"),
&[],
Some(&credentials),
)
.await
.unwrap();
);
let inspect_while_storage_pending = async {
store_hit.await.unwrap();
let published = store
.get_message_by_name::<Provider>("default", "openai-local")
.await
.unwrap()
.unwrap();
assert_eq!(published, stored_first);
let resolved = resolve_provider_environment_with_credentials(
&store,
&catalog,
"default",
&["openai-local".to_string()],
&credentials,
)
.await
.unwrap();
assert_eq!(
resolved.get("OPENAI_API_KEY").map(String::as_str),
Some("sk-first")
);
release_store.send(()).unwrap();
};
let (updated, ()) = tokio::join!(update, inspect_while_storage_pending);
let updated = updated.unwrap();
assert_eq!(
updated
.credentials
@@ -8691,13 +8764,18 @@ mod tests {
.unwrap()
.unwrap();
assert!(stored_second.credentials.is_empty());
assert_eq!(
assert_ne!(
stored_second
.credential_handles
.get("OPENAI_API_KEY")
.map(|handle| handle.handle.as_str()),
Some(first_handle.as_str())
);
assert_eq!(
stored_second.metadata.as_ref().unwrap().resource_version,
stored_first.metadata.as_ref().unwrap().resource_version + 1
);
assert_eq!(credentials.stored_credential_count(), Some(1));
let result = resolve_provider_environment_with_credentials(
&store,
@@ -8711,6 +8789,376 @@ mod tests {
assert_eq!(result.get("OPENAI_API_KEY"), Some(&"sk-second".to_string()));
}
#[tokio::test]
async fn update_provider_credential_store_failure_preserves_published_revision() {
let store = test_store().await;
let config = openshell_core::Config::new(None).with_credential_drivers(["test-static"]);
let credentials = crate::credentials::CredentialRuntime::from_config(&config).unwrap();
let catalog = ProviderProfileSources::with_default_sources()
.snapshot_catalog(&store, "default")
.await
.unwrap();
create_provider_record_validating(
&store,
"default",
&catalog,
provider_with_credential_value("openai-local", "openai", "OPENAI_API_KEY", "sk-first"),
Some(&credentials),
)
.await
.unwrap();
let before = store
.get_message_by_name::<Provider>("default", "openai-local")
.await
.unwrap()
.unwrap();
credentials.fail_next_store();
let error = update_provider_record_validating(
&store,
"default",
&catalog,
provider_with_credential_value("openai-local", "openai", "OPENAI_API_KEY", "sk-failed"),
&[],
Some(&credentials),
)
.await
.unwrap_err();
assert_eq!(error.code(), Code::Unavailable);
assert_eq!(
error.message(),
"credential storage failed before provider publication"
);
let after = store
.get_message_by_name::<Provider>("default", "openai-local")
.await
.unwrap()
.unwrap();
assert_eq!(after, before);
assert_eq!(credentials.stored_credential_count(), Some(1));
let resolved = resolve_provider_environment_with_credentials(
&store,
&catalog,
"default",
&["openai-local".to_string()],
&credentials,
)
.await
.unwrap();
assert_eq!(
resolved.get("OPENAI_API_KEY").map(String::as_str),
Some("sk-first")
);
}
#[tokio::test]
async fn update_provider_credential_publication_cas_loser_preserves_winner() {
let store = test_store().await;
let config = openshell_core::Config::new(None).with_credential_drivers(["test-static"]);
let credentials = crate::credentials::CredentialRuntime::from_config(&config).unwrap();
let catalog = ProviderProfileSources::with_default_sources()
.snapshot_catalog(&store, "default")
.await
.unwrap();
create_provider_record_validating(
&store,
"default",
&catalog,
provider_with_credential_value("openai-local", "openai", "OPENAI_API_KEY", "sk-first"),
Some(&credentials),
)
.await
.unwrap();
let initial = store
.get_message_by_name::<Provider>("default", "openai-local")
.await
.unwrap()
.unwrap();
let (store_hit, release_store) = credentials.gate_next_store();
// Pause one writer after it reads the provider version. A second writer
// publishes while it waits, so only a database CAS can reject the loser.
let loser = update_provider_record_validating(
&store,
"default",
&catalog,
provider_with_credential_value("openai-local", "openai", "OPENAI_API_KEY", "sk-loser"),
&[],
Some(&credentials),
);
let winner = async {
store_hit.await.unwrap();
let result = update_provider_record_validating(
&store,
"default",
&catalog,
provider_with_credential_value(
"openai-local",
"openai",
"OPENAI_API_KEY",
"sk-winner",
),
&[],
Some(&credentials),
)
.await;
release_store.send(()).unwrap();
result
};
let (loser_result, winner_result) = tokio::join!(loser, winner);
assert_eq!(loser_result.unwrap_err().code(), Code::Aborted);
winner_result.unwrap();
let published = store
.get_message_by_name::<Provider>("default", "openai-local")
.await
.unwrap()
.unwrap();
assert_eq!(
published.metadata.as_ref().unwrap().resource_version,
initial.metadata.as_ref().unwrap().resource_version + 1
);
assert_eq!(credentials.stored_credential_count(), Some(1));
let resolved = resolve_provider_environment_with_credentials(
&store,
&catalog,
"default",
&["openai-local".to_string()],
&credentials,
)
.await
.unwrap();
assert_eq!(
resolved.get("OPENAI_API_KEY").map(String::as_str),
Some("sk-winner")
);
}
#[tokio::test]
async fn update_provider_receipts_freeze_targets_before_credential_publication() {
let state = test_server_state().await;
handle_create_provider(
&state,
authed_request(CreateProviderRequest {
request_id: String::new(),
provider: Some(provider_with_credential_value(
"openai-local",
"openai",
"OPENAI_API_KEY",
"sk-first",
)),
workspace_scope: Some(openshell_core::proto::workspace_selector("default")),
}),
)
.await
.unwrap();
for name in ["first", "second", "late"] {
state
.store
.put_message(&Sandbox {
metadata: Some(openshell_core::proto::datamodel::v1::ObjectMeta {
id: name.to_string(),
name: name.to_string(),
workspace: "default".to_string(),
..Default::default()
}),
spec: Some(SandboxSpec {
providers: if name == "late" {
Vec::new()
} else {
vec!["openai-local".to_string()]
},
provider_attachment_epoch: format!("epoch-{name}"),
..Default::default()
}),
..Default::default()
})
.await
.unwrap();
}
let (store_hit, release_store) = state.credentials.gate_next_store();
let update = handle_update_provider(
&state,
authed_request(UpdateProviderRequest {
request_id: String::new(),
provider: Some(provider_with_credential_value(
"openai-local",
"openai",
"OPENAI_API_KEY",
"sk-second",
)),
credential_expiration_times: HashMap::new(),
clear_credential_expiration_keys: Vec::new(),
workspace_scope: Some(openshell_core::proto::workspace_selector("default")),
}),
);
let attach_from_another_replica = async {
store_hit.await.unwrap();
// Another gateway's attachment is outside the captured target set,
// even when its write reaches the database before provider publish.
state
.store
.update_message_cas::<Sandbox, _>("late", 0, |sandbox| {
let spec = sandbox.spec.as_mut().unwrap();
spec.providers.push("openai-local".to_string());
spec.provider_attachment_epoch = "epoch-late-attached".to_string();
})
.await
.unwrap();
release_store.send(()).unwrap();
};
let (response, ()) = tokio::join!(update, attach_from_another_replica);
let response = response.unwrap().into_inner();
let provider = response.provider.as_ref().unwrap();
let targets: HashSet<_> = response
.target_receipts
.iter()
.map(|receipt| receipt.desired.as_ref().unwrap().sandbox_name.as_str())
.collect();
assert_eq!(targets, HashSet::from(["first", "second"]));
assert!(!response.mutation_id.is_empty());
for receipt in &response.target_receipts {
assert_eq!(receipt.mutation_id, response.mutation_id);
assert_eq!(
receipt.kind,
openshell_core::proto::ProviderMutationKind::Update as i32
);
let desired = receipt.desired.as_ref().unwrap();
assert_eq!(desired.provider_id, provider.object_id());
assert_eq!(
desired.provider_resource_version,
provider.metadata.as_ref().unwrap().resource_version
);
assert_eq!(
desired.attachment_epoch,
format!("epoch-{}", desired.sandbox_name)
);
}
}
#[tokio::test]
async fn update_provider_retirement_failure_preserves_publication_and_receipts() {
let state = test_server_state().await;
handle_create_provider(
&state,
authed_request(CreateProviderRequest {
request_id: String::new(),
provider: Some(provider_with_credential_value(
"openai-local",
"openai",
"OPENAI_API_KEY",
"sk-first",
)),
workspace_scope: Some(openshell_core::proto::workspace_selector("default")),
}),
)
.await
.unwrap();
let initial = state
.store
.get_message_by_name::<Provider>("default", "openai-local")
.await
.unwrap()
.unwrap();
let initial_version = initial.metadata.as_ref().unwrap().resource_version;
let mut target_epochs = HashMap::new();
for name in ["s1", "s2"] {
let attachment_epoch = uuid::Uuid::new_v4().to_string();
state
.store
.put_message(&Sandbox {
metadata: Some(openshell_core::proto::datamodel::v1::ObjectMeta {
id: uuid::Uuid::new_v4().to_string(),
name: name.to_string(),
workspace: "default".to_string(),
..Default::default()
}),
spec: Some(SandboxSpec {
providers: vec!["openai-local".to_string()],
provider_attachment_epoch: attachment_epoch.clone(),
policy: Some(openshell_policy::restrictive_default_policy()),
..Default::default()
}),
..Default::default()
})
.await
.unwrap();
target_epochs.insert(name, attachment_epoch);
}
// Only old-object retirement fails: staging and the provider CAS must
// still publish the replacement and preserve every selected receipt.
state.credentials.fail_next_delete();
let result = handle_update_provider(
&state,
authed_request(UpdateProviderRequest {
request_id: String::new(),
provider: Some(provider_with_credential_value(
"openai-local",
"openai",
"OPENAI_API_KEY",
"sk-second",
)),
credential_expiration_times: HashMap::new(),
clear_credential_expiration_keys: Vec::new(),
workspace_scope: Some(openshell_core::proto::workspace_selector("default")),
}),
)
.await;
let published = state
.store
.get_message_by_name::<Provider>("default", "openai-local")
.await
.unwrap()
.unwrap();
let published_version = published.metadata.as_ref().unwrap().resource_version;
assert_eq!(published_version, initial_version + 1);
assert_ne!(published.credential_handles, initial.credential_handles);
let resolved = state
.credentials
.resolve_provider_handles(&published, crate::persistence::current_time_ms())
.await
.unwrap();
assert_eq!(resolved.values["OPENAI_API_KEY"], "sk-second");
assert_eq!(state.credentials.stored_credential_count(), Some(2));
let response = result
.expect("published credential rotation must return its mutation receipts")
.into_inner();
let provider = response.provider.as_ref().unwrap();
assert_eq!(provider.object_id(), published.object_id());
assert_eq!(
provider.metadata.as_ref().unwrap().resource_version,
published_version
);
assert!(!response.mutation_id.is_empty());
let targets: HashSet<_> = response
.target_receipts
.iter()
.map(|receipt| receipt.desired.as_ref().unwrap().sandbox_name.as_str())
.collect();
assert_eq!(targets, HashSet::from(["s1", "s2"]));
assert_eq!(response.target_receipts.len(), target_epochs.len());
for receipt in &response.target_receipts {
assert!(!receipt.receipt_id.is_empty());
assert_eq!(receipt.mutation_id, response.mutation_id);
assert_eq!(
receipt.kind,
openshell_core::proto::ProviderMutationKind::Update as i32
);
let desired = receipt.desired.as_ref().unwrap();
assert_eq!(desired.provider_id, published.object_id());
assert_eq!(desired.provider_resource_version, published_version);
assert_eq!(
desired.attachment_epoch,
target_epochs[desired.sandbox_name.as_str()]
);
assert!(!desired.policy_hash.is_empty());
}
}
#[tokio::test]
async fn update_provider_record_clears_credential_expiration_by_key() {
let store = test_store().await;
@@ -11024,6 +11472,10 @@ mod tests {
assert_eq!(result.get("FRESH_TOKEN"), Some(&"fresh".to_string()));
assert!(!result.contains_key("STALE_TOKEN"));
assert!(!result.contains_key("EPOCH_TOKEN"));
assert_eq!(
result.readiness_reason,
openshell_core::proto::ProviderReadinessReason::CredentialExpired
);
assert_eq!(
result.credential_expiration_times.get("FRESH_TOKEN"),
Some(&(now_ms + 60_000))
@@ -0,0 +1,628 @@
// SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
// SPDX-License-Identifier: Apache-2.0
//! Provider views of common configuration operations and session-bound evidence.
#![allow(clippy::result_large_err)] // The RPC boundary returns tonic status values.
use std::sync::Arc;
use std::time::{Duration, Instant, SystemTime};
use openshell_core::proto::{
GetSandboxProviderStatusRequest, GetSandboxProviderStatusResponse, Provider,
ProviderDesiredIdentity, ProviderMutationKind, ProviderMutationReceipt,
ProviderReadinessObservation, ProviderReadinessReason, ProviderReadinessState,
ProviderReadinessStatus, ReportProviderReadinessRequest, ReportProviderReadinessResponse,
Sandbox, SandboxPhase, SupervisorHello,
};
use openshell_core::{ObjectId, ObjectName, ObjectWorkspace};
use tonic::{Request, Response, Status};
use uuid::Uuid;
use crate::ServerState;
use crate::auth::guard::{enforce_sandbox_scope, ensure_sandbox_principal_scope};
use crate::auth::workspace_authz::{MinWorkspaceRole, authorize_workspace_selector};
use crate::config_update_operation;
use crate::persistence::ObjectType;
const REPORT_INTERVAL_SECONDS: u32 = 5;
const OBSERVATION_TTL_SECONDS: u32 = 15;
/// Installation evidence owned by one live `ConnectSupervisor` session.
/// Replacing or removing that session discards this evidence with it.
#[derive(Clone, Debug)]
pub struct ProviderReadinessEvidence {
network_instance_id: String,
supported: bool,
process_instance_id: Option<String>,
last_seen: Instant,
observation: Option<ProviderReadinessObservation>,
observed_time: Option<prost_types::Timestamp>,
}
impl ProviderReadinessEvidence {
/// Require the instance last accepted into the persisted sandbox lifecycle.
/// A local session alone cannot prove ownership across gateway replicas.
pub(crate) fn belongs_to_instance(&self, active_instance_id: &str) -> bool {
!active_instance_id.is_empty() && self.network_instance_id == active_instance_id
}
/// Capture capability and instance identity from the authenticated hello.
pub(crate) fn from_hello(hello: &SupervisorHello) -> Result<Self, Status> {
if hello.supports_provider_readiness {
canonical_uuid(&hello.sandbox_id)?;
canonical_uuid(&hello.instance_id)?;
}
Ok(Self {
network_instance_id: hello.instance_id.clone(),
supported: hello.supports_provider_readiness,
process_instance_id: None,
last_seen: Instant::now(),
observation: None,
observed_time: None,
})
}
/// Accept an ordered report after the registry verifies session ownership.
/// Retrying an identical report never extends its original acceptance time.
pub(crate) fn accept(
&mut self,
observation: ProviderReadinessObservation,
) -> Result<(), Status> {
if !self.supported {
return Err(Status::failed_precondition(
"supervisor did not advertise provider readiness support",
));
}
if observation.sequence == 0 {
return Err(Status::invalid_argument("report sequence is required"));
}
if let Some(last) = self.observation.as_ref() {
if observation.sequence == last.sequence {
return if &observation == last {
Ok(())
} else {
Err(Status::invalid_argument(
"a provider report sequence must reuse the identical observation",
))
};
}
if observation.sequence < last.sequence {
return Err(Status::failed_precondition(
"provider readiness observation is out of order",
));
}
}
if observation.session_id.is_empty() {
return Err(Status::failed_precondition(
"supervisor session identity is required",
));
}
ProviderReadinessReason::try_from(observation.reason)
.map_err(|_| Status::invalid_argument("unknown readiness reason"))?;
if !observation.attachment_epoch.is_empty() {
canonical_uuid(&observation.attachment_epoch)?;
}
if observation.policy_hash.len() > 128
|| !observation
.policy_hash
.bytes()
.all(|byte| byte.is_ascii_hexdigit())
{
return Err(Status::invalid_argument("invalid policy fingerprint"));
}
if observation.launch_environment_installed && observation.process_instance_id.is_empty() {
return Err(Status::invalid_argument(
"process instance identity is required",
));
}
let observed_time = now_timestamp()?;
if !observation.process_instance_id.is_empty() {
canonical_uuid(&observation.process_instance_id)?;
if self
.process_instance_id
.as_ref()
.is_some_and(|process_id| process_id != &observation.process_instance_id)
{
return Err(Status::failed_precondition("process instance changed"));
}
// The supervisor reports its authenticated boundary's process
// identity. A boundary replacement requires a new control session.
self.process_instance_id = Some(observation.process_instance_id.clone());
}
self.last_seen = Instant::now();
self.observed_time = Some(observed_time);
self.observation = Some(observation);
Ok(())
}
}
fn registry_unavailable() -> Status {
Status::unavailable("provider readiness state is unavailable")
}
fn now_timestamp() -> Result<prost_types::Timestamp, Status> {
openshell_core::time::timestamp_from_system_time(SystemTime::now())
.map_err(|error| Status::internal(format!("create provider readiness timestamp: {error}")))
}
fn canonical_uuid(value: &str) -> Result<(), Status> {
let parsed = Uuid::parse_str(value)
.map_err(|_| Status::invalid_argument("invalid readiness instance identity"))?;
if parsed.is_nil() || parsed.to_string() != value {
return Err(Status::invalid_argument(
"invalid readiness instance identity",
));
}
Ok(())
}
/// Persist an immutable receipt for the target frozen before a source mutation.
/// Expected provider identity prevents a competing update from acquiring this
/// receipt even when snapshot reads happen after the source CAS.
pub(super) async fn record_provider_mutation(
state: &Arc<ServerState>,
sandbox: &Sandbox,
provider_name: &str,
kind: ProviderMutationKind,
expected_provider: Option<(&str, u64)>,
mutation_id: &str,
) -> Result<ProviderMutationReceipt, Status> {
let mut desired = target_identity(sandbox, expected_provider);
let snapshot_reason = match load_current_snapshot(state, sandbox, provider_name).await {
Ok((snapshot, _)) if same_authority(&desired, &snapshot) => {
desired = snapshot;
ProviderReadinessReason::Unspecified
}
Ok(_) => ProviderReadinessReason::SnapshotMismatch,
Err(_) => ProviderReadinessReason::CredentialsWithheld,
};
let receipt = ProviderMutationReceipt {
receipt_id: Uuid::new_v4().to_string(),
mutation_id: mutation_id.to_string(),
provider_name: provider_name.to_string(),
workspace: sandbox.object_workspace().to_string(),
kind: kind.into(),
desired: Some(desired),
persisted_time: Some(now_timestamp()?),
};
config_update_operation::record_provider_operation(
state.store.as_ref(),
receipt,
snapshot_reason,
)
.await
}
fn target_identity(sandbox: &Sandbox, provider: Option<(&str, u64)>) -> ProviderDesiredIdentity {
// The empty initial epoch remains an opaque identity paired with the
// sandbox UUID. Every provider-set mutation atomically replaces the epoch.
ProviderDesiredIdentity {
sandbox_id: sandbox.object_id().to_string(),
sandbox_name: sandbox.object_name().to_string(),
attachment_epoch: sandbox
.spec
.as_ref()
.map(|spec| spec.provider_attachment_epoch.clone())
.unwrap_or_default(),
provider_id: provider.map_or_else(String::new, |(id, _)| id.to_string()),
provider_resource_version: provider.map_or(0, |(_, version)| version),
..Default::default()
}
}
fn same_authority(left: &ProviderDesiredIdentity, right: &ProviderDesiredIdentity) -> bool {
left.sandbox_id == right.sandbox_id
&& left.sandbox_name == right.sandbox_name
&& left.attachment_epoch == right.attachment_epoch
&& left.provider_id == right.provider_id
&& left.provider_resource_version == right.provider_resource_version
}
async fn current_target_identity(
state: &Arc<ServerState>,
sandbox: &Sandbox,
provider_name: &str,
) -> Result<ProviderDesiredIdentity, Status> {
let attached = sandbox
.spec
.as_ref()
.is_some_and(|spec| spec.providers.iter().any(|name| name == provider_name));
let provider = if attached {
state
.store
.get_by_name(
Provider::object_type(),
sandbox.object_workspace(),
provider_name,
)
.await
.map_err(|_| registry_unavailable())?
} else {
None
};
Ok(target_identity(
sandbox,
provider
.as_ref()
.map(|provider| (provider.id.as_str(), provider.resource_version)),
))
}
/// Bracket independently stored config and credential records with exact
/// identity reads. Inconsistent reads remain pending; revisions are never
/// treated as counters or compared with greater-than ordering.
async fn load_current_snapshot(
state: &Arc<ServerState>,
sandbox: &Sandbox,
provider_name: &str,
) -> Result<(ProviderDesiredIdentity, ProviderReadinessReason), Status> {
let mut desired = current_target_identity(state, sandbox, provider_name).await?;
let config = super::policy::load_sandbox_config(state, sandbox).await?;
let environment =
super::policy::load_sandbox_provider_environment(state, sandbox, true).await?;
let current = state
.store
.get_message::<Sandbox>(sandbox.object_id())
.await
.map_err(|_| registry_unavailable())?
.ok_or_else(|| Status::not_found("sandbox not found"))?;
let after = current_target_identity(state, &current, provider_name).await?;
let final_config = super::policy::load_sandbox_config(state, &current).await?;
if !same_authority(&desired, &after)
|| config.provider_attachment_epoch != environment.provider_attachment_epoch
|| config.provider_env_revision != environment.provider_env_revision
|| (config.policy.is_some() && config.policy_hash != environment.policy_hash)
|| config.config_revision != final_config.config_revision
|| config.provider_env_revision != final_config.provider_env_revision
|| config.policy_hash != final_config.policy_hash
{
return Err(Status::aborted("provider readiness snapshot changed"));
}
desired.provider_env_revision = config.provider_env_revision;
desired.config_revision = config.config_revision;
desired.policy_hash = config.policy_hash;
let reason = if config.policy.is_none() {
ProviderReadinessReason::LocalPolicy
} else {
ProviderReadinessReason::try_from(environment.readiness_reason)
.unwrap_or(ProviderReadinessReason::CredentialsWithheld)
};
Ok((desired, reason))
}
/// Return a redacted operator view. Receipt-less lookup reconstructs current
/// intent only; it never claims that a mutation lost before persistence finished.
pub(super) async fn handle_get_sandbox_provider_status(
state: &Arc<ServerState>,
request: Request<GetSandboxProviderStatusRequest>,
) -> Result<Response<GetSandboxProviderStatusResponse>, Status> {
let principal = super::extract_principal(&request)?;
let request = request.into_inner();
let authz = authorize_workspace_selector(
&state.store,
&state.admin_role,
&principal,
request.workspace_scope.as_ref(),
MinWorkspaceRole::User,
)
.await?;
let workspace = super::workspace::resolve_workspace(state.store.as_ref(), &authz.workspace)
.await?
.name;
// Validate the selector before it can contribute to a durable observation.
if request.provider_name.len() > super::MAX_NAME_LEN {
return Err(Status::invalid_argument(
"provider_name exceeds maximum length",
));
}
let sandbox = state
.store
.get_message_by_name::<Sandbox>(&workspace, &request.sandbox_name)
.await
.map_err(|_| registry_unavailable())?
.ok_or_else(|| Status::not_found("sandbox not found"))?;
let receipt_id = if request.receipt_id.is_empty() {
if request.provider_name.is_empty() {
return Err(Status::invalid_argument(
"provider_name or receipt_id is required",
));
}
// Only existing workspace providers can create observation operations.
// A detached provider still has a useful absent-reference target;
// historical receipts remain queryable after the provider is deleted.
state
.store
.get_by_name(Provider::object_type(), &workspace, &request.provider_name)
.await
.map_err(|_| registry_unavailable())?
.ok_or_else(|| Status::not_found("provider not found"))?;
let current = current_target_identity(state, &sandbox, &request.provider_name).await?;
let provider = (!current.provider_id.is_empty()).then_some((
current.provider_id.as_str(),
current.provider_resource_version,
));
record_provider_mutation(
state,
&sandbox,
&request.provider_name,
ProviderMutationKind::Observe,
provider,
&Uuid::new_v4().to_string(),
)
.await?
.receipt_id
} else {
canonical_uuid(&request.receipt_id)?;
request.receipt_id
};
let stored = config_update_operation::get_provider_operation(
state.store.as_ref(),
&receipt_id,
&workspace,
)
.await?;
let receipt = stored.receipt;
if receipt.workspace != workspace
|| receipt
.desired
.as_ref()
.is_none_or(|desired| desired.sandbox_id != sandbox.object_id())
|| (!request.provider_name.is_empty() && receipt.provider_name != request.provider_name)
{
return Err(Status::not_found("provider receipt not found"));
}
let current_authority =
current_target_identity(state, &sandbox, &receipt.provider_name).await?;
let (current, current_reason) = if let Ok(snapshot) =
load_current_snapshot(state, &sandbox, &receipt.provider_name).await
{
snapshot
} else {
// A temporarily unreadable config cannot prove supersession or
// completion. Metadata changes remain comparable independently.
let mut fallback = current_authority;
if let Some(desired) = receipt.desired.as_ref() {
fallback.provider_env_revision = desired.provider_env_revision;
fallback.config_revision = desired.config_revision;
fallback.policy_hash.clone_from(&desired.policy_hash);
}
(fallback, ProviderReadinessReason::CredentialsWithheld)
};
// Refresh lifecycle identity after configuration reads. Another gateway
// replica may have accepted a new supervisor while this replica retained
// the predecessor's local connection and installation evidence.
let active_sandbox = state
.store
.get_message::<Sandbox>(sandbox.object_id())
.await
.map_err(|_| registry_unavailable())?
.ok_or_else(|| Status::not_found("sandbox not found"))?;
let active_instance_id = active_sandbox
.status
.as_ref()
.map_or("", |status| status.main_process_instance_id.as_str());
let session = state
.supervisor_sessions
.provider_readiness(sandbox.object_id())?;
let mut status = evaluate_status(
receipt,
stored.snapshot_reason,
&current,
current_reason,
active_sandbox.phase() == SandboxPhase::Ready as i32,
active_instance_id,
session.as_ref(),
)?;
config_update_operation::observe_provider_status(state.store.as_ref(), &mut status).await?;
Ok(Response::new(GetSandboxProviderStatusResponse {
status: Some(status),
}))
}
fn observation_matches(
observation: &ProviderReadinessObservation,
desired: &ProviderDesiredIdentity,
) -> bool {
observation.attachment_epoch == desired.attachment_epoch
&& observation.provider_env_revision == desired.provider_env_revision
&& observation.config_revision == desired.config_revision
&& observation.policy_hash == desired.policy_hash
}
fn failure_state(reason: ProviderReadinessReason) -> ProviderReadinessState {
match reason {
ProviderReadinessReason::CredentialsWithheld
| ProviderReadinessReason::CredentialExpired => ProviderReadinessState::Withheld,
ProviderReadinessReason::CredentialInstallFailed
| ProviderReadinessReason::PolicyActivationFailed
| ProviderReadinessReason::ProcessInstallFailed
| ProviderReadinessReason::UnsupportedSupervisor
| ProviderReadinessReason::LocalPolicy => ProviderReadinessState::Failed,
_ => ProviderReadinessState::Pending,
}
}
fn evaluate_status(
receipt: ProviderMutationReceipt,
snapshot_reason: ProviderReadinessReason,
current: &ProviderDesiredIdentity,
current_reason: ProviderReadinessReason,
running: bool,
active_instance_id: &str,
session: Option<&ProviderReadinessEvidence>,
) -> Result<ProviderReadinessStatus, Status> {
let mut status = ProviderReadinessStatus {
receipt: Some(receipt.clone()),
state: ProviderReadinessState::Persisted.into(),
reason: ProviderReadinessReason::WaitingForSupervisor.into(),
observed: session.and_then(|session| session.observation.clone()),
network_instance_id: session
.map(|session| session.network_instance_id.clone())
.unwrap_or_default(),
observed_time: session.and_then(|session| session.observed_time),
evaluated_time: Some(now_timestamp()?),
operation: None,
};
let set = |status: &mut ProviderReadinessStatus,
state: ProviderReadinessState,
reason: ProviderReadinessReason| {
status.state = state.into();
status.reason = reason.into();
};
let Some(desired) = receipt.desired.as_ref() else {
set(
&mut status,
ProviderReadinessState::Failed,
ProviderReadinessReason::SnapshotMismatch,
);
return Ok(status);
};
if !same_authority(desired, current)
|| (snapshot_reason == ProviderReadinessReason::Unspecified && desired != current)
{
set(
&mut status,
ProviderReadinessState::Superseded,
ProviderReadinessReason::DesiredStateChanged,
);
} else if snapshot_reason != ProviderReadinessReason::Unspecified {
set(&mut status, ProviderReadinessState::Failed, snapshot_reason);
} else if current_reason != ProviderReadinessReason::Unspecified {
set(&mut status, failure_state(current_reason), current_reason);
} else if !running {
set(
&mut status,
ProviderReadinessState::Pending,
ProviderReadinessReason::SupervisorDisconnected,
);
} else if let Some(session) = session {
if !session.belongs_to_instance(active_instance_id) {
set(
&mut status,
ProviderReadinessState::Pending,
ProviderReadinessReason::SupervisorDisconnected,
);
} else if session.last_seen.elapsed()
>= Duration::from_secs(u64::from(OBSERVATION_TTL_SECONDS))
{
set(
&mut status,
ProviderReadinessState::Pending,
ProviderReadinessReason::SupervisorLeaseExpired,
);
} else if !session.supported {
set(
&mut status,
ProviderReadinessState::Failed,
ProviderReadinessReason::UnsupportedSupervisor,
);
} else if let Some(observation) = session.observation.as_ref() {
let reason = ProviderReadinessReason::try_from(observation.reason)
.unwrap_or(ProviderReadinessReason::SnapshotMismatch);
// Installation failures describe one snapshot. A retained failure
// from an older poll must not fail a newly persisted authority.
if !observation_matches(observation, desired) {
set(
&mut status,
ProviderReadinessState::Pending,
ProviderReadinessReason::SnapshotMismatch,
);
} else if reason != ProviderReadinessReason::Unspecified {
set(&mut status, failure_state(reason), reason);
} else if !observation.credentials_installed {
set(
&mut status,
ProviderReadinessState::Pending,
ProviderReadinessReason::WaitingForCredentials,
);
} else if !observation.policy_active {
set(
&mut status,
ProviderReadinessState::Pending,
ProviderReadinessReason::WaitingForPolicy,
);
} else if !observation.launch_environment_installed
|| observation.process_instance_id.is_empty()
{
set(
&mut status,
ProviderReadinessState::Pending,
ProviderReadinessReason::WaitingForProcess,
);
} else {
let state = if desired.provider_id.is_empty() {
ProviderReadinessState::Revoked
} else {
ProviderReadinessState::Ready
};
set(&mut status, state, ProviderReadinessReason::Unspecified);
}
}
}
Ok(status)
}
fn authorize_provider_readiness<T>(
request: &Request<T>,
sandbox_id: &str,
) -> Result<crate::auth::principal::Principal, Status> {
let principal = enforce_sandbox_scope(request, sandbox_id)?;
ensure_sandbox_principal_scope(&principal, sandbox_id)?;
Ok(principal)
}
/// Accept installation evidence only from the sandbox's current supervisor.
/// Authorization is independent from operator permission to inspect progress.
pub(super) async fn handle_report_provider_readiness(
state: &Arc<ServerState>,
request: Request<ReportProviderReadinessRequest>,
) -> Result<Response<ReportProviderReadinessResponse>, Status> {
let sandbox_id = request.get_ref().sandbox_id.clone();
let principal = authorize_provider_readiness(&request, &sandbox_id)?;
canonical_uuid(&sandbox_id)?;
let observation = request
.into_inner()
.observation
.ok_or_else(|| Status::invalid_argument("provider readiness observation is required"))?;
canonical_uuid(&observation.session_id)?;
// A valid session must still belong to an existing sandbox. The registry
// performs the final current-session check atomically with accepting evidence,
// so a reconnect during this read cannot publish into its replacement.
let sandbox =
super::sandbox::fetch_and_authorize_sandbox(state, &principal, &sandbox_id).await?;
if sandbox.phase() != SandboxPhase::Ready as i32 {
return Err(Status::failed_precondition(
"sandbox supervisor is not ready",
));
}
let active_instance_id = sandbox
.status
.as_ref()
.map_or("", |status| status.main_process_instance_id.as_str());
let accepted_sequence = observation.sequence;
state.supervisor_sessions.accept_provider_readiness(
&sandbox_id,
active_instance_id,
observation,
)?;
Ok(Response::new(ReportProviderReadinessResponse {
accepted_sequence,
report_interval: Some(
openshell_core::time::duration_from_std(Duration::from_secs(u64::from(
REPORT_INTERVAL_SECONDS,
)))
.map_err(|error| Status::internal(format!("create report interval: {error}")))?,
),
observation_ttl: Some(
openshell_core::time::duration_from_std(Duration::from_secs(u64::from(
OBSERVATION_TTL_SECONDS,
)))
.map_err(|error| Status::internal(format!("create observation TTL: {error}")))?,
),
}))
}
#[cfg(test)]
#[path = "provider_readiness_tests.rs"]
mod tests;
File diff suppressed because it is too large Load Diff
+136 -20
View File
@@ -29,12 +29,12 @@ use openshell_core::proto::{
ExecSandboxEvent, ExecSandboxExit, ExecSandboxInput, ExecSandboxRequest, ExecSandboxStderr,
ExecSandboxStdout, GetSandboxRequest, GetSandboxTemplateRequest, ListSandboxProvidersRequest,
ListSandboxProvidersResponse, ListSandboxTemplatesRequest, ListSandboxTemplatesResponse,
ListSandboxesRequest, ListSandboxesResponse, Provider, ResourceRequirements,
RevokeSshSessionRequest, RevokeSshSessionResponse, SandboxResources, SandboxResponse,
SandboxSpec, SandboxStreamEvent, SandboxTemplateResponse, SandboxWorkloadTemplate,
SandboxWorkloadTemplateProvenance, SshRelayTarget, StartSandboxRequest, StopSandboxRequest,
TcpForwardFrame, TcpForwardInit, TcpRelayTarget, WatchSandboxRequest, relay_open,
tcp_forward_init,
ListSandboxesRequest, ListSandboxesResponse, Provider, ProviderMutationKind,
ResourceRequirements, RevokeSshSessionRequest, RevokeSshSessionResponse, SandboxResources,
SandboxResponse, SandboxSpec, SandboxStreamEvent, SandboxTemplateResponse,
SandboxWorkloadTemplate, SandboxWorkloadTemplateProvenance, SshRelayTarget,
StartSandboxRequest, StopSandboxRequest, TcpForwardFrame, TcpForwardInit, TcpRelayTarget,
WatchSandboxRequest, relay_open, tcp_forward_init,
};
use openshell_core::proto::{
BeginRootfsTarStagingRequest, BeginRootfsTarStagingResponse, Sandbox, SandboxPhase,
@@ -172,6 +172,8 @@ pub(super) async fn handle_create_sandbox(
request: Request<CreateSandboxRequest>,
) -> Result<Response<SandboxResponse>, Status> {
let create_request = request.get_ref().clone();
// Sandbox creation retains large configuration values across awaits.
// Box the inner future to keep this wrapper small for every caller.
let result = Box::pin(handle_create_sandbox_inner(state, request)).await;
let created_sandbox = result
.as_ref()
@@ -373,6 +375,10 @@ async fn handle_create_sandbox_inner(
(resolved, Some(provenance))
};
// Attachment identity belongs to the gateway. Accepting an epoch from a
// create request or workload template could revive stale installation proof.
spec.provider_attachment_epoch = uuid::Uuid::new_v4().to_string();
// Leave an omitted command empty rather than persisting a concrete shell:
// the sandbox boundary resolves the default login shell against the agent image
// (bash when present, otherwise /bin/sh on minimal images like Alpine),
@@ -1067,6 +1073,11 @@ pub(super) async fn handle_attach_sandbox_provider(
request: Request<AttachSandboxProviderRequest>,
) -> Result<Response<AttachSandboxProviderResponse>, Status> {
let principal = super::extract_principal(&request)?;
#[cfg(test)]
let attach_wait_probe = request
.extensions()
.get::<Arc<tokio::sync::Notify>>()
.cloned();
let request = request.into_inner();
let authz = authorize_workspace_selector(
&state.store,
@@ -1093,20 +1104,27 @@ pub(super) async fn handle_attach_sandbox_provider(
)));
}
get_provider_record(state.store.as_ref(), &workspace, &request.provider_name)
.await
.map_err(|err| {
if err.code() == tonic::Code::NotFound {
Status::failed_precondition(format!(
"provider '{}' not found",
request.provider_name
))
} else {
err
}
})?;
// The receipt must capture the provider revision selected by this
// serialized mutation, after any preceding credential update has finished.
#[cfg(test)]
if let Some(probe) = attach_wait_probe {
probe.notify_one();
}
let _sandbox_sync_guard = state.compute.sandbox_sync_guard().await;
let provider_record =
get_provider_record(state.store.as_ref(), &workspace, &request.provider_name)
.await
.map_err(|err| {
if err.code() == tonic::Code::NotFound {
Status::failed_precondition(format!(
"provider '{}' not found",
request.provider_name
))
} else {
err
}
})?;
let sandbox = sandbox_by_name(state, &workspace, &request.sandbox_name).await?;
let sandbox_id = sandbox
.metadata
@@ -1172,6 +1190,7 @@ pub(super) async fn handle_attach_sandbox_provider(
let provider_name = request.provider_name.clone();
let attached = Arc::new(AtomicBool::new(false));
let attached_clone = attached.clone();
let mutation_id = uuid::Uuid::new_v4().to_string();
let sandbox = state
.store
@@ -1179,16 +1198,22 @@ pub(super) async fn handle_attach_sandbox_provider(
&sandbox_id,
request.expected_resource_version,
|sandbox| {
attached_clone.store(false, Ordering::Relaxed);
let Some(ref mut spec) = sandbox.spec else {
// Spec should always exist post-creation; if missing, fail CAS to surface error
return;
};
if spec.provider_attachment_epoch.is_empty() {
spec.provider_attachment_epoch.clone_from(&mutation_id);
}
dedupe_provider_names(&mut spec.providers);
if !spec.providers.iter().any(|name| name == &provider_name)
&& spec.providers.len() < MAX_PROVIDERS
{
spec.providers.push(provider_name.clone());
spec.provider_attachment_epoch.clone_from(&mutation_id);
attached_clone.store(true, Ordering::Relaxed);
}
},
@@ -1197,6 +1222,18 @@ pub(super) async fn handle_attach_sandbox_provider(
.map_err(|e| super::persistence_error_to_status(e, "attach sandbox provider"))?;
let attached = attached.load(Ordering::Relaxed);
let receipt = super::provider_readiness::record_provider_mutation(
state,
&sandbox,
&request.provider_name,
ProviderMutationKind::Attach,
Some((
provider_record.object_id(),
provider_record.get_resource_version(),
)),
&mutation_id,
)
.await?;
info!(
sandbox_name = %request.sandbox_name,
@@ -1208,6 +1245,7 @@ pub(super) async fn handle_attach_sandbox_provider(
Ok(Response::new(AttachSandboxProviderResponse {
sandbox: Some(sandbox),
attached,
receipt: Some(receipt),
}))
}
@@ -1271,6 +1309,7 @@ pub(super) async fn handle_detach_sandbox_provider(
let provider_name = request.provider_name.clone();
let detached = Arc::new(AtomicBool::new(false));
let detached_clone = detached.clone();
let mutation_id = uuid::Uuid::new_v4().to_string();
let sandbox = state
.store
@@ -1278,14 +1317,20 @@ pub(super) async fn handle_detach_sandbox_provider(
&sandbox_id,
request.expected_resource_version,
|sandbox| {
detached_clone.store(false, Ordering::Relaxed);
let Some(ref mut spec) = sandbox.spec else {
// Spec should always exist post-creation; if missing, fail CAS to surface error
return;
};
if spec.provider_attachment_epoch.is_empty() {
spec.provider_attachment_epoch.clone_from(&mutation_id);
}
let before_len = spec.providers.len();
spec.providers.retain(|name| name != &provider_name);
if spec.providers.len() != before_len {
spec.provider_attachment_epoch.clone_from(&mutation_id);
detached_clone.store(true, Ordering::Relaxed);
// Only dedupe after making a change
dedupe_provider_names(&mut spec.providers);
@@ -1296,6 +1341,15 @@ pub(super) async fn handle_detach_sandbox_provider(
.map_err(|e| super::persistence_error_to_status(e, "detach sandbox provider"))?;
let detached = detached.load(Ordering::Relaxed);
let receipt = super::provider_readiness::record_provider_mutation(
state,
&sandbox,
&request.provider_name,
ProviderMutationKind::Detach,
None,
&mutation_id,
)
.await?;
info!(
sandbox_name = %request.sandbox_name,
@@ -1307,6 +1361,7 @@ pub(super) async fn handle_detach_sandbox_provider(
Ok(Response::new(DetachSandboxProviderResponse {
sandbox: Some(sandbox),
detached,
receipt: Some(receipt),
}))
}
@@ -5407,6 +5462,7 @@ mod tests {
"template",
"resource_requirements",
],
&["provider_attachment_epoch"],
);
}
@@ -5414,6 +5470,7 @@ mod tests {
message_name: &str,
copied_from_create_request: &[&str],
rejected_template_workload_overrides: &[&str],
generated_by_gateway: &[&str],
) {
let pool = prost_reflect::DescriptorPool::decode(openshell_core::FILE_DESCRIPTOR_SET)
.expect("decode descriptor set");
@@ -5423,8 +5480,16 @@ mod tests {
let classified: std::collections::HashSet<&str> = copied_from_create_request
.iter()
.chain(rejected_template_workload_overrides.iter())
.chain(generated_by_gateway.iter())
.copied()
.collect();
assert_eq!(
classified.len(),
copied_from_create_request.len()
+ rejected_template_workload_overrides.len()
+ generated_by_gateway.len(),
"every field must have exactly one create-time owner"
);
let actual: std::collections::HashSet<String> = message
.fields()
.map(|field| field.name().to_string())
@@ -5435,7 +5500,8 @@ mod tests {
classified.contains(field.as_str()),
"{message_name}.{field} is not classified for template-backed sandbox creates. \
Add it to copied_from_create_request when callers own the create-time value, \
or to rejected_template_workload_overrides when the workload template owns it."
to rejected_template_workload_overrides when the workload template owns it, \
or to generated_by_gateway when the gateway replaces the caller's value."
);
}
@@ -5448,6 +5514,56 @@ mod tests {
}
}
#[tokio::test]
async fn create_sandbox_ignores_caller_provider_attachment_epoch() {
let state = test_server_state().await;
handle_create_sandbox_template(
&state,
authed_request(CreateSandboxTemplateRequest {
request_id: String::new(),
template: Some(test_workload_template("epoch-template")),
workspace_scope: Some(openshell_core::proto::workspace_selector("default")),
}),
)
.await
.unwrap();
let supplied_epoch = uuid::Uuid::new_v4().to_string();
let mut generated_epochs = std::collections::HashSet::new();
for (name, workload_template_name) in
[("direct-epoch", ""), ("template-epoch", "epoch-template")]
{
let created = handle_create_sandbox(
&state,
authed_request(CreateSandboxRequest {
name: name.to_string(),
spec: Some(SandboxSpec {
provider_attachment_epoch: supplied_epoch.clone(),
..Default::default()
}),
workload_template_name: workload_template_name.to_string(),
workspace_scope: Some(openshell_core::proto::workspace_selector("default")),
..Default::default()
}),
)
.await
.unwrap()
.into_inner()
.sandbox
.unwrap();
let epoch = &created.spec.as_ref().unwrap().provider_attachment_epoch;
assert_ne!(epoch, &supplied_epoch);
assert!(uuid::Uuid::parse_str(epoch).is_ok());
assert!(generated_epochs.insert(epoch.clone()));
let stored = state
.store
.get_message::<Sandbox>(created.object_id())
.await
.unwrap()
.unwrap();
assert_eq!(&stored.spec.unwrap().provider_attachment_epoch, epoch);
}
}
#[tokio::test]
async fn create_sandbox_from_workload_template_resolves_workload_and_preserves_governance() {
let state = test_server_state().await;
+1
View File
@@ -18,6 +18,7 @@ pub mod certgen;
pub mod cli;
mod compute;
pub mod config_file;
mod config_update_operation;
mod credentials;
mod defaults;
mod gateway_listener;
+65 -29
View File
@@ -116,13 +116,13 @@ mod tests {
use std::collections::{BTreeMap, BTreeSet, VecDeque};
const STORAGE_V1_SCHEMA_SHA256: &str =
"574bf5fcff731bd6e3fd84ed3f124161035bd236ef0fb7e32b4d8a8c55ceba5e";
"d68401809d8cea445c35233ef32412bbd041cb2ac5acaf368a0d0bf74d2ddf17";
const PUBLIC_RPC_SCHEMA_SHA256: &str =
"1c72f65167a5ceb00cd17043079b92324ca55d19b547c86827b51bd7376cb23c";
"418130a28d1a5398d3d7f73f11b7f025d027c70f215be320dfe8e58371769f83";
const DURABLE_SCHEMA_SHA256: &str =
"65066c0b0eef57a4c708f20fcbbb8e8f47376da9f4bf73dfc3bca0b3df174ba8";
"557ca283c55fd46b213d5573b950ba8604cc3f4b31bad3e433eb9c5f9975138d";
const PUBLIC_DURABLE_OVERLAP_SHA256: &str =
"39e8aaf0d1fbc86906c49a9e7f60641a3ce203d3130799c8065e09acf9d53ddf";
"f541c25bb3e1e5806865bc61470c66d10c16cca7399bf5909c77dc384d469171";
// A persisted Sandbox without endpoint status retains its lifecycle fields;
// the absent repeated field decodes empty and needs no database rewrite.
const SANDBOX_WITHOUT_ENDPOINT_STATUS: &str = "0a1e0a0a73616e64626f782d6964120773616e64626f783a0764656661756c741a2b0a0773616e64626f782a0d0a05526561647912045472756530023807420d73757065727669736f722d6964";
@@ -138,9 +138,10 @@ mod tests {
"0a0472756c651a07666978747572652d0000403f3a0b6578616d706c652e636f6d40bb035002";
const V0_0_116_POLICY_RECORD: &str = "0a09706f6c6963792d6964120a73616e64626f782d6964180222030102032a0673686132353632066c6f616465643a046e6f6e6540fa0148ac0252110a06736f75726365120766697874757265";
const V0_0_116_DRAFT_RECORD: &str = "0a086368756e6b2d6964120a73616e64626f782d69641802220770656e64696e672a0472756c65320204053a076669787475726549000000000000e83f50de02589003620b6578616d706c652e636f6d68bb037801";
const STORAGE_MESSAGE_NAMES: [&str; 8] = [
const STORAGE_MESSAGE_NAMES: [&str; 9] = [
"DraftChunkPayload",
"PolicyRevisionPayload",
"StoredConfigUpdateOperation",
"StoredDraftChunk",
"StoredPolicyRevision",
"StoredProviderCredentialRefreshState",
@@ -148,12 +149,21 @@ mod tests {
"StoredProviderProfile",
"StoredRefreshMaterialDeletion",
];
const DURABLE_ROOTS: [&str; 12] = [
const PROVIDER_READINESS_RPC_SIGNATURES: [&str; 2] = [
"openshell.v1.OpenShell/GetSandboxProviderStatus|.openshell.v1.GetSandboxProviderStatusRequest|.openshell.v1.GetSandboxProviderStatusResponse|false|false",
"openshell.v1.OpenShell/ReportProviderReadiness|.openshell.v1.ReportProviderReadinessRequest|.openshell.v1.ReportProviderReadinessResponse|false|false",
];
// Synthetic SandboxSpec bytes with log level, provider, and command fields,
// emitted before the gateway-owned attachment epoch field was introduced.
const PRE_READINESS_SANDBOX_SPEC: &str =
"0a04696e666f421273796e7468657469632d70726f766964657262046563686f";
const DURABLE_ROOTS: &[&str] = &[
".openshell.datamodel.v1.Provider",
".openshell.datamodel.v1.Workspace",
".openshell.sandbox.v1.SandboxPolicy",
".openshell.storage.v1.DraftChunkPayload",
".openshell.storage.v1.PolicyRevisionPayload",
".openshell.storage.v1.StoredConfigUpdateOperation",
".openshell.storage.v1.StoredProviderCredentialRefreshStateV2",
".openshell.storage.v1.StoredProviderProfile",
".openshell.v1.Sandbox",
@@ -495,19 +505,34 @@ mod tests {
}
}
methods.sort();
assert_eq!(compiled_method_count, 101, "classify every compiled RPC");
assert_eq!(methods.len(), 75, "inventory every public gateway RPC");
for signature in PROVIDER_READINESS_RPC_SIGNATURES {
assert!(
methods.iter().any(|method| method == signature),
"provider readiness RPC is missing or changed: {signature}"
);
}
assert_eq!(
compiled_method_count,
101 + PROVIDER_READINESS_RPC_SIGNATURES.len(),
"classify every compiled RPC"
);
assert_eq!(
methods.len(),
75 + PROVIDER_READINESS_RPC_SIGNATURES.len(),
"inventory every public gateway RPC"
);
assert_eq!(
methods
.iter()
.filter(|method| method.starts_with("openshell.v1.OpenShell/"))
.count(),
75
75 + PROVIDER_READINESS_RPC_SIGNATURES.len()
);
assert!(methods.iter().all(|method| !method.contains(".storage.")));
let public_closure = schema_closure(&index, public_roots);
let durable_closure = schema_closure(&index, DURABLE_ROOTS.into_iter().map(str::to_string));
let durable_closure =
schema_closure(&index, DURABLE_ROOTS.iter().copied().map(str::to_string));
let overlap_messages = public_closure
.messages
.intersection(&durable_closure.messages)
@@ -533,28 +558,39 @@ mod tests {
let durable_inventory_hash = schema_fingerprint(&index, &durable_closure);
let overlap_hash = format!("{:x}", Sha256::digest(overlap_inventory.as_bytes()));
// Report the complete measured inventory on failure so one schema
// change exposes every affected boundary in the same focused run.
assert_eq!(
(public_closure.messages.len(), public_closure.enums.len()),
(283, 14)
(
(public_closure.messages.len(), public_closure.enums.len()),
(durable_closure.messages.len(), durable_closure.enums.len()),
(overlap_messages.len(), overlap_enums.len()),
public_inventory_hash.as_str(),
durable_inventory_hash.as_str(),
overlap_hash.as_str(),
),
(
(294, 20),
(90, 15),
(78, 15),
PUBLIC_RPC_SCHEMA_SHA256,
DURABLE_SCHEMA_SHA256,
PUBLIC_DURABLE_OVERLAP_SHA256
),
"the public/durable schema inventory changed; review API and storage ownership, preserve prior-payload decoding, and update the reviewed fingerprints"
);
assert_eq!(
(durable_closure.messages.len(), durable_closure.enums.len()),
(83, 9)
);
assert_eq!((overlap_messages.len(), overlap_enums.len()), (73, 9));
}
assert_eq!(
public_inventory_hash, PUBLIC_RPC_SCHEMA_SHA256,
"the public RPC schema closure changed; review API compatibility and update the inventory and architecture/gateway.md"
);
assert_eq!(
durable_inventory_hash, DURABLE_SCHEMA_SHA256,
"a durable protobuf root or transitive dependency changed; record migration handling and a prior-version fixture before updating this fingerprint"
);
assert_eq!(
overlap_hash, PUBLIC_DURABLE_OVERLAP_SHA256,
"the public/durable protobuf overlap changed; review both API and storage compatibility before updating this inventory"
);
#[test]
fn pre_readiness_sandbox_spec_decodes_with_initial_attachment_epoch() {
let spec = openshell_core::proto::SandboxSpec::decode(
legacy_bytes(PRE_READINESS_SANDBOX_SPEC).as_slice(),
)
.expect("prior sandbox spec must decode");
assert_eq!(spec.log_level, "info");
assert_eq!(spec.providers, ["synthetic-provider"]);
assert_eq!(spec.command, ["echo"]);
assert!(spec.provider_attachment_epoch.is_empty());
}
fn legacy_bytes(encoded: &str) -> Vec<u8> {
@@ -14,14 +14,16 @@ use tracing::{debug, info, warn};
use uuid::Uuid;
use openshell_core::proto::{
GatewayMessage, RelayFrame, RelayInit, RelayOpen, ReportMainProcessExitRequest,
ReportMainProcessExitResponse, Sandbox, SandboxPhase, SessionAccepted, SshRelayTarget,
SupervisorMessage, gateway_message, relay_open, supervisor_message,
GatewayMessage, ProviderReadinessObservation, RelayFrame, RelayInit, RelayOpen,
ReportMainProcessExitRequest, ReportMainProcessExitResponse, Sandbox, SandboxPhase,
SessionAccepted, SshRelayTarget, SupervisorMessage, gateway_message, relay_open,
supervisor_message,
};
use openshell_core::transport_errors::is_expected_transport_close_status;
use crate::ServerState;
use crate::auth::principal::Principal;
use crate::grpc::provider_readiness::ProviderReadinessEvidence;
use crate::persistence::ObjectId;
const HEARTBEAT_INTERVAL_SECS: u32 = 15;
@@ -71,6 +73,9 @@ struct LiveSession {
/// gateway restart invalidates every session and startup reconciliation
/// resets any persisted endpoint result before requests are served.
endpoint_report_cursor: Option<EndpointReportCursor>,
/// Installation evidence belongs to this connection and is never restored
/// from persistence or inherited by a replacement supervisor session.
provider_readiness: Option<ProviderReadinessEvidence>,
#[allow(dead_code)]
connected_at: Instant,
}
@@ -155,6 +160,7 @@ impl SupervisorSessionRegistry {
terminal_delivery_finalized: false,
endpoint_status_initialized: false,
endpoint_report_cursor: None,
provider_readiness: None,
connected_at: Instant::now(),
},
);
@@ -267,6 +273,79 @@ impl SupervisorSessionRegistry {
.is_some_and(|session| session.session_id == session_id)
}
/// Bind the authenticated hello's installation capability to its session.
/// Initialization is single-use and cannot erase accepted observations.
pub(crate) fn initialize_provider_readiness(
&self,
sandbox_id: &str,
session_id: &str,
evidence: ProviderReadinessEvidence,
) -> Result<(), Status> {
let mut sessions = self
.sessions
.lock()
.map_err(|_| Status::unavailable("supervisor session state is unavailable"))?;
let session = sessions
.get_mut(sandbox_id)
.filter(|session| session.session_id == session_id)
.ok_or_else(|| Status::failed_precondition("supervisor session was replaced"))?;
if session.provider_readiness.is_some() {
return Err(Status::failed_precondition(
"provider readiness is already initialized",
));
}
session.provider_readiness = Some(evidence);
Ok(())
}
/// Accept installation evidence only while its session owns this sandbox.
/// Session comparison and publication share one lock so reconnects cannot
/// transfer a predecessor's evidence into the replacement session.
pub(crate) fn accept_provider_readiness(
&self,
sandbox_id: &str,
active_instance_id: &str,
observation: ProviderReadinessObservation,
) -> Result<(), Status> {
let mut sessions = self
.sessions
.lock()
.map_err(|_| Status::unavailable("supervisor session state is unavailable"))?;
let evidence = sessions
.get_mut(sandbox_id)
.filter(|session| {
session.session_id == observation.session_id && session.endpoint_status_initialized
})
.and_then(|session| session.provider_readiness.as_mut())
.ok_or_else(|| {
Status::permission_denied(
"provider readiness requires the active supervisor session",
)
})?;
if !evidence.belongs_to_instance(active_instance_id) {
return Err(Status::failed_precondition(
"provider readiness requires the current sandbox instance",
));
}
evidence.accept(observation)
}
/// Snapshot installation evidence from the current initialized session.
/// Disconnect and replacement discard the previous connection's state.
pub(crate) fn provider_readiness(
&self,
sandbox_id: &str,
) -> Result<Option<ProviderReadinessEvidence>, Status> {
let sessions = self
.sessions
.lock()
.map_err(|_| Status::unavailable("supervisor session state is unavailable"))?;
Ok(sessions
.get(sandbox_id)
.filter(|session| session.endpoint_status_initialized)
.and_then(|session| session.provider_readiness.clone()))
}
/// Mark the current session as the observation authority after its
/// public endpoint results have been durably reset.
pub(crate) fn initialize_endpoint_status_authority(
@@ -870,6 +949,9 @@ pub async fn handle_connect_supervisor(
crate::auth::guard::ensure_sandbox_principal_scope(principal, &sandbox_id)?;
}
require_persisted_sandbox(&state.store, &sandbox_id).await?;
// Validate readiness identities before replacing a healthy session. Older
// supervisors remain usable but cannot assert provider installation.
let provider_readiness = ProviderReadinessEvidence::from_hello(&hello)?;
let session_id = Uuid::new_v4().to_string();
info!(
@@ -919,6 +1001,16 @@ pub async fn handle_connect_supervisor(
"supervisor session was replaced during endpoint status initialization",
));
}
if let Err(error) = state.supervisor_sessions.initialize_provider_readiness(
&sandbox_id,
&session_id,
provider_readiness,
) {
state
.supervisor_sessions
.remove_if_current(&sandbox_id, &session_id);
return Err(error);
}
// Step 3: Send SessionAccepted.
let accepted = GatewayMessage {
@@ -221,6 +221,22 @@ impl OpenShell for TestOpenShell {
Ok(Response::new(GetGatewayConfigResponse::default()))
}
async fn get_sandbox_provider_status(
&self,
_request: tonic::Request<openshell_core::proto::GetSandboxProviderStatusRequest>,
) -> Result<Response<openshell_core::proto::GetSandboxProviderStatusResponse>, Status> {
Err(Status::unimplemented(
"provider readiness is not exercised by this mock",
))
}
async fn report_provider_readiness(
&self,
_request: tonic::Request<openshell_core::proto::ReportProviderReadinessRequest>,
) -> Result<Response<openshell_core::proto::ReportProviderReadinessResponse>, Status> {
Err(Status::unimplemented("provider readiness"))
}
async fn get_sandbox_provider_environment(
&self,
_request: tonic::Request<GetSandboxProviderEnvironmentRequest>,
@@ -248,6 +248,22 @@ impl OpenShell for RelayGateway {
) -> Result<Response<openshell_core::proto::GetGatewayConfigResponse>, Status> {
Err(Status::unimplemented("unused"))
}
async fn get_sandbox_provider_status(
&self,
_request: tonic::Request<openshell_core::proto::GetSandboxProviderStatusRequest>,
) -> Result<Response<openshell_core::proto::GetSandboxProviderStatusResponse>, Status> {
Err(Status::unimplemented(
"provider readiness is not exercised by this mock",
))
}
async fn report_provider_readiness(
&self,
_request: tonic::Request<openshell_core::proto::ReportProviderReadinessRequest>,
) -> Result<Response<openshell_core::proto::ReportProviderReadinessResponse>, Status> {
Err(Status::unimplemented("provider readiness"))
}
async fn get_sandbox_provider_environment(
&self,
_: tonic::Request<openshell_core::proto::GetSandboxProviderEnvironmentRequest>,
+12 -8
View File
@@ -175,7 +175,7 @@ pub(crate) fn test_opa_query_count() -> u64 {
}
/// Generation guard captured when an HTTP tunnel or request path starts.
#[derive(Clone)]
#[derive(Clone, Debug)]
pub struct PolicyGenerationGuard {
captured_generation: u64,
current_generation: Arc<AtomicU64>,
@@ -714,7 +714,7 @@ impl OpaEngine {
/// validation guarantees as initial load. Atomically replaces the inner
/// engine on success; on failure the previous engine is untouched (LKG).
pub fn reload_from_proto(&self, proto: &ProtoSandboxPolicy) -> Result<()> {
self.reload_from_proto_with_pid(proto, 0)
self.reload_from_proto_with_pid(proto, 0).map(|_| ())
}
/// Reload policy from a proto with symlink resolution.
@@ -722,11 +722,12 @@ impl OpaEngine {
/// When `entrypoint_pid` is non-zero, binary paths that are symlinks
/// inside the container filesystem are resolved and added as additional
/// match entries. See [`from_proto_with_pid`] for details.
/// Returns evidence tied to the generation installed by this call.
pub fn reload_from_proto_with_pid(
&self,
proto: &ProtoSandboxPolicy,
entrypoint_pid: u32,
) -> Result<()> {
) -> Result<PolicyGenerationGuard> {
// Build a complete new engine through the same validated pipeline.
let new = Self::from_proto_with_pid(proto, entrypoint_pid)?;
let new_engine = new
@@ -742,8 +743,10 @@ impl OpaEngine {
.fail_closed_reason
.write()
.map_err(|_| miette::miette!("OPA fail-closed state lock poisoned"))? = None;
self.advance_generation();
Ok(())
let generation = self.advance_generation();
// Capture evidence while the installation lock still excludes a
// competing reload; reading the generation later can bind the wrong policy.
self.generation_guard(generation)
}
/// Reload the policy and middleware registry as one runtime generation.
@@ -752,12 +755,13 @@ impl OpaEngine {
/// engine and runner are then swapped while holding both locks, followed by
/// a single generation increment. A preparation or lock failure leaves the
/// live pair and generation untouched.
/// Returns evidence tied to that combined installation.
pub fn reload_policy_and_middleware_from_proto_with_pid(
&self,
proto: &ProtoSandboxPolicy,
entrypoint_pid: u32,
registry: MiddlewareRegistry,
) -> Result<()> {
) -> Result<PolicyGenerationGuard> {
let new = Self::from_proto_with_pid(proto, entrypoint_pid)?;
let new_engine = new
.engine
@@ -780,8 +784,8 @@ impl OpaEngine {
.fail_closed_reason
.write()
.map_err(|_| miette::miette!("OPA fail-closed state lock poisoned"))? = None;
self.advance_generation();
Ok(())
let generation = self.advance_generation();
self.generation_guard(generation)
}
/// Publish a deny-all quarantine generation without activating any part
@@ -8134,6 +8134,7 @@ network_policies:
},
);
let snapshot = ProviderCredentialSnapshot {
installation_id: String::new(),
revision: 42,
child_env: std::collections::HashMap::new(),
dynamic_credentials,
@@ -274,7 +274,7 @@ pub async fn run_networking(
"Container filesystem accessible, resolving policy binary symlinks"
);
match resolve_engine.reload_from_proto_with_pid(&resolve_proto, pid) {
Ok(()) => {
Ok(_) => {
info!(
pid = pid,
"Policy binary symlink resolution complete \
@@ -394,6 +394,7 @@ async fn run_single_session(
payload: Some(supervisor_message::Payload::Hello(SupervisorHello {
sandbox_id: config.sandbox_id.clone(),
instance_id: config.instance_id.clone(),
supports_provider_readiness: true,
})),
})
.await
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,837 @@
// SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
// SPDX-License-Identifier: Apache-2.0
//! Installation evidence reported under the accepted supervisor session.
//!
//! Desired-state cursors never establish success. Credentials, policy, and
//! the authenticated workload boundary must acknowledge the same authority.
use std::sync::Arc;
use std::time::Duration;
use openshell_core::proto::{
ProviderReadinessObservation, ProviderReadinessReason as Reason,
ReportProviderReadinessResponse,
};
use openshell_core::provider_credentials::ProviderCredentialState;
use openshell_isolation_interface::contract::{BoundaryExec, ProviderEnvironmentInstallation};
use openshell_supervisor_network::opa::PolicyGenerationGuard;
use tokio::sync::watch;
/// Authority captured before consuming a provider response's credential fields.
#[derive(Clone, Debug, Default, PartialEq, Eq)]
pub struct EnvironmentIdentity {
pub(crate) attachment_epoch: String,
pub(crate) revision: u64,
pub(crate) policy_hash: String,
}
impl EnvironmentIdentity {
/// Preserve the authority delivered with a credential snapshot before consuming it.
pub(crate) fn from_environment(
result: &openshell_core::grpc_client::ProviderEnvironmentResult,
) -> Self {
Self {
attachment_epoch: result.provider_attachment_epoch.clone(),
revision: result.provider_env_revision,
policy_hash: result.policy_hash.clone(),
}
}
/// Capture the requested authority without claiming that it is installed.
pub(crate) fn from_settings(result: &openshell_core::grpc_client::SettingsPollResult) -> Self {
Self {
attachment_epoch: result.provider_attachment_epoch.clone(),
revision: result.provider_env_revision,
policy_hash: result.policy_hash.clone(),
}
}
}
#[derive(Clone, Debug)]
struct InstalledPolicy {
epoch: String,
hash: String,
config_revision: u64,
generation: PolicyGenerationGuard,
}
#[derive(Clone, Debug)]
struct FailedPolicy {
identity: EnvironmentIdentity,
config_revision: u64,
}
#[derive(Clone, Debug)]
struct State {
identity: EnvironmentIdentity,
installation_id: String,
credential_reason: Reason,
expires_at_ms: Option<i64>,
policy: Option<InstalledPolicy>,
policy_failure: Option<FailedPolicy>,
process: Option<ProviderEnvironmentInstallation>,
process_reason: Reason,
}
/// Shared installation tracker; each transition wakes the independent reporter.
#[derive(Clone, Debug)]
pub struct Tracker {
state: watch::Sender<State>,
}
impl Tracker {
/// Start without credential, policy, or process installation evidence.
pub(crate) fn new() -> Self {
let (state, _) = watch::channel(State {
identity: EnvironmentIdentity::default(),
installation_id: String::new(),
credential_reason: Reason::WaitingForCredentials,
expires_at_ms: None,
policy: None,
policy_failure: None,
process: None,
process_reason: Reason::WaitingForProcess,
});
Self { state }
}
/// Invalidate old success before a replacement or fail-closed clear begins.
pub(crate) fn credentials_failed(&self, identity: EnvironmentIdentity, reason: Reason) {
self.state.send_modify(|state| {
state.identity = identity;
state.installation_id.clear();
state.credential_reason = reason;
state.expires_at_ms = None;
state.process = None;
state.process_reason = Reason::WaitingForProcess;
});
}
/// Record a completed credential installation and require a fresh process acknowledgment.
pub(crate) fn credentials_installed(
&self,
identity: EnvironmentIdentity,
credentials: &ProviderCredentialState,
expires_at_ms: Option<i64>,
) {
let snapshot = credentials.snapshot();
self.state.send_modify(|state| {
state.credential_reason =
if !identity.policy_hash.is_empty() && snapshot.revision == identity.revision {
Reason::Unspecified
} else {
Reason::SnapshotMismatch
};
state.identity = identity;
state.installation_id.clone_from(&snapshot.installation_id);
state.expires_at_ms = expires_at_ms;
state.process = None;
state.process_reason = Reason::WaitingForProcess;
});
}
/// Retry failed installations even when the requested fingerprint is unchanged.
pub(crate) fn needs_environment(&self, identity: &EnvironmentIdentity) -> bool {
let state = self.state.borrow();
state.credential_reason != Reason::Unspecified || state.identity != *identity
}
/// Bind policy evidence to the exact generation returned by runtime installation.
pub(crate) fn policy_activated(
&self,
identity: &EnvironmentIdentity,
config_revision: u64,
generation: PolicyGenerationGuard,
) {
self.state.send_modify(|state| {
state.policy = Some(InstalledPolicy {
epoch: identity.attachment_epoch.clone(),
hash: identity.policy_hash.clone(),
config_revision,
generation,
});
state.policy_failure = None;
});
}
/// Report the rejected desired configuration without acknowledging its policy.
pub(crate) fn policy_install_failed(
&self,
identity: EnvironmentIdentity,
config_revision: u64,
) {
self.state.send_modify(|state| {
// Failed desired identity must not masquerade as evidence for the
// last installed policy, even when the failure retains that policy.
state.policy_failure = Some(FailedPolicy {
identity,
config_revision,
});
});
}
fn process_installed(
&self,
installed: ProviderEnvironmentInstallation,
credentials: &ProviderCredentialState,
) {
let current = credentials.snapshot();
self.state.send_if_modified(|state| {
if installed.installation_id != current.installation_id
|| installed.installation_id != state.installation_id
|| installed.revision != state.identity.revision
{
return false;
}
let changed = state.process.as_ref() != Some(&installed)
|| state.process_reason != Reason::Unspecified;
state.process = Some(installed);
state.process_reason = Reason::Unspecified;
changed
});
}
fn process_failed(&self) {
self.state.send_if_modified(|state| {
let changed =
state.process.is_some() || state.process_reason != Reason::ProcessInstallFailed;
state.process = None;
state.process_reason = Reason::ProcessInstallFailed;
changed
});
}
/// Read current evidence, rechecking credential expiry and policy generation.
pub(crate) fn observation(
&self,
credentials: &ProviderCredentialState,
) -> ProviderReadinessObservation {
let state = self.state.borrow();
let current = credentials.snapshot();
let identity = state
.policy_failure
.as_ref()
.map_or(&state.identity, |failure| &failure.identity);
let expired = state
.expires_at_ms
.is_some_and(|expiry| expiry <= openshell_core::time::now_ms());
let credentials_installed = state.credential_reason == Reason::Unspecified
&& !expired
&& state.identity == *identity
&& current.installation_id == state.installation_id;
let policy_active = state.policy_failure.is_none()
&& state.policy.as_ref().is_some_and(|policy| {
!policy.hash.is_empty()
&& policy.hash == identity.policy_hash
&& policy.epoch == identity.attachment_epoch
&& !policy.generation.is_stale()
});
let launch_environment_installed = credentials_installed
&& state.process.as_ref().is_some_and(|process| {
process.installation_id == state.installation_id
&& process.revision == identity.revision
});
let reason = if state.policy_failure.is_some() {
Reason::PolicyActivationFailed
} else if expired {
Reason::CredentialExpired
} else if state.credential_reason != Reason::Unspecified {
state.credential_reason
} else if !credentials_installed {
Reason::WaitingForCredentials
} else if !policy_active {
Reason::WaitingForPolicy
} else if !launch_environment_installed {
if state.process_reason == Reason::Unspecified {
Reason::WaitingForProcess
} else {
state.process_reason
}
} else {
Reason::Unspecified
};
ProviderReadinessObservation {
attachment_epoch: identity.attachment_epoch.clone(),
provider_env_revision: identity.revision,
config_revision: state.policy_failure.as_ref().map_or_else(
|| {
state
.policy
.as_ref()
.map_or(0, |policy| policy.config_revision)
},
|failure| failure.config_revision,
),
policy_hash: identity.policy_hash.clone(),
credentials_installed,
policy_active,
launch_environment_installed,
process_instance_id: state
.process
.as_ref()
.map_or_else(String::new, |process| process.session_id.to_string()),
reason: reason.into(),
..Default::default()
}
}
/// Report while the owner holds this handle; reconnects use the accepted session.
pub(crate) fn start_reporter(
&self,
endpoint: String,
sandbox_id: String,
credentials: ProviderCredentialState,
sessions: watch::Receiver<Option<String>>,
boundary: Arc<dyn BoundaryExec>,
) -> Reporter {
let tracker = self.clone();
Reporter(tokio::spawn(async move {
tracker
.run_reporter(
credentials,
sessions,
boundary,
GatewayReporter {
endpoint,
sandbox_id,
},
)
.await;
}))
}
async fn run_reporter<C: ReportClient>(
&self,
credentials: ProviderCredentialState,
mut sessions: watch::Receiver<Option<String>>,
boundary: Arc<dyn BoundaryExec>,
client: C,
) {
let tracker = self;
let mut updates = tracker.state.subscribe();
let mut active_session = None;
let mut sequence = 0_u64;
let mut interval = tokio::time::interval(Duration::from_secs(5));
interval.set_missed_tick_behavior(tokio::time::MissedTickBehavior::Skip);
loop {
let guard = tracker
.state
.borrow()
.policy
.as_ref()
.map(|policy| policy.generation.clone())
.filter(|guard| !guard.is_stale());
tokio::select! {
changed = sessions.changed() => { if changed.is_err() { return; } }
changed = updates.changed() => { if changed.is_err() { return; } }
_ = interval.tick() => {}
() = async {
if let Some(guard) = guard { guard.wait_until_stale().await; }
else { std::future::pending::<()>().await; }
} => {}
}
let session = sessions.borrow_and_update().clone();
if session != active_session {
sequence = 0;
active_session = session.clone();
}
let Some(session) = session else {
continue;
};
// Synchronization is also a boundary liveness check. A stopped
// boundary cannot keep renewing successful launch evidence.
let synchronization = tokio::time::timeout(
Duration::from_secs(10),
boundary.synchronize_provider_environment(),
);
let installed = tokio::select! {
changed = sessions.changed() => { if changed.is_err() { return; } interval.reset_immediately(); continue; }
installed = synchronization => installed,
};
match installed {
Ok(Ok(installed)) => tracker.process_installed(installed, &credentials),
_ => tracker.process_failed(),
}
let Some(next_sequence) = sequence.checked_add(1) else {
return;
};
sequence = next_sequence;
let mut observation = tracker.observation(&credentials);
observation.session_id.clone_from(&session);
observation.sequence = sequence;
let report = tokio::time::timeout(Duration::from_secs(5), client.report(observation));
let response = tokio::select! {
changed = sessions.changed() => { if changed.is_err() { return; } interval.reset_immediately(); continue; }
response = report => response,
};
if let Ok(Ok(response)) = response
&& !report_acknowledged(&response, sequence)
{
// A gateway rejection does not invalidate the boundary's
// installation. Retry on the normal cadence or a real state
// change without creating a self-triggered failure/retry loop.
tracing::warn!("Provider readiness report was not acknowledged");
}
}
}
}
/// An acknowledgment must bind this report and a valid positive evidence lease.
fn report_acknowledged(response: &ReportProviderReadinessResponse, sequence: u64) -> bool {
response.accepted_sequence == sequence
&& response
.observation_ttl
.as_ref()
.and_then(|ttl| openshell_core::time::duration_to_std(ttl).ok())
.is_some_and(|ttl| !ttl.is_zero())
}
#[tonic::async_trait]
trait ReportClient: Send + Sync {
async fn report(
&self,
observation: ProviderReadinessObservation,
) -> miette::Result<ReportProviderReadinessResponse>;
}
struct GatewayReporter {
endpoint: String,
sandbox_id: String,
}
#[tonic::async_trait]
impl ReportClient for GatewayReporter {
async fn report(
&self,
observation: ProviderReadinessObservation,
) -> miette::Result<ReportProviderReadinessResponse> {
openshell_core::grpc_client::report_provider_readiness(
&self.endpoint,
&self.sandbox_id,
observation,
)
.await
}
}
/// Cancels reporting before supervisor teardown can renew stale evidence.
pub struct Reporter(tokio::task::JoinHandle<()>);
impl Drop for Reporter {
fn drop(&mut self) {
self.0.abort();
}
}
#[cfg(test)]
mod tests {
use super::*;
use std::collections::HashMap;
#[test]
fn acknowledgment_requires_matching_sequence_and_valid_positive_ttl() {
let mut response = ReportProviderReadinessResponse {
accepted_sequence: 7,
..Default::default()
};
assert!(!report_acknowledged(&response, 7));
for (seconds, nanos, valid) in [
(15, 0, true),
(0, 1, true),
(0, 0, false),
(-1, 0, false),
(0, -1, false),
(1, -1, false),
(0, 1_000_000_000, false),
(315_576_000_001, 0, false),
] {
response.observation_ttl = Some(prost_types::Duration { seconds, nanos });
assert_eq!(
report_acknowledged(&response, 7),
valid,
"seconds={seconds}, nanos={nanos}"
);
}
response.observation_ttl = Some(prost_types::Duration {
seconds: 15,
nanos: 0,
});
assert!(!report_acknowledged(&response, 8));
}
struct CapturingReporter(tokio::sync::mpsc::UnboundedSender<ProviderReadinessObservation>);
#[tonic::async_trait]
impl ReportClient for CapturingReporter {
async fn report(
&self,
observation: ProviderReadinessObservation,
) -> miette::Result<ReportProviderReadinessResponse> {
let sequence = observation.sequence;
self.0.send(observation).unwrap();
Ok(ReportProviderReadinessResponse {
accepted_sequence: sequence,
report_interval: Some(
openshell_core::time::duration_from_std(Duration::from_secs(5)).unwrap(),
),
observation_ttl: Some(
openshell_core::time::duration_from_std(Duration::from_secs(15)).unwrap(),
),
})
}
}
struct DelayedReporter(
tokio::sync::mpsc::UnboundedSender<(
ProviderReadinessObservation,
tokio::sync::oneshot::Sender<ReportProviderReadinessResponse>,
)>,
);
#[tonic::async_trait]
impl ReportClient for DelayedReporter {
async fn report(
&self,
observation: ProviderReadinessObservation,
) -> miette::Result<ReportProviderReadinessResponse> {
let (response, received) = tokio::sync::oneshot::channel();
self.0.send((observation, response)).unwrap();
received
.await
.map_err(|_| miette::miette!("report response channel closed"))
}
}
struct DelayedBoundary(
tokio::sync::mpsc::UnboundedSender<
tokio::sync::oneshot::Sender<ProviderEnvironmentInstallation>,
>,
);
#[tonic::async_trait]
impl BoundaryExec for DelayedBoundary {
async fn exec(
&self,
_spec: openshell_isolation_interface::contract::ExecSpec,
) -> Result<
openshell_isolation_interface::contract::ExecSession,
openshell_isolation_interface::contract::BackendError,
> {
unreachable!("reporting must never launch a workload process")
}
async fn synchronize_provider_environment(
&self,
) -> Result<
ProviderEnvironmentInstallation,
openshell_isolation_interface::contract::BackendError,
> {
let (sender, receiver) = tokio::sync::oneshot::channel();
self.0.send(sender).unwrap();
receiver.await.map_err(|_| {
openshell_isolation_interface::contract::BackendError::Unavailable(
"boundary disconnected".into(),
)
})
}
}
async fn next_installation_report(
installations: &mut tokio::sync::mpsc::UnboundedReceiver<
tokio::sync::oneshot::Sender<ProviderEnvironmentInstallation>,
>,
reports: &mut tokio::sync::mpsc::UnboundedReceiver<(
ProviderReadinessObservation,
tokio::sync::oneshot::Sender<ReportProviderReadinessResponse>,
)>,
installed: &ProviderEnvironmentInstallation,
wait: Duration,
) -> (
ProviderReadinessObservation,
tokio::sync::oneshot::Sender<ReportProviderReadinessResponse>,
) {
tokio::time::timeout(wait, async {
installations
.recv()
.await
.unwrap()
.send(installed.clone())
.unwrap();
reports.recv().await.unwrap()
})
.await
.expect("reporter did not synchronize and report before the deadline")
}
#[tokio::test]
async fn invalid_report_acknowledgments_retry_without_installation_feedback() {
let tracker = Tracker::new();
let credentials = ProviderCredentialState::from_child_env_snapshot(6, HashMap::new());
tracker.credentials_installed(identity(), &credentials, None);
let policy = openshell_policy::restrictive_default_policy();
let engine = openshell_supervisor_network::opa::OpaEngine::from_proto(&policy).unwrap();
tracker.policy_activated(&identity(), 12, engine.generation_guard(0).unwrap());
let mut installed = installation(&credentials);
let (sessions, receiver) = watch::channel(Some("current-session".to_string()));
let (boundary_tx, mut boundary_rx) = tokio::sync::mpsc::unbounded_channel();
let (report_tx, mut report_rx) = tokio::sync::mpsc::unbounded_channel();
let reporting_tracker = tracker.clone();
let reporting_credentials = credentials.clone();
let task = tokio::spawn(async move {
reporting_tracker
.run_reporter(
reporting_credentials,
receiver,
Arc::new(DelayedBoundary(boundary_tx)),
DelayedReporter(report_tx),
)
.await;
});
// The first installation changes local evidence once. Its watch update
// may cause one additional report, but an invalid gateway reply must
// not keep toggling that evidence and scheduling more installations.
for sequence in 1..=2 {
let (report, response) = next_installation_report(
&mut boundary_rx,
&mut report_rx,
&installed,
Duration::from_secs(1),
)
.await;
assert_eq!(report.sequence, sequence);
assert!(report.launch_environment_installed);
response
.send(ReportProviderReadinessResponse::default())
.unwrap();
}
assert!(
tokio::time::timeout(Duration::from_millis(100), boundary_rx.recv())
.await
.is_err(),
"invalid report replies must not repeatedly reinstall an unchanged environment"
);
// A real credential installation still wakes the reporter immediately.
credentials.install_child_env_snapshot(6, HashMap::new());
tracker.credentials_installed(identity(), &credentials, None);
installed.installation_id = credentials.snapshot().installation_id.clone();
let (report, response) = next_installation_report(
&mut boundary_rx,
&mut report_rx,
&installed,
Duration::from_secs(1),
)
.await;
assert_eq!(report.sequence, 3);
assert!(report.launch_environment_installed);
// A policy update received while the report is pending must survive
// processing its reply and cause a report of the new configuration.
tracker.policy_activated(&identity(), 13, engine.generation_guard(0).unwrap());
response
.send(ReportProviderReadinessResponse::default())
.unwrap();
let (report, response) = next_installation_report(
&mut boundary_rx,
&mut report_rx,
&installed,
Duration::from_secs(1),
)
.await;
assert_eq!(report.sequence, 4);
assert_eq!(report.config_revision, 13);
response
.send(ReportProviderReadinessResponse::default())
.unwrap();
assert!(
tokio::time::timeout(Duration::from_millis(100), boundary_rx.recv())
.await
.is_err()
);
sessions.send_replace(Some("replacement-session".to_string()));
let (report, response) = next_installation_report(
&mut boundary_rx,
&mut report_rx,
&installed,
Duration::from_secs(1),
)
.await;
assert_eq!(report.session_id, "replacement-session");
assert_eq!(report.sequence, 1);
response
.send(ReportProviderReadinessResponse::default())
.unwrap();
assert!(
tokio::time::timeout(Duration::from_millis(100), boundary_rx.recv())
.await
.is_err()
);
// Without further state changes, the periodic liveness check retries.
let (report, response) = next_installation_report(
&mut boundary_rx,
&mut report_rx,
&installed,
Duration::from_secs(6),
)
.await;
assert_eq!(report.session_id, "replacement-session");
assert_eq!(report.sequence, 2);
assert!(report.launch_environment_installed);
response
.send(ReportProviderReadinessResponse::default())
.unwrap();
task.abort();
let _ = task.await;
}
#[tokio::test]
async fn reconnect_cancels_delayed_boundary_ack_and_starts_a_new_report_sequence() {
let tracker = Tracker::new();
let credentials = ProviderCredentialState::from_child_env_snapshot(6, HashMap::new());
tracker.credentials_installed(identity(), &credentials, None);
let installed = installation(&credentials);
let (sessions, receiver) = watch::channel(Some("old-session".to_string()));
let (boundary_tx, mut boundary_rx) = tokio::sync::mpsc::unbounded_channel();
let (report_tx, mut report_rx) = tokio::sync::mpsc::unbounded_channel();
let task = tokio::spawn(async move {
tracker
.run_reporter(
credentials,
receiver,
Arc::new(DelayedBoundary(boundary_tx)),
CapturingReporter(report_tx),
)
.await;
});
let old_ack = tokio::time::timeout(Duration::from_secs(1), boundary_rx.recv())
.await
.unwrap()
.unwrap();
sessions.send_replace(Some("current-session".to_string()));
let current_ack = tokio::time::timeout(Duration::from_secs(1), boundary_rx.recv())
.await
.unwrap()
.unwrap();
assert!(
old_ack.send(installed.clone()).is_err(),
"replaced session must cancel its pending installation"
);
current_ack.send(installed.clone()).unwrap();
let report = tokio::time::timeout(Duration::from_secs(1), report_rx.recv())
.await
.unwrap()
.unwrap();
assert_eq!(report.session_id, "current-session");
assert_eq!(report.sequence, 1);
assert!(report.launch_environment_installed);
// The installation transition wakes a second report, whose sequence
// must advance even though its desired authority is unchanged.
let next_ack = tokio::time::timeout(Duration::from_secs(1), boundary_rx.recv())
.await
.unwrap()
.unwrap();
next_ack.send(installed).unwrap();
let report = tokio::time::timeout(Duration::from_secs(1), report_rx.recv())
.await
.unwrap()
.unwrap();
assert_eq!(report.sequence, 2);
assert!(
tokio::time::timeout(Duration::from_millis(30), boundary_rx.recv())
.await
.is_err(),
"an unchanged installation acknowledgment must not trigger another report"
);
sessions.send_replace(None);
assert!(
tokio::time::timeout(Duration::from_millis(30), report_rx.recv())
.await
.is_err()
);
task.abort();
let _ = task.await;
}
fn identity() -> EnvironmentIdentity {
EnvironmentIdentity {
attachment_epoch: "epoch".into(),
revision: 6,
policy_hash: "policy".into(),
}
}
fn installation(credentials: &ProviderCredentialState) -> ProviderEnvironmentInstallation {
ProviderEnvironmentInstallation {
installation_id: credentials.snapshot().installation_id.clone(),
revision: 6,
session_id: uuid::Uuid::new_v4().to_string().parse().unwrap(),
}
}
#[test]
fn same_revision_repair_rejects_the_previous_boundary_acknowledgment() {
let tracker = Tracker::new();
let credentials = ProviderCredentialState::from_child_env_snapshot(6, HashMap::new());
tracker.credentials_installed(identity(), &credentials, None);
let old = installation(&credentials);
credentials
.install_child_env_snapshot(6, HashMap::from([("TOKEN".into(), "restored".into())]));
tracker.credentials_installed(identity(), &credentials, None);
tracker.process_installed(old, &credentials);
assert!(
!tracker
.observation(&credentials)
.launch_environment_installed
);
tracker.process_installed(installation(&credentials), &credentials);
assert!(
tracker
.observation(&credentials)
.launch_environment_installed
);
}
#[test]
fn policy_failure_and_generation_change_invalidate_installed_evidence() {
let tracker = Tracker::new();
let credentials = ProviderCredentialState::from_child_env_snapshot(6, HashMap::new());
let policy = openshell_policy::restrictive_default_policy();
let engine = openshell_supervisor_network::opa::OpaEngine::from_proto(&policy).unwrap();
tracker.credentials_installed(identity(), &credentials, None);
tracker.policy_activated(&identity(), 12, engine.generation_guard(0).unwrap());
tracker.process_installed(installation(&credentials), &credentials);
assert!(tracker.observation(&credentials).policy_active);
engine.enter_fail_closed("test quarantine").unwrap();
assert!(!tracker.observation(&credentials).policy_active);
let mut rejected = identity();
rejected.policy_hash = "rejected".into();
tracker.policy_install_failed(rejected, 13);
let observed = tracker.observation(&credentials);
assert_eq!(observed.config_revision, 13);
assert_eq!(observed.policy_hash, "rejected");
assert!(!observed.credentials_installed);
assert!(!observed.launch_environment_installed);
}
#[test]
fn expired_credentials_and_disconnected_boundary_never_complete() {
let tracker = Tracker::new();
let credentials = ProviderCredentialState::from_child_env_snapshot(6, HashMap::new());
tracker.credentials_installed(identity(), &credentials, Some(1));
tracker.process_installed(installation(&credentials), &credentials);
assert_eq!(
tracker.observation(&credentials).reason,
i32::from(Reason::CredentialExpired)
);
tracker.credentials_installed(identity(), &credentials, None);
tracker.process_installed(installation(&credentials), &credentials);
tracker.process_failed();
assert!(
!tracker
.observation(&credentials)
.launch_environment_installed
);
}
}
+45 -8
View File
@@ -40,7 +40,7 @@ Provider profiles include these user-facing features:
- Provider instances whose submitted credentials can be stored by a configured gateway credential driver.
- Profile-backed credential discovery for explicit `openshell provider create --from-existing` and `openshell provider update --from-existing` flows. The built-in `google-vertex-ai` profile also supplements discovery with Vertex config env vars such as `VERTEX_AI_PROJECT_ID` and `VERTEX_AI_REGION`.
- Just-in-time effective policy composition from sandbox policy plus attached provider profiles.
- Runtime sandbox provider lifecycle commands under `openshell sandbox provider list|attach|detach`.
- Runtime sandbox provider commands under `openshell sandbox provider list|attach|detach|status`, with an option to wait until the sandbox applies a change.
- Credential refresh configuration with `openshell provider refresh status|configure|rotate|delete`.
- Credential expiry metadata with `openshell provider update --credential-expires-at`; values accept Unix epoch milliseconds or ISO/RFC3339 timestamps.
- Dynamic token grants that use the sandbox's SPIFFE JWT-SVID as an OAuth2 client assertion and inject short-lived tokens into supported headers for matching profile endpoints.
@@ -968,29 +968,66 @@ Composition follows these rules:
## Attach and Detach Providers
Attach an existing provider to a running sandbox:
Attach an existing provider and wait until the sandbox can use it:
```shell
openshell sandbox provider attach provider-demo work-github
openshell sandbox provider attach provider-demo work-github --wait --timeout 30
```
Detach a provider:
Detach a provider and wait until the sandbox stops resolving its credentials:
```shell
openshell sandbox provider detach provider-demo work-github
openshell sandbox provider detach provider-demo work-github --wait --timeout 30
```
Attach and detach are idempotent. Attach validates that the provider exists before mutating the sandbox, and provider deletion fails while the provider is attached to any sandbox.
### Inspect Provider Readiness
Without `--wait`, attach, detach, and update confirm that the gateway saved the change. Add `--wait` when the next step depends on that change taking effect. The command succeeds after the sandbox applies the matching credentials and policy, and updates the environment used by new processes.
Each result includes a change record called a `receipt` in the API. Its ID lets you check the same change later, including after a timeout:
```shell
openshell sandbox provider status provider-demo work-github --output json
openshell sandbox provider status provider-demo work-github --receipt RECEIPT_ID --wait --timeout 30
```
Replace `RECEIPT_ID` with the `receipt_id` returned by attach, detach, or update. The ID always refers to the original change. If a later change replaces it, the status is `superseded` and the wait ends without reporting success.
If attach, detach, or update reports `CONFIG_OPERATION_STORAGE_UNCERTAIN`, the change may already be saved even though the gateway could not record its readiness receipt. Do not blindly retry the mutation: inspect the provider and sandbox state and reconcile the saved change first. A receipt may be unavailable, so an error does not establish that the mutation was rolled back or that the sandbox is ready.
Readiness states have these meanings:
- `persisted`: the gateway saved the change; the command has not checked the sandbox yet.
- `pending`: the sandbox has not confirmed all parts of the change.
- `ready`: the sandbox applied the requested provider credentials, policy, and environment for new processes.
- `revoked`: the detached provider's credentials no longer resolve, and new processes do not receive its credential references.
- `withheld`: a credential or supervisor configuration prevents the sandbox from applying the change.
- `failed`: the sandbox could not install the credentials, policy, or process environment.
- `superseded`: a later change replaced the original request.
The default wait is 30 seconds and the maximum is 3600 seconds. The timeout starts after the gateway saves the change and includes time spent requesting status. If time runs out, the command exits with an error and reports `wait_outcome: timed_out` with the last known state. The change remains saved and may finish applying later. Check its ID again to learn the result.
An unavailable supervisor, an expired report, a failed installation, or a missing process-environment acknowledgment cannot produce `ready` or `revoked`. After a reconnect or restart, the current authenticated supervisor must report its installed state again.
JSON and YAML output include change IDs, requested and installed revisions, timestamps, reason categories, and a result for each sandbox. Revision strings identify configurations; compare them for equality rather than numerical order. Status output excludes credential values, credential references, authorization headers, and raw installation errors.
The `persisted_time`, `observed_time`, and `evaluated_time` fields use RFC 3339 timestamp strings. An absent `observed_time` means the current supervisor session has not supplied accepted evidence. API operation timestamps use protobuf `Timestamp`, and report intervals and observation lifetimes use protobuf `Duration`, following the [protobuf time representation](/reference/protobuf-time-types).
The API's `operation` field records the common operation's historical outcome; its `operation_id` equals the receipt ID. Use the provider `state` and the wait result for current readiness. A disconnected supervisor can make the live state pending even after the operation previously applied, and a newer change can supersede the live result. Historical completion does not override those checks.
### Runtime Limitations
Provider attach and detach update the persisted sandbox provider list. Running sandboxes poll for provider environment revisions and effective policy changes.
Running sandboxes periodically check for provider and policy changes. Use `--wait` or `sandbox provider status` to confirm when a saved change has taken effect.
The policy effect applies to future effective policy reads after the sandbox observes the update. The credential environment effect applies only to new process launches after the update is observed, such as later SSH, exec, or SFTP sessions.
Already-running processes keep the placeholder environment they started with. OpenShell does not mutate a live process environment after provider attach, detach, or credential update. The proxy resolves existing placeholders against current credentials and bindings, so rotation, expiry, endpoint changes, and detach take effect without restarting the process. If a long-running process needs a newly attached provider credential placeholder, restart that process or launch a new process after the sandbox has observed the provider update.
Already-running processes keep the placeholder environment they started with. For ordinary static credentials, an existing reference retains its selected revision after a value update; readiness does not retarget that reference to the replacement value. After an update wait succeeds, launch a new process to receive the updated credential reference. Gateway-managed refresh follows its own reference lifecycle. Expiry, endpoint authorization, and acknowledged detach continue to apply when the proxy resolves retained references.
Detaching a provider removes its provider policy layer from future effective policy reads, revokes resolution for its existing placeholders, and removes its credential placeholders from future process environments. It does not remove the placeholder strings from already-running process environments.
For a static provider, the sequence is attach, wait, then launch client A; update, wait, then launch client B. B's first request uses the newly installed credential. Readiness does not establish that A has stopped using its retained revision or that an old upstream key can be retired.
An acknowledged detachment removes its provider policy layer from the active effective policy, revokes future resolution for its existing placeholders, and removes its credential placeholders from future process environments. It does not remove strings from already-running process environments or undo requests already forwarded upstream.
OpenShell rejects provider updates and refresh configuration when they would make two providers attached to the same sandbox expose the same active credential environment key. It also rejects attached provider sets with ambiguous dynamic token grants at equal host/path specificity. Use provider-specific credential names and make one dynamic grant selector more specific when one sandbox needs multiple providers with overlapping upstream concepts.
+1
View File
@@ -145,6 +145,7 @@ tokens are first-execution preconditions, not conditions to execute again.
|---|---|
| Workspace, membership, template, provider, and profile resources | Load the original UUID at its recorded version. Missing or changed resources make replay unavailable. Provider credentials remain redacted. |
| Sandbox resources | Load the original UUID with its current state, including newer status or configuration. Attach/detach flags describe the original operation. A missing original sandbox makes replay unavailable. |
| Provider readiness receipts | Return the original attach, detach, or update receipts and mutation ID, including the original update target set. Later attachment changes do not replace these receipts. Missing operation evidence makes replay unavailable without executing the mutation again. |
| Service endpoint | Require the original endpoint version and sandbox identity. Return the recorded URL. |
| Refresh status | Return current status for the original provider, refresh record, and authorization epoch. Reconfiguration or removal makes replay unavailable. Rotation is not repeated. |
| Policy and config | Return the original versions, counts, and nonsecret annotations while the original sandbox exists. Global config has no sandbox guard. |
+12
View File
@@ -178,6 +178,18 @@ Update a provider's credentials:
openshell provider update my-claude --from-existing
```
To wait until attached sandboxes apply the updated credentials, add `--wait`:
```shell
openshell provider update my-claude --from-existing --wait --timeout 30 --output json
```
The update records which sandboxes are attached before saving the new credentials. Its result includes a change ID and outcome for each of those sandboxes. Sandboxes attached later are outside this wait. The timeout applies to the whole group; if any selected sandbox fails or times out, the command exits with an error and still reports each sandbox's outcome. When no sandboxes are attached, it reports the saved change and an empty target list.
Credential refresh status tells you whether OpenShell obtained credentials. [Provider readiness](/providers/profiles#inspect-provider-readiness) tells you whether the sandbox applied the credentials, policy, and environment for new processes. It does not test access to a backend model or cancel requests already sent upstream.
After a static credential update completes, launch a new client process to use the updated reference. An existing process keeps its revision-scoped reference; readiness does not make that old reference resolve the new value. Acknowledged detach revokes retained references and removes them from environments for future launches.
Set or clear a credential expiry timestamp:
```shell
+5
View File
@@ -108,6 +108,11 @@ name = "provider_refresh_handles"
path = "tests/provider_refresh_handles.rs"
required-features = ["e2e-podman"]
[[test]]
name = "provider_readiness"
path = "tests/provider_readiness.rs"
required-features = ["e2e-docker"]
[[test]]
name = "vm_gateway_start"
path = "tests/vm_gateway_start.rs"
+46 -4
View File
@@ -808,6 +808,7 @@ async fn workspace_admin_cannot_manage_another_workspace_members() {
async fn workspace_admin_cannot_manage_another_workspace_providers() {
const WORKSPACE_A: &str = "oidc-wsa-xprov-a";
const WORKSPACE_B: &str = "oidc-wsa-xprov-b";
const PROVIDER: &str = "oidc-wsa-xprovider";
let (admin, workspace_admin, _user_b) =
prepare_isolated_workspaces_with_admin(WORKSPACE_A, WORKSPACE_B).await;
@@ -818,7 +819,7 @@ async fn workspace_admin_cannot_manage_another_workspace_providers() {
"provider",
"create",
"--name",
"oidc-wsa-xprovider",
PROVIDER,
"--type",
"openai",
"--credential",
@@ -826,9 +827,50 @@ async fn workspace_admin_cannot_manage_another_workspace_providers() {
],
)
.await;
assert_non_member_denial(
&denied,
"manage another workspace's providers as workspace admin",
let diagnostic = combined_output(&denied);
let compact: String = diagnostic
.chars()
.filter(|character| !character.is_whitespace() && *character != '│')
.collect();
// Profile lookup redacts backend diagnostics. Check the permission code and
// safe recovery guidance without requiring the server's membership details.
assert!(
!denied.status.success()
&& compact.contains("PERMISSION_DENIED")
&& compact.contains("verifyworkspacemembershipandrequiredpermissions"),
"cross-workspace provider creation did not report a safe permission denial:\n{diagnostic}"
);
assert!(!diagnostic.contains("e2e-test-value"));
// Query with an independent authorized identity so a failed command alone
// cannot hide a provider created before the denial was returned.
let listed = assert_workspace_allowed(
&admin,
WORKSPACE_B,
&["provider", "list", "--output", "json"],
"verify denied creation left the target workspace empty",
)
.await;
// CLI startup diagnostics may precede the JSON object on stdout.
let stdout = String::from_utf8(listed.stdout).expect("provider list output should be UTF-8");
let json_start = stdout
.find('{')
.expect("provider list output should contain JSON");
let json_end = stdout
.rfind('}')
.expect("provider list output should contain a complete JSON object");
let listing: Value = serde_json::from_str(&stdout[json_start..=json_end])
.expect("provider list --output json should return JSON on stdout");
assert_eq!(
listing["next_page_token"], "",
"provider listing is incomplete"
);
assert!(
listing["providers"]
.as_array()
.expect("provider collection")
.is_empty(),
"denied creation added a provider to the isolated target workspace"
);
delete_workspace(&admin, WORKSPACE_B).await;
File diff suppressed because it is too large Load Diff
+246
View File
@@ -158,6 +158,16 @@ service OpenShell {
};
}
// Inspect the installed authority for one sandbox provider mutation.
rpc GetSandboxProviderStatus(GetSandboxProviderStatusRequest)
returns (GetSandboxProviderStatusResponse) {
option (openshell.options.v1.authorization) = {
auth_mode: "bearer"
scope: "sandbox:read"
workspace_role: "user"
};
}
// Delete a sandbox by name.
rpc DeleteSandbox(DeleteSandboxRequest) returns (DeleteSandboxResponse) {
option (openshell.options.v1.authorization) = {
@@ -484,6 +494,15 @@ service OpenShell {
};
}
// Report installed provider state for the current ConnectSupervisor session.
// Replacing or losing that session invalidates its observations.
rpc ReportProviderReadiness(ReportProviderReadinessRequest)
returns (ReportProviderReadinessResponse) {
option (openshell.options.v1.authorization) = {
auth_mode: "sandbox"
};
}
// Get provider environment for a sandbox (called by sandbox supervisor at startup).
rpc GetSandboxProviderEnvironment(GetSandboxProviderEnvironmentRequest)
returns (GetSandboxProviderEnvironmentResponse) {
@@ -953,6 +972,9 @@ message SandboxSpec {
repeated string command = 12;
// Allocate a retained pseudo-terminal for the main process.
bool tty = 13;
// Gateway-owned attachment identity, changed atomically with the provider set.
// Equality only: detach and reattach must not revive an older receipt.
string provider_attachment_epoch = 14;
}
message ResourceRequirements {
@@ -1400,6 +1422,8 @@ message AttachSandboxProviderResponse {
Sandbox sandbox = 1;
// True when the provider was newly attached. False means it was already attached.
bool attached = 2;
// Persisted intent; readiness requires current supervisor observations.
ProviderMutationReceipt receipt = 3;
}
// Detach provider from sandbox response.
@@ -1407,6 +1431,215 @@ message DetachSandboxProviderResponse {
Sandbox sandbox = 1;
// True when the provider was removed. False means it was not attached.
bool detached = 2;
// Revocation is complete only when this receipt reports REVOKED.
ProviderMutationReceipt receipt = 3;
}
// Operation whose installed authority is tracked by a receipt.
enum ProviderMutationKind {
PROVIDER_MUTATION_KIND_UNSPECIFIED = 0;
PROVIDER_MUTATION_KIND_ATTACH = 1;
PROVIDER_MUTATION_KIND_DETACH = 2;
PROVIDER_MUTATION_KIND_UPDATE = 3;
// Reconstructed status for existing desired state without a mutation receipt.
PROVIDER_MUTATION_KIND_OBSERVE = 4;
}
// Readiness states describe persisted intent separately from installed state.
enum ProviderReadinessState {
PROVIDER_READINESS_STATE_UNSPECIFIED = 0;
PROVIDER_READINESS_STATE_PERSISTED = 1;
PROVIDER_READINESS_STATE_PENDING = 2;
PROVIDER_READINESS_STATE_READY = 3;
PROVIDER_READINESS_STATE_WITHHELD = 4;
PROVIDER_READINESS_STATE_REVOKED = 5;
PROVIDER_READINESS_STATE_FAILED = 6;
PROVIDER_READINESS_STATE_SUPERSEDED = 7;
}
// Closed reason categories are safe to display. Raw installation errors are
// never part of the readiness protocol.
enum ProviderReadinessReason {
PROVIDER_READINESS_REASON_UNSPECIFIED = 0;
PROVIDER_READINESS_REASON_WAITING_FOR_SUPERVISOR = 1;
PROVIDER_READINESS_REASON_WAITING_FOR_CREDENTIALS = 2;
PROVIDER_READINESS_REASON_WAITING_FOR_POLICY = 3;
PROVIDER_READINESS_REASON_WAITING_FOR_PROCESS = 4;
PROVIDER_READINESS_REASON_UNSUPPORTED_SUPERVISOR = 5;
PROVIDER_READINESS_REASON_CREDENTIALS_WITHHELD = 6;
PROVIDER_READINESS_REASON_CREDENTIAL_INSTALL_FAILED = 7;
PROVIDER_READINESS_REASON_POLICY_ACTIVATION_FAILED = 8;
PROVIDER_READINESS_REASON_PROCESS_INSTALL_FAILED = 9;
PROVIDER_READINESS_REASON_SUPERVISOR_DISCONNECTED = 10;
PROVIDER_READINESS_REASON_SUPERVISOR_LEASE_EXPIRED = 11;
PROVIDER_READINESS_REASON_DESIRED_STATE_CHANGED = 12;
PROVIDER_READINESS_REASON_CREDENTIAL_EXPIRED = 13;
PROVIDER_READINESS_REASON_LOCAL_POLICY = 14;
PROVIDER_READINESS_REASON_SNAPSHOT_MISMATCH = 15;
}
// Exact desired authority. Revisions are opaque identities, never ordered.
message ProviderDesiredIdentity {
string sandbox_id = 1;
string sandbox_name = 2;
string attachment_epoch = 3;
// Empty for a detached provider.
string provider_id = 4;
uint64 provider_resource_version = 5;
uint64 provider_env_revision = 6;
uint64 config_revision = 7;
string policy_hash = 8;
}
// Component whose desired state is tracked by a durable update operation.
enum ConfigComponent {
CONFIG_COMPONENT_UNSPECIFIED = 0;
CONFIG_COMPONENT_SANDBOX_CONFIG = 1;
CONFIG_COMPONENT_PROVIDER_ENVIRONMENT = 2;
}
// Identifies one component snapshot revision. Revisions are equality tokens,
// not members of one shared ordering domain.
message ConfigSnapshotRevision {
oneof component {
SandboxConfigRevision sandbox_config = 1;
uint64 provider_environment = 2;
// Complete desired authority for one sandbox-scoped provider mutation.
ProviderDesiredIdentity provider_target = 3;
}
}
// Identity needed to correlate effective sandbox configuration with the
// policy-history row whose apply status the gateway records.
message SandboxConfigRevision {
uint64 config_revision = 1;
uint32 policy_version = 2;
openshell.sandbox.v1.PolicySource policy_source = 3;
uint32 global_policy_version = 4;
// Monotonic revision of the sandbox-scoped settings row. This disambiguates
// setting operations whose effective config fingerprint is equality-only.
uint64 settings_revision = 5;
}
// Result of applying a component revision at its owning runtime boundary.
enum ConfigApplyOutcome {
CONFIG_APPLY_OUTCOME_UNSPECIFIED = 0;
CONFIG_APPLY_OUTCOME_APPLIED = 1;
CONFIG_APPLY_OUTCOME_IGNORED_DUPLICATE = 2;
CONFIG_APPLY_OUTCOME_IGNORED_STALE = 3;
CONFIG_APPLY_OUTCOME_RETAINED_LOCAL_OVERRIDE = 4;
CONFIG_APPLY_OUTCOME_DEGRADED = 5;
CONFIG_APPLY_OUTCOME_FAILED_RETAINED_LAST_KNOWN_GOOD = 6;
CONFIG_APPLY_OUTCOME_FAILED_CLOSED = 7;
CONFIG_APPLY_OUTCOME_UNSUPPORTED = 8;
}
// Durable lifecycle of one desired-state update operation.
enum ConfigUpdateOperationState {
CONFIG_UPDATE_OPERATION_STATE_UNSPECIFIED = 0;
CONFIG_UPDATE_OPERATION_STATE_PENDING = 1;
CONFIG_UPDATE_OPERATION_STATE_APPLIED = 2;
CONFIG_UPDATE_OPERATION_STATE_INACTIVE = 3;
CONFIG_UPDATE_OPERATION_STATE_FAILED = 4;
CONFIG_UPDATE_OPERATION_STATE_SUPERSEDED = 5;
CONFIG_UPDATE_OPERATION_STATE_CANCELLED = 6;
}
// Durable progress for one sandbox-scoped desired-state mutation. Snapshot
// contents and credentials are never stored in this resource.
message ConfigUpdateOperation {
string operation_id = 1;
string sandbox_id = 2;
ConfigComponent component = 3;
ConfigSnapshotRevision target_revision = 4;
ConfigUpdateOperationState state = 5;
ConfigApplyOutcome outcome = 6;
string sanitized_error = 7;
reserved 8, 9, 10;
reserved "created_at_ms", "updated_at_ms", "completed_at_ms";
google.protobuf.Timestamp created_time = 108;
google.protobuf.Timestamp updated_time = 109;
// Absent until the operation reaches a terminal state.
google.protobuf.Timestamp completed_time = 110;
}
// Immutable, secret-free record of one sandbox's intended provider mutation.
message ProviderMutationReceipt {
string receipt_id = 1;
// Shared by all sandbox receipts from one provider update.
string mutation_id = 2;
string provider_name = 3;
string workspace = 4;
ProviderMutationKind kind = 5;
ProviderDesiredIdentity desired = 6;
reserved 7;
reserved "persisted_at_ms";
google.protobuf.Timestamp persisted_time = 107;
}
// Installed state reported by the current supervisor. Process installation is
// acknowledged by the authenticated sandbox boundary after replacing its
// environment for future process launches.
message ProviderReadinessObservation {
// Gateway-issued identifier from the current ConnectSupervisor response.
string session_id = 1;
// Monotonic only within this connection; unrelated to revision fingerprints.
uint64 sequence = 2;
string attachment_epoch = 3;
uint64 provider_env_revision = 4;
uint64 config_revision = 5;
string policy_hash = 6;
bool credentials_installed = 7;
bool policy_active = 8;
bool launch_environment_installed = 9;
string process_instance_id = 10;
ProviderReadinessReason reason = 11;
}
// Operator view of desired and observed state; contains no credential material.
message ProviderReadinessStatus {
ProviderMutationReceipt receipt = 1;
ProviderReadinessState state = 2;
ProviderReadinessReason reason = 3;
ProviderReadinessObservation observed = 4;
string network_instance_id = 5;
reserved 6, 7;
reserved "observed_at_ms", "evaluated_at_ms";
// Absent until the current supervisor session supplies accepted evidence.
google.protobuf.Timestamp observed_time = 106;
google.protobuf.Timestamp evaluated_time = 107;
// Durable operation for the receipt, including its terminal apply outcome.
ConfigUpdateOperation operation = 8;
}
// Query an immutable receipt, or reconstruct the current desired state when
// receipt_id is empty. The sandbox identity must match the receipt.
message GetSandboxProviderStatusRequest {
string sandbox_name = 1;
string provider_name = 2;
string receipt_id = 3;
openshell.datamodel.v1.WorkspaceSelector workspace_scope = 4;
}
message GetSandboxProviderStatusResponse {
ProviderReadinessStatus status = 1;
}
// Installation evidence from the authenticated supervisor for this sandbox.
// A caller-provided instance identifier alone never establishes authority.
message ReportProviderReadinessRequest {
string sandbox_id = 1;
ProviderReadinessObservation observation = 2;
}
// Acknowledges accepted evidence without granting a separate session authority.
// Identical retries do not extend the evidence's original acceptance time.
message ReportProviderReadinessResponse {
uint64 accepted_sequence = 1;
reserved 2, 3;
reserved "report_interval_seconds", "observation_ttl_seconds";
google.protobuf.Duration report_interval = 102;
google.protobuf.Duration observation_ttl = 103;
}
// Delete sandbox response.
@@ -1846,6 +2079,11 @@ message DeleteProviderRequest {
// Provider response.
message ProviderResponse {
openshell.datamodel.v1.Provider provider = 1;
// Selection-time sandbox target set for an update, with one receipt per target.
// Sandboxes attached later are outside this operation's readiness result.
repeated ProviderMutationReceipt target_receipts = 2;
// Identifies the update even when its target set is empty.
string mutation_id = 3;
}
// List providers response.
@@ -2312,6 +2550,12 @@ message GetSandboxProviderEnvironmentResponse {
// Environment variables that contain provider configuration rather than
// credentials and therefore do not require endpoint-scoped resolution.
repeated string non_secret_environment_keys = 6;
// Attachment identity captured with the returned provider records.
string provider_attachment_epoch = 7;
// Effective policy identity used to derive this snapshot's endpoint bindings.
string policy_hash = 8;
// Nonzero when material was withheld; installing an empty map is not readiness.
ProviderReadinessReason readiness_reason = 9;
}
message ExchangeProviderSubjectTokenRequest {
@@ -2628,6 +2872,8 @@ message SupervisorHello {
string sandbox_id = 1;
// Supervisor instance ID (e.g. boot id or process epoch).
string instance_id = 2;
// The supervisor can report credential, policy, and launch-environment installation.
bool supports_provider_readiness = 4;
}
// Gateway accepts the supervisor session.
+3
View File
@@ -399,6 +399,9 @@ message GetSandboxConfigResponse {
// False also covers older gateways that do not advertise this capability;
// supervisors preserve their legacy unauthenticated connection behavior.
bool extension_authentication_enabled = 12;
// Gateway-owned attachment identity captured with this desired configuration.
// Compare for equality; reattachment invalidates previous installation evidence.
string provider_attachment_epoch = 14;
}
// Connection details for one operator-registered supervisor middleware service.
@@ -33,7 +33,12 @@ func TestConverterCoversAllProtoFields_SandboxSpec(t *testing.T) {
"tty": true,
}
assertAllFieldsCovered(t, (&pb.SandboxSpec{}).ProtoReflect().Descriptor(), handled, nil)
// The gateway owns this identity. Provider status exposes it through the
// raw API; callers must not supply it when constructing a sandbox spec.
skipped := fieldSet{
"provider_attachment_epoch": true,
}
assertAllFieldsCovered(t, (&pb.SandboxSpec{}).ProtoReflect().Descriptor(), handled, skipped)
}
func TestConverterCoversAllProtoFields_SandboxTemplate(t *testing.T) {
File diff suppressed because it is too large Load Diff
@@ -37,6 +37,7 @@ const (
OpenShell_ListSandboxProviders_FullMethodName = "/openshell.v1.OpenShell/ListSandboxProviders"
OpenShell_AttachSandboxProvider_FullMethodName = "/openshell.v1.OpenShell/AttachSandboxProvider"
OpenShell_DetachSandboxProvider_FullMethodName = "/openshell.v1.OpenShell/DetachSandboxProvider"
OpenShell_GetSandboxProviderStatus_FullMethodName = "/openshell.v1.OpenShell/GetSandboxProviderStatus"
OpenShell_DeleteSandbox_FullMethodName = "/openshell.v1.OpenShell/DeleteSandbox"
OpenShell_StopSandbox_FullMethodName = "/openshell.v1.OpenShell/StopSandbox"
OpenShell_StartSandbox_FullMethodName = "/openshell.v1.OpenShell/StartSandbox"
@@ -71,6 +72,7 @@ const (
OpenShell_ListSandboxPolicies_FullMethodName = "/openshell.v1.OpenShell/ListSandboxPolicies"
OpenShell_ReportPolicyStatus_FullMethodName = "/openshell.v1.OpenShell/ReportPolicyStatus"
OpenShell_ReportEndpointStatus_FullMethodName = "/openshell.v1.OpenShell/ReportEndpointStatus"
OpenShell_ReportProviderReadiness_FullMethodName = "/openshell.v1.OpenShell/ReportProviderReadiness"
OpenShell_GetSandboxProviderEnvironment_FullMethodName = "/openshell.v1.OpenShell/GetSandboxProviderEnvironment"
OpenShell_ExchangeProviderSubjectToken_FullMethodName = "/openshell.v1.OpenShell/ExchangeProviderSubjectToken"
OpenShell_GetSandboxLogs_FullMethodName = "/openshell.v1.OpenShell/GetSandboxLogs"
@@ -148,6 +150,8 @@ type OpenShellClient interface {
AttachSandboxProvider(ctx context.Context, in *AttachSandboxProviderRequest, opts ...grpc.CallOption) (*AttachSandboxProviderResponse, error)
// Detach a provider record from an existing sandbox.
DetachSandboxProvider(ctx context.Context, in *DetachSandboxProviderRequest, opts ...grpc.CallOption) (*DetachSandboxProviderResponse, error)
// Inspect the installed authority for one sandbox provider mutation.
GetSandboxProviderStatus(ctx context.Context, in *GetSandboxProviderStatusRequest, opts ...grpc.CallOption) (*GetSandboxProviderStatusResponse, error)
// Delete a sandbox by name.
DeleteSandbox(ctx context.Context, in *DeleteSandboxRequest, opts ...grpc.CallOption) (*DeleteSandboxResponse, error)
// Stop a sandbox while retaining its persistent state.
@@ -224,6 +228,9 @@ type OpenShellClient interface {
ReportPolicyStatus(ctx context.Context, in *ReportPolicyStatusRequest, opts ...grpc.CallOption) (*ReportPolicyStatusResponse, error)
// Replace the gateway's observed tool server endpoint status for one sandbox.
ReportEndpointStatus(ctx context.Context, in *ReportEndpointStatusRequest, opts ...grpc.CallOption) (*ReportEndpointStatusResponse, error)
// Report installed provider state for the current ConnectSupervisor session.
// Replacing or losing that session invalidates its observations.
ReportProviderReadiness(ctx context.Context, in *ReportProviderReadinessRequest, opts ...grpc.CallOption) (*ReportProviderReadinessResponse, error)
// Get provider environment for a sandbox (called by sandbox supervisor at startup).
GetSandboxProviderEnvironment(ctx context.Context, in *GetSandboxProviderEnvironmentRequest, opts ...grpc.CallOption) (*GetSandboxProviderEnvironmentResponse, error)
// Exchange a stored provider subject token for an intermediate token scoped
@@ -459,6 +466,16 @@ func (c *openShellClient) DetachSandboxProvider(ctx context.Context, in *DetachS
return out, nil
}
func (c *openShellClient) GetSandboxProviderStatus(ctx context.Context, in *GetSandboxProviderStatusRequest, opts ...grpc.CallOption) (*GetSandboxProviderStatusResponse, error) {
cOpts := append([]grpc.CallOption{grpc.StaticMethod()}, opts...)
out := new(GetSandboxProviderStatusResponse)
err := c.cc.Invoke(ctx, OpenShell_GetSandboxProviderStatus_FullMethodName, in, out, cOpts...)
if err != nil {
return nil, err
}
return out, nil
}
func (c *openShellClient) DeleteSandbox(ctx context.Context, in *DeleteSandboxRequest, opts ...grpc.CallOption) (*DeleteSandboxResponse, error) {
cOpts := append([]grpc.CallOption{grpc.StaticMethod()}, opts...)
out := new(DeleteSandboxResponse)
@@ -814,6 +831,16 @@ func (c *openShellClient) ReportEndpointStatus(ctx context.Context, in *ReportEn
return out, nil
}
func (c *openShellClient) ReportProviderReadiness(ctx context.Context, in *ReportProviderReadinessRequest, opts ...grpc.CallOption) (*ReportProviderReadinessResponse, error) {
cOpts := append([]grpc.CallOption{grpc.StaticMethod()}, opts...)
out := new(ReportProviderReadinessResponse)
err := c.cc.Invoke(ctx, OpenShell_ReportProviderReadiness_FullMethodName, in, out, cOpts...)
if err != nil {
return nil, err
}
return out, nil
}
func (c *openShellClient) GetSandboxProviderEnvironment(ctx context.Context, in *GetSandboxProviderEnvironmentRequest, opts ...grpc.CallOption) (*GetSandboxProviderEnvironmentResponse, error) {
cOpts := append([]grpc.CallOption{grpc.StaticMethod()}, opts...)
out := new(GetSandboxProviderEnvironmentResponse)
@@ -1150,6 +1177,8 @@ type OpenShellServer interface {
AttachSandboxProvider(context.Context, *AttachSandboxProviderRequest) (*AttachSandboxProviderResponse, error)
// Detach a provider record from an existing sandbox.
DetachSandboxProvider(context.Context, *DetachSandboxProviderRequest) (*DetachSandboxProviderResponse, error)
// Inspect the installed authority for one sandbox provider mutation.
GetSandboxProviderStatus(context.Context, *GetSandboxProviderStatusRequest) (*GetSandboxProviderStatusResponse, error)
// Delete a sandbox by name.
DeleteSandbox(context.Context, *DeleteSandboxRequest) (*DeleteSandboxResponse, error)
// Stop a sandbox while retaining its persistent state.
@@ -1226,6 +1255,9 @@ type OpenShellServer interface {
ReportPolicyStatus(context.Context, *ReportPolicyStatusRequest) (*ReportPolicyStatusResponse, error)
// Replace the gateway's observed tool server endpoint status for one sandbox.
ReportEndpointStatus(context.Context, *ReportEndpointStatusRequest) (*ReportEndpointStatusResponse, error)
// Report installed provider state for the current ConnectSupervisor session.
// Replacing or losing that session invalidates its observations.
ReportProviderReadiness(context.Context, *ReportProviderReadinessRequest) (*ReportProviderReadinessResponse, error)
// Get provider environment for a sandbox (called by sandbox supervisor at startup).
GetSandboxProviderEnvironment(context.Context, *GetSandboxProviderEnvironmentRequest) (*GetSandboxProviderEnvironmentResponse, error)
// Exchange a stored provider subject token for an intermediate token scoped
@@ -1363,6 +1395,9 @@ func (UnimplementedOpenShellServer) AttachSandboxProvider(context.Context, *Atta
func (UnimplementedOpenShellServer) DetachSandboxProvider(context.Context, *DetachSandboxProviderRequest) (*DetachSandboxProviderResponse, error) {
return nil, status.Error(codes.Unimplemented, "method DetachSandboxProvider not implemented")
}
func (UnimplementedOpenShellServer) GetSandboxProviderStatus(context.Context, *GetSandboxProviderStatusRequest) (*GetSandboxProviderStatusResponse, error) {
return nil, status.Error(codes.Unimplemented, "method GetSandboxProviderStatus not implemented")
}
func (UnimplementedOpenShellServer) DeleteSandbox(context.Context, *DeleteSandboxRequest) (*DeleteSandboxResponse, error) {
return nil, status.Error(codes.Unimplemented, "method DeleteSandbox not implemented")
}
@@ -1465,6 +1500,9 @@ func (UnimplementedOpenShellServer) ReportPolicyStatus(context.Context, *ReportP
func (UnimplementedOpenShellServer) ReportEndpointStatus(context.Context, *ReportEndpointStatusRequest) (*ReportEndpointStatusResponse, error) {
return nil, status.Error(codes.Unimplemented, "method ReportEndpointStatus not implemented")
}
func (UnimplementedOpenShellServer) ReportProviderReadiness(context.Context, *ReportProviderReadinessRequest) (*ReportProviderReadinessResponse, error) {
return nil, status.Error(codes.Unimplemented, "method ReportProviderReadiness not implemented")
}
func (UnimplementedOpenShellServer) GetSandboxProviderEnvironment(context.Context, *GetSandboxProviderEnvironmentRequest) (*GetSandboxProviderEnvironmentResponse, error) {
return nil, status.Error(codes.Unimplemented, "method GetSandboxProviderEnvironment not implemented")
}
@@ -1819,6 +1857,24 @@ func _OpenShell_DetachSandboxProvider_Handler(srv interface{}, ctx context.Conte
return interceptor(ctx, in, info, handler)
}
func _OpenShell_GetSandboxProviderStatus_Handler(srv interface{}, ctx context.Context, dec func(interface{}) error, interceptor grpc.UnaryServerInterceptor) (interface{}, error) {
in := new(GetSandboxProviderStatusRequest)
if err := dec(in); err != nil {
return nil, err
}
if interceptor == nil {
return srv.(OpenShellServer).GetSandboxProviderStatus(ctx, in)
}
info := &grpc.UnaryServerInfo{
Server: srv,
FullMethod: OpenShell_GetSandboxProviderStatus_FullMethodName,
}
handler := func(ctx context.Context, req interface{}) (interface{}, error) {
return srv.(OpenShellServer).GetSandboxProviderStatus(ctx, req.(*GetSandboxProviderStatusRequest))
}
return interceptor(ctx, in, info, handler)
}
func _OpenShell_DeleteSandbox_Handler(srv interface{}, ctx context.Context, dec func(interface{}) error, interceptor grpc.UnaryServerInterceptor) (interface{}, error) {
in := new(DeleteSandboxRequest)
if err := dec(in); err != nil {
@@ -2402,6 +2458,24 @@ func _OpenShell_ReportEndpointStatus_Handler(srv interface{}, ctx context.Contex
return interceptor(ctx, in, info, handler)
}
func _OpenShell_ReportProviderReadiness_Handler(srv interface{}, ctx context.Context, dec func(interface{}) error, interceptor grpc.UnaryServerInterceptor) (interface{}, error) {
in := new(ReportProviderReadinessRequest)
if err := dec(in); err != nil {
return nil, err
}
if interceptor == nil {
return srv.(OpenShellServer).ReportProviderReadiness(ctx, in)
}
info := &grpc.UnaryServerInfo{
Server: srv,
FullMethod: OpenShell_ReportProviderReadiness_FullMethodName,
}
handler := func(ctx context.Context, req interface{}) (interface{}, error) {
return srv.(OpenShellServer).ReportProviderReadiness(ctx, req.(*ReportProviderReadinessRequest))
}
return interceptor(ctx, in, info, handler)
}
func _OpenShell_GetSandboxProviderEnvironment_Handler(srv interface{}, ctx context.Context, dec func(interface{}) error, interceptor grpc.UnaryServerInterceptor) (interface{}, error) {
in := new(GetSandboxProviderEnvironmentRequest)
if err := dec(in); err != nil {
@@ -2911,6 +2985,10 @@ var OpenShell_ServiceDesc = grpc.ServiceDesc{
MethodName: "DetachSandboxProvider",
Handler: _OpenShell_DetachSandboxProvider_Handler,
},
{
MethodName: "GetSandboxProviderStatus",
Handler: _OpenShell_GetSandboxProviderStatus_Handler,
},
{
MethodName: "DeleteSandbox",
Handler: _OpenShell_DeleteSandbox_Handler,
@@ -3035,6 +3113,10 @@ var OpenShell_ServiceDesc = grpc.ServiceDesc{
MethodName: "ReportEndpointStatus",
Handler: _OpenShell_ReportEndpointStatus_Handler,
},
{
MethodName: "ReportProviderReadiness",
Handler: _OpenShell_ReportProviderReadiness_Handler,
},
{
MethodName: "GetSandboxProviderEnvironment",
Handler: _OpenShell_GetSandboxProviderEnvironment_Handler,
+15 -4
View File
@@ -1826,8 +1826,11 @@ type GetSandboxConfigResponse struct {
// False also covers older gateways that do not advertise this capability;
// supervisors preserve their legacy unauthenticated connection behavior.
ExtensionAuthenticationEnabled bool `protobuf:"varint,12,opt,name=extension_authentication_enabled,json=extensionAuthenticationEnabled,proto3" json:"extension_authentication_enabled,omitempty"`
unknownFields protoimpl.UnknownFields
sizeCache protoimpl.SizeCache
// Gateway-owned attachment identity captured with this desired configuration.
// Compare for equality; reattachment invalidates previous installation evidence.
ProviderAttachmentEpoch string `protobuf:"bytes,14,opt,name=provider_attachment_epoch,json=providerAttachmentEpoch,proto3" json:"provider_attachment_epoch,omitempty"`
unknownFields protoimpl.UnknownFields
sizeCache protoimpl.SizeCache
}
func (x *GetSandboxConfigResponse) Reset() {
@@ -1944,6 +1947,13 @@ func (x *GetSandboxConfigResponse) GetExtensionAuthenticationEnabled() bool {
return false
}
func (x *GetSandboxConfigResponse) GetProviderAttachmentEpoch() string {
if x != nil {
return x.ProviderAttachmentEpoch
}
return ""
}
// Connection details for one operator-registered supervisor middleware service.
// V1 supports plaintext and server-authenticated TLS gRPC.
type SupervisorMiddlewareService struct {
@@ -2207,7 +2217,7 @@ const file_sandbox_proto_rawDesc = "" +
"\x05value\"\x86\x01\n" +
"\x10EffectiveSetting\x128\n" +
"\x05value\x18\x01 \x01(\v2\".openshell.sandbox.v1.SettingValueR\x05value\x128\n" +
"\x05scope\x18\x02 \x01(\x0e2\".openshell.sandbox.v1.SettingScopeR\x05scope\"\xd1\x06\n" +
"\x05scope\x18\x02 \x01(\x0e2\".openshell.sandbox.v1.SettingScopeR\x05scope\"\x8d\a\n" +
"\x18GetSandboxConfigResponse\x12;\n" +
"\x06policy\x18\x01 \x01(\v2#.openshell.sandbox.v1.SandboxPolicyR\x06policy\x12\x18\n" +
"\aversion\x18\x02 \x01(\rR\aversion\x12\x1f\n" +
@@ -2222,7 +2232,8 @@ const file_sandbox_proto_rawDesc = "" +
"\tworkspace\x18\n" +
" \x01(\tR\tworkspace\x12C\n" +
"\x1epolicy_validation_failure_mode\x18\v \x01(\tR\x1bpolicyValidationFailureMode\x12H\n" +
" extension_authentication_enabled\x18\f \x01(\bR\x1eextensionAuthenticationEnabled\x1ac\n" +
" extension_authentication_enabled\x18\f \x01(\bR\x1eextensionAuthenticationEnabled\x12:\n" +
"\x19provider_attachment_epoch\x18\x0e \x01(\tR\x17providerAttachmentEpoch\x1ac\n" +
"\rSettingsEntry\x12\x10\n" +
"\x03key\x18\x01 \x01(\tR\x03key\x12<\n" +
"\x05value\x18\x02 \x01(\v2&.openshell.sandbox.v1.EffectiveSettingR\x05value:\x028\x01\"\xd2\x02\n" +
+9 -7
View File
@@ -62,21 +62,23 @@ binding error.
```bash
openshell sandbox provider list <sandbox>
openshell sandbox provider attach <sandbox> <provider>
openshell sandbox provider attach <sandbox> <provider> --wait --timeout 30
```
Launch a new process after attaching a provider so it inherits newly available
credential placeholders:
Save the change's `receipt_id` and use `openshell sandbox provider status <sandbox> <provider> --receipt <receipt-id> --wait --timeout 30` to check when it takes effect. Success confirms that the sandbox applied the credentials, policy, and environment for new processes. If the result is pending, failed, withheld, or superseded, inspect its reason before launching the client.
Launch the client after the attachment wait succeeds so it receives the updated environment:
```bash
openshell sandbox exec <sandbox> -- env
openshell sandbox exec <sandbox> -- <client-command>
```
Do not print or copy credential values into diagnostic output. Detaching a
provider revokes its policy and credential access:
After updating an ordinary static provider, wait for that change and launch a new client. An existing process keeps its revision-scoped reference; readiness does not make the old reference resolve the replacement value. Diagnose managed-refresh credentials according to their own lifecycle.
Keep credentials and issued references out of diagnostic output. Acknowledged detach revokes future credential resolution and removes the reference from future process environments. Requests already forwarded may still finish:
```bash
openshell sandbox provider detach <sandbox> <provider>
openshell sandbox provider detach <sandbox> <provider> --wait --timeout 30
```
### 4. Verify Native Client Configuration
+11 -6
View File
@@ -164,6 +164,10 @@ openshell provider profile import --file ./my-profile.yaml
### List, inspect, update, delete
Use `openshell sandbox provider status --help` and the attach, detach, and update help to find the installed version's wait options. Add `--wait` when the next step depends on a provider change taking effect. Without it, a successful command only confirms that the gateway saved the change. Save the returned `receipt_id` to check that same change later, and inspect the result for every selected sandbox. Credential refresh status confirms that OpenShell obtained credentials; provider status confirms that the sandbox applied them, activated the policy, and updated the environment for new processes. If the status is `superseded`, explain that a later change replaced the request and inspect that change separately.
If attach, detach, or update reports `CONFIG_OPERATION_STORAGE_UNCERTAIN`, explain that the change may already be saved and its readiness receipt may be unavailable. Do not blindly retry the mutation. Inspect the provider and sandbox state and reconcile the saved change before deciding on another mutation; the error proves neither rollback nor readiness.
```bash
openshell provider list
openshell provider list --output json
@@ -378,8 +382,9 @@ provider instead of passing API keys, tokens, or other secrets to `sandbox exec`
```bash
openshell sandbox provider list my-sandbox
openshell sandbox provider list my-sandbox --output json
openshell sandbox provider attach my-sandbox my-github
openshell sandbox provider detach my-sandbox my-github
openshell sandbox provider attach my-sandbox my-github --wait --timeout 30
openshell sandbox provider status my-sandbox my-github --output json
openshell sandbox provider detach my-sandbox my-github --wait --timeout 30
```
Structured attachment output contains provider names, types, and sorted
@@ -723,14 +728,14 @@ endpoint, create the provider, and attach it only to sandboxes that need it:
```bash
openshell provider profile import -f ./inference-provider.yaml
openshell provider create --name model-provider --type <profile-id> --credential <KEY>
openshell sandbox provider attach work-session model-provider
openshell sandbox provider attach work-session model-provider --wait --timeout 30
openshell sandbox exec work-session -- <client-command>
```
The application owns the native base URL, model, request shape, and timeout.
Launch a new process after attaching a provider so it inherits the provider
credential placeholder. Use the `debug-inference` skill for endpoint, policy,
credential-binding, or migration failures.
Launch a new process after attachment readiness so it inherits the installed provider environment. Use the `debug-inference` skill for endpoint, policy, credential-binding, or migration failures.
For an ordinary static provider update, wait for the update and launch a new client process to obtain the new reference. Do not claim that readiness updates the environment of an existing process or retargets its old reference. Acknowledged detach revokes retained references and removes them from future process environments.
## Workflow 8: Gateway Management