mirror of
https://github.com/NVIDIA/OpenShell.git
synced 2026-10-04 08:28:19 +08:00
* feat(credentials): add provider credential storage drivers Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(credentials): harden credential update handling Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(credentials): harden credential driver security, correctness, and performance Address review findings from the credential storage drivers PR: - Route additional_credentials through the driver on refresh to prevent silent data loss for multi-credential providers (e.g. AWS STS) - Clean up stored credential handles on CAS failure during refresh to prevent orphaned secrets in external backends - Enforce namespace validation in the Kubernetes Secrets driver to prevent cross-namespace credential access when allow_reference_namespace is not enabled - Cache Vault Kubernetes auth tokens with 80% TTL to avoid re-authenticating on every credential operation - Parallelize resolve_credentials in all three drivers using try_join_all for faster sandbox startup - Add existingSecret support for the KEK Secret to fix helm template/GitOps workflows where lookup returns empty and regenerates the key - Document RBAC blast radius for the Kubernetes Secrets credential driver and recommend a dedicated namespace Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com> Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> * fix(credentials): add optimistic concurrency, fix thundering herd, parallelize operations Use resourceVersion optimistic concurrency with retry loop for K8s Secret ownership checks to prevent TOCTOU races. Switch Vault token cache from RwLock to Mutex with double-check pattern to prevent thundering herd on cache miss. Parallelize credential store and delete operations across independent keys using try_join_all. Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com> Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> * fix(credentials): handle partial failures, add delete retry, consolidate cleanup Replace try_join_all with join_all in credential store/delete operations to handle partial failures — successfully-stored handles are cleaned up when another key fails. Add retry loop with conflict detection to db-credstore delete_credential, matching the K8s driver pattern. Consolidate 4 manual cleanup_pre_stored_provider_credentials call sites into a single error handler using an async block. Remove inconsistent .trim() from db-credstore validate_handle_owner. Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com> Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> * fix(credentials): fix retry loop guard and remove unprotected validation Remove attempt-count guard from 409/Aborted match arms in retry loops so the post-loop Status::aborted error is reachable after exhausting retries. Previously, last-attempt conflicts fell through to the catch-all error arm, producing misleading Status::unavailable errors. Remove duplicate validation calls that ran after prepare_provider_credential_update but outside the cleanup-protected async block, which would leak pre-stored handles on failure. Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com> Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> * fix(credentials): add workspace/provider UUID to credential backend paths Include workspace and provider ID in credential backend object paths to ensure cross-workspace uniqueness and prevent credential collision (GATOR-1806c9be-01). - Updated credential driver proto to include workspace and provider_id fields - Modified Vault driver to include workspace/provider_id in managed_secret_path - Modified Kubernetes Secrets driver to include workspace/provider_id in credential_owner_id and managed_secret_name - Updated all credential runtime calls to pass workspace/provider_id - Updated tests to use the new signatures This prevents two workspaces sharing the same external credential store from colliding on provider names, which was a critical security issue (CWE-639). * fix(credentials): preserve provider-level expiration for handle-backed credentials Compute effective expiration from both provider and driver values using the earliest non-zero timestamp and skip expired values before insertion (GATOR-1806c9be-02). - Modified resolve_provider_handles to check provider credential_expires_at_ms - Skip expired credentials during resolution instead of returning them - Use effective expiration (min of provider and driver) in resolution results - Fix inference.rs to preserve earliest expiration when merging This ensures handle-backed credentials respect the same expiration semantics as inline credentials. * fix(credentials): stage refresh changes under new handles before validation Stage credential replacements under new immutable handles instead of reusing existing handles to prevent overwriting committed values before validation/CAS (GATOR-1806c9be-03). - Stage credentials with empty existing_handles map to force new handle creation - Validate and CAS before the new values are committed to backend storage - Delete old handles only after successful CAS - On CAS failure, delete only the newly staged handles - This prevents CWE-362/CWE-367 race conditions where failed refreshes could still modify or delete the active credential The fix ensures that a rejected refresh cannot modify the backend object still referenced by the committed provider record. * fix(credentials): add timeouts to credential driver RPCs Apply configured timeouts to both startup capability negotiation and runtime RPCs to prevent indefinite hangs (GATOR-1806c9be-05). - Add DEFAULT_CREDENTIAL_DRIVER_RPC_TIMEOUT_SECS constant (30s) - Apply timeout to GetCapabilities during startup connection - Apply timeout to all runtime RPCs (store, delete, resolve) - Use tokio::time::timeout to bound the entire GetCapabilities operation during startup, not just the socket connection - Return contextual deadline errors on timeout This prevents a faulty or overloaded driver from hanging gateway operations indefinitely. * fix(credentials): fix test to use consistent workspace/provider identity The Kubernetes auth Vault resolve test was constructing a managed path with test-workspace/test-provider-id but sending default/prov-123 in the request, causing validation to reject the request (GATOR-18e32351-01). - Update test to use test-workspace and test-provider-id in the request to match the logical_path construction - This ensures the test exercises the intended code path and validates Kubernetes auth resolution properly The test now passes and correctly validates identity enforcement. * fix(credentials): use unique staging ID for refresh to avoid overwrites Stage refresh replacements under genuinely distinct immutable handles using a unique staging ID to prevent overwriting committed values (GATOR-1806c9be-03). - Generate a unique staging ID using UUID for each refresh operation - Use this staging ID when storing credentials instead of the real provider ID - Pass the same staging ID during cleanup on failure to delete only staged objects - This ensures deterministic paths (Vault) and object names (K8s) don't collide with the committed provider's credentials The fix prevents failed refreshes from silently replacing active credentials or breaking providers by deleting still-referenced backend objects. * fix(credentials): wrap credential driver RPCs in local timeouts Add local tokio::time::timeout wrappers around credential driver RPCs to bound non-compliant or stalled UDS peers (GATOR-1806c9be-05). - Wrap StoreCredential, DeleteCredential, and ResolveCredentials in local timeouts - Return contextual deadline_exceeded errors when timeouts occur - Keep existing gRPC timeout metadata for compliant implementations - GetCapabilities during startup was already wrapped in previous commit This ensures a faulty local driver cannot hang gateway operations indefinitely, even if it accepts the connection but never responds to the RPC. * fix(credentials): preserve ownership for staged refreshes Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(credentials): bound startup capability probe Signed-off-by: Seth Jennings <sjenning@redhat.com> * test(provider): authenticate credential handler requests Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(ci): grant actions read to credential driver e2e Signed-off-by: Seth Jennings <sjenning@redhat.com> --------- Signed-off-by: Taylor Mutch <taylormutch@gmail.com> Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com> Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> Signed-off-by: Seth Jennings <sjenning@redhat.com> Co-authored-by: Taylor Mutch <taylormutch@gmail.com> Co-authored-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>
156 lines
5.8 KiB
YAML
156 lines
5.8 KiB
YAML
# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
|
# SPDX-License-Identifier: Apache-2.0
|
|
|
|
# Local dev: builds gateway + supervisor images via tasks/scripts/docker-build-image.sh,
|
|
# which first stages Rust binaries natively on the host (using cargo / cargo-zigbuild
|
|
# when cross-compiling) and then builds the image from the prebuilt binary. This
|
|
# mirrors CI and is faster than compiling inside Docker on every rebuild because
|
|
# the host's cargo target cache and sccache are reused across iterations.
|
|
#
|
|
# Run from repo root:
|
|
# mise run helm:skaffold:dev
|
|
# mise run helm:skaffold:run
|
|
#
|
|
# See https://skaffold.dev/docs/deployers/helm/ (setValueTemplates, IMAGE_* fields).
|
|
apiVersion: skaffold/v4beta14
|
|
kind: Config
|
|
metadata:
|
|
name: openshell
|
|
build:
|
|
local:
|
|
push: false
|
|
tagPolicy:
|
|
gitCommit: {}
|
|
artifacts:
|
|
- image: openshell/gateway
|
|
context: ../../..
|
|
custom:
|
|
buildCommand: |
|
|
CONTAINER_ENGINE_TARGET=local-k8s-cluster \
|
|
IMAGE_NAME="${IMAGE%:*}" \
|
|
IMAGE_TAG="${IMAGE##*:}" \
|
|
tasks/scripts/docker-build-image.sh gateway
|
|
dependencies:
|
|
paths:
|
|
- Cargo.toml
|
|
- Cargo.lock
|
|
- crates/**
|
|
- proto/**
|
|
- deploy/docker/Dockerfile.gateway
|
|
- tasks/scripts/docker-build-image.sh
|
|
- tasks/scripts/stage-prebuilt-binaries.sh
|
|
- image: openshell/supervisor
|
|
context: ../../..
|
|
custom:
|
|
buildCommand: |
|
|
CONTAINER_ENGINE_TARGET=local-k8s-cluster \
|
|
IMAGE_NAME="${IMAGE%:*}" \
|
|
IMAGE_TAG="${IMAGE##*:}" \
|
|
tasks/scripts/docker-build-image.sh supervisor
|
|
dependencies:
|
|
paths:
|
|
- Cargo.toml
|
|
- Cargo.lock
|
|
- crates/**
|
|
- proto/**
|
|
- deploy/docker/Dockerfile.supervisor
|
|
- tasks/scripts/docker-build-image.sh
|
|
- tasks/scripts/stage-prebuilt-binaries.sh
|
|
deploy:
|
|
helm:
|
|
releases:
|
|
# cert-manager — comment this in and add values-cert-manager.yaml below
|
|
# when you want cert-manager to manage TLS while certgen handles JWT.
|
|
# Requires cert-manager CRDs to be installed in the cluster first.
|
|
#- name: cert-manager
|
|
# remoteChart: oci://quay.io/jetstack/charts/cert-manager
|
|
# version: v1.20.2
|
|
# namespace: cert-manager
|
|
# createNamespace: true
|
|
# setValues:
|
|
# crds.enabled: true
|
|
# Envoy Gateway — Kubernetes Gateway API implementation.
|
|
# Installs the Gateway API CRDs and the "eg" GatewayClass.
|
|
# Required when grpcRoute.enabled is true in the openshell release.
|
|
#- name: envoy-gateway
|
|
# remoteChart: oci://docker.io/envoyproxy/gateway-helm
|
|
# version: v1.7.2
|
|
# namespace: envoy-gateway-system
|
|
# createNamespace: true
|
|
# # wait ensures Gateway API CRDs are registered before the openshell
|
|
# # release attempts to create Gateway and HTTPRoute resources.
|
|
# wait: true
|
|
# SPIRE — installs SPIRE Server, Agent, Controller Manager, CSI Driver,
|
|
# and OIDC Discovery Provider using the SPIFFE hardened charts.
|
|
# Uncomment both releases and ci/values-spire.yaml below to use
|
|
# SPIFFE JWT-SVIDs for dynamic provider token grants.
|
|
#- name: spire-crds
|
|
# repo: https://spiffe.github.io/helm-charts-hardened/
|
|
# remoteChart: spire-crds
|
|
# version: 0.5.0
|
|
# namespace: spire
|
|
# createNamespace: true
|
|
# wait: true
|
|
#- name: spire
|
|
# repo: https://spiffe.github.io/helm-charts-hardened/
|
|
# remoteChart: spire
|
|
# version: 0.29.0
|
|
# namespace: spire
|
|
# createNamespace: true
|
|
# valuesFiles:
|
|
# - ci/values-spire-stack.yaml
|
|
# wait: true
|
|
- name: openshell
|
|
chartPath: .
|
|
namespace: openshell
|
|
createNamespace: true
|
|
valuesFiles:
|
|
- values.yaml
|
|
- ci/values-skaffold.yaml
|
|
# Add ci/values-cert-manager.yaml here (and uncomment the cert-manager
|
|
# release above) to switch TLS generation to cert-manager.
|
|
#- ci/values-cert-manager.yaml
|
|
# To enable OIDC with a local Keycloak instance, run the one-time
|
|
# setup task first, then uncomment the line below:
|
|
# mise run keycloak:k8s:setup
|
|
#- ci/values-keycloak.yaml
|
|
# To enable the Gateway API HTTPRoute (requires Envoy Gateway above):
|
|
#- ci/values-gateway.yaml
|
|
# To enable SPIFFE/SPIRE provider token grants (requires the
|
|
# spire-crds and spire releases above):
|
|
#- ci/values-spire.yaml
|
|
# To exercise the Kubernetes supervisor sidecar topology:
|
|
#- ci/values-sidecar.yaml
|
|
# To test multi-replica external PostgreSQL behavior:
|
|
#- ci/values-high-availability.yaml
|
|
setValueTemplates:
|
|
image.repository: '{{.IMAGE_REPO_openshell_gateway}}'
|
|
image.tag: '{{.IMAGE_TAG_openshell_gateway}}'
|
|
supervisor.image.repository: '{{.IMAGE_REPO_openshell_supervisor}}'
|
|
supervisor.image.tag: '{{.IMAGE_TAG_openshell_supervisor}}'
|
|
profiles:
|
|
- name: sidecar
|
|
patches:
|
|
- op: add
|
|
path: /deploy/helm/releases/0/valuesFiles/-
|
|
value: ci/values-sidecar.yaml
|
|
- name: sidecar-mtls
|
|
patches:
|
|
- op: add
|
|
path: /deploy/helm/releases/0/valuesFiles/-
|
|
value: ci/values-sidecar.yaml
|
|
- op: add
|
|
path: /deploy/helm/releases/0/setValues
|
|
value:
|
|
server.disableTls: "false"
|
|
- name: credential-driver-kubernetes-secrets
|
|
patches:
|
|
- op: add
|
|
path: /deploy/helm/releases/0/valuesFiles/-
|
|
value: ci/values-credential-driver-kubernetes-secrets.yaml
|
|
- name: credential-driver-vault
|
|
patches:
|
|
- op: add
|
|
path: /deploy/helm/releases/0/valuesFiles/-
|
|
value: ci/values-credential-driver-vault.yaml
|