8 Commits
Author SHA1 Message Date
Seth Jennings 2493d415c2 feat(extensions)!: normalize protocol negotiation (#3352)
* feat(extensions)!: normalize protocol negotiation

Closes #3057

Introduce a shared extension handshake, enforce protocol and capability compatibility across extension families, and expose immutable negotiated snapshots through gateway info and the Go SDK.

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(credentials): fail fast on negotiation errors

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(extensions): validate gateway handshake metadata

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(go-sdk): re-export extension kind constants

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(extensions): fail fast on credential handshake rejection

Signed-off-by: Seth Jennings <sjenning@redhat.com>

---------

Signed-off-by: Seth Jennings <sjenning@redhat.com>
2026-09-21 13:19:39 -07:00
Drew Newberry d91b1999a0 feat(api)!: use sandbox names as canonical RPC references (#3272)
* feat(api)!: use sandbox names as canonical references

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(cli): update forward color fixture for workspace scope

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(supervisor): use sandbox names for settings lookup

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(e2e): use canonical sandbox request fields

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(e2e): use canonical sandbox receipt field

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(server): harden sandbox mutation handling

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(ci): update rebased sandbox references

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(api)!: standardize canonical entity references

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(api): codify protobuf API conventions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(api): preserve workspace selector semantics

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(api): restore workspace selector parity

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(api): preserve descriptive name fields

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(api): update e2e request fixtures

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(core): omit workspace selector during bootstrap

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* chore(api): remove proto convention checker

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* chore(api): refresh schema fingerprints after rebase

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(cli): use canonical provider receipt field

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-18 11:39:55 -07:00
Derek Carr d68b7069c3 refactor(proto)!: use well-known time types (#3113)
* refactor(proto)!: use well-known time types

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(proto): preserve time migration behavior

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(proto): preserve timestamp boundary semantics

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(proto): convert sandbox token expiry to timestamp

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(proto): preserve time compatibility semantics

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(proto): preserve exact endpoint and profile times

Signed-off-by: Derek Carr <decarr@redhat.com>

* test(e2e): use duration for interactive exec timeout

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(sdk-go)!: remove legacy profile duration fields

BREAKING CHANGE: Go provider profile callers must use RefreshBefore, MaxLifetime, and CacheTTL with ProfileDuration instead of the whole-second fields.

Signed-off-by: Derek Carr <decarr@redhat.com>

---------

Signed-off-by: Derek Carr <decarr@redhat.com>
2026-09-16 19:51:45 +00:00
Seth Jennings dbe36eaf85 fix(security): harden Vault credential transport (#3329)
Reject non-loopback plaintext Vault endpoints, disable redirects, and support private CA bundles without weakening hostname verification. Update Helm configuration, documentation, operator skills, and regression coverage for OSSR-002.

Signed-off-by: Seth Jennings <sjenning@redhat.com>
2026-09-15 19:56:09 +00:00
Simon Scatton 48c449d8c8 chore(deps): replace ring with AWS-LC (#3243)
* chore(deps): replace ring with AWS-LC

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* fix(lint): address warnings after dependency upgrades

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* fix(tls): limit provider initialization to reqwest clients

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

---------

Signed-off-by: Simon Scatton <sscatton@nvidia.com>
2026-09-09 17:57:42 +00:00
Simon Scatton 9505ca5ed1 chore: remove Bazel build support (#2840)
Signed-off-by: Simon Scatton <sscatton@nvidia.com>
2026-08-20 15:08:09 +00:00
Simon Scatton d85339d621 build(bazel): add credential driver targets (#2649)
* build(bazel): add credential driver targets

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

* fix(bazel): sync default CA root features

Signed-off-by: Simon Scatton <sscatton@nvidia.com>

---------

Signed-off-by: Simon Scatton <sscatton@nvidia.com>
2026-08-07 09:40:24 +00:00
5548405fcb feat(credentials): add provider credential storage drivers (#2437)
* feat(credentials): add provider credential storage drivers

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(credentials): harden credential update handling

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(credentials): harden credential driver security, correctness, and performance

Address review findings from the credential storage drivers PR:

- Route additional_credentials through the driver on refresh to prevent
  silent data loss for multi-credential providers (e.g. AWS STS)
- Clean up stored credential handles on CAS failure during refresh to
  prevent orphaned secrets in external backends
- Enforce namespace validation in the Kubernetes Secrets driver to
  prevent cross-namespace credential access when allow_reference_namespace
  is not enabled
- Cache Vault Kubernetes auth tokens with 80% TTL to avoid re-authenticating
  on every credential operation
- Parallelize resolve_credentials in all three drivers using try_join_all
  for faster sandbox startup
- Add existingSecret support for the KEK Secret to fix helm template/GitOps
  workflows where lookup returns empty and regenerates the key
- Document RBAC blast radius for the Kubernetes Secrets credential driver
  and recommend a dedicated namespace

Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com>
Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>

* fix(credentials): add optimistic concurrency, fix thundering herd, parallelize operations

Use resourceVersion optimistic concurrency with retry loop for K8s
Secret ownership checks to prevent TOCTOU races. Switch Vault token
cache from RwLock to Mutex with double-check pattern to prevent
thundering herd on cache miss. Parallelize credential store and delete
operations across independent keys using try_join_all.

Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com>
Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>

* fix(credentials): handle partial failures, add delete retry, consolidate cleanup

Replace try_join_all with join_all in credential store/delete operations
to handle partial failures — successfully-stored handles are cleaned up
when another key fails. Add retry loop with conflict detection to
db-credstore delete_credential, matching the K8s driver pattern.
Consolidate 4 manual cleanup_pre_stored_provider_credentials call sites
into a single error handler using an async block. Remove inconsistent
.trim() from db-credstore validate_handle_owner.

Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com>
Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>

* fix(credentials): fix retry loop guard and remove unprotected validation

Remove attempt-count guard from 409/Aborted match arms in retry loops
so the post-loop Status::aborted error is reachable after exhausting
retries. Previously, last-attempt conflicts fell through to the
catch-all error arm, producing misleading Status::unavailable errors.

Remove duplicate validation calls that ran after
prepare_provider_credential_update but outside the cleanup-protected
async block, which would leak pre-stored handles on failure.

Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com>
Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>

* fix(credentials): add workspace/provider UUID to credential backend paths

Include workspace and provider ID in credential backend object paths to ensure
cross-workspace uniqueness and prevent credential collision (GATOR-1806c9be-01).

- Updated credential driver proto to include workspace and provider_id fields
- Modified Vault driver to include workspace/provider_id in managed_secret_path
- Modified Kubernetes Secrets driver to include workspace/provider_id in
  credential_owner_id and managed_secret_name
- Updated all credential runtime calls to pass workspace/provider_id
- Updated tests to use the new signatures

This prevents two workspaces sharing the same external credential store from
colliding on provider names, which was a critical security issue (CWE-639).

* fix(credentials): preserve provider-level expiration for handle-backed credentials

Compute effective expiration from both provider and driver values using the
earliest non-zero timestamp and skip expired values before insertion
(GATOR-1806c9be-02).

- Modified resolve_provider_handles to check provider credential_expires_at_ms
- Skip expired credentials during resolution instead of returning them
- Use effective expiration (min of provider and driver) in resolution results
- Fix inference.rs to preserve earliest expiration when merging

This ensures handle-backed credentials respect the same expiration semantics
as inline credentials.

* fix(credentials): stage refresh changes under new handles before validation

Stage credential replacements under new immutable handles instead of reusing
existing handles to prevent overwriting committed values before validation/CAS
(GATOR-1806c9be-03).

- Stage credentials with empty existing_handles map to force new handle creation
- Validate and CAS before the new values are committed to backend storage
- Delete old handles only after successful CAS
- On CAS failure, delete only the newly staged handles
- This prevents CWE-362/CWE-367 race conditions where failed refreshes could
  still modify or delete the active credential

The fix ensures that a rejected refresh cannot modify the backend object still
referenced by the committed provider record.

* fix(credentials): add timeouts to credential driver RPCs

Apply configured timeouts to both startup capability negotiation and runtime
RPCs to prevent indefinite hangs (GATOR-1806c9be-05).

- Add DEFAULT_CREDENTIAL_DRIVER_RPC_TIMEOUT_SECS constant (30s)
- Apply timeout to GetCapabilities during startup connection
- Apply timeout to all runtime RPCs (store, delete, resolve)
- Use tokio::time::timeout to bound the entire GetCapabilities operation
  during startup, not just the socket connection
- Return contextual deadline errors on timeout

This prevents a faulty or overloaded driver from hanging gateway operations
indefinitely.

* fix(credentials): fix test to use consistent workspace/provider identity

The Kubernetes auth Vault resolve test was constructing a managed path with
test-workspace/test-provider-id but sending default/prov-123 in the request,
causing validation to reject the request (GATOR-18e32351-01).

- Update test to use test-workspace and test-provider-id in the request to
  match the logical_path construction
- This ensures the test exercises the intended code path and validates
  Kubernetes auth resolution properly

The test now passes and correctly validates identity enforcement.

* fix(credentials): use unique staging ID for refresh to avoid overwrites

Stage refresh replacements under genuinely distinct immutable handles using
a unique staging ID to prevent overwriting committed values (GATOR-1806c9be-03).

- Generate a unique staging ID using UUID for each refresh operation
- Use this staging ID when storing credentials instead of the real provider ID
- Pass the same staging ID during cleanup on failure to delete only staged objects
- This ensures deterministic paths (Vault) and object names (K8s) don't collide
  with the committed provider's credentials

The fix prevents failed refreshes from silently replacing active credentials
or breaking providers by deleting still-referenced backend objects.

* fix(credentials): wrap credential driver RPCs in local timeouts

Add local tokio::time::timeout wrappers around credential driver RPCs to
bound non-compliant or stalled UDS peers (GATOR-1806c9be-05).

- Wrap StoreCredential, DeleteCredential, and ResolveCredentials in local timeouts
- Return contextual deadline_exceeded errors when timeouts occur
- Keep existing gRPC timeout metadata for compliant implementations
- GetCapabilities during startup was already wrapped in previous commit

This ensures a faulty local driver cannot hang gateway operations indefinitely,
even if it accepts the connection but never responds to the RPC.

* fix(credentials): preserve ownership for staged refreshes

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(credentials): bound startup capability probe

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* test(provider): authenticate credential handler requests

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(ci): grant actions read to credential driver e2e

Signed-off-by: Seth Jennings <sjenning@redhat.com>

---------

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com>
Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>
Signed-off-by: Seth Jennings <sjenning@redhat.com>
Co-authored-by: Taylor Mutch <taylormutch@gmail.com>
Co-authored-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>
2026-08-05 16:29:41 +00:00