* feat(ci): detect breaking protobuf changes
Compare the proto module against the PR or merge-group base and report Buf violations in Branch Checks. Add local reproduction and fixture coverage.
Closes#3794
Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
* fix(ci): pin protobuf check container image
Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
* fix(ci): qualify protobuf compatibility by release train
Signed-off-by: Simon Scatton <sscatton@nvidia.com>
* refactor(ci): reuse protobuf compatibility action
Signed-off-by: Simon Scatton <sscatton@nvidia.com>
* refactor(ci): run protobuf checks as a Nix app with one ref
Signed-off-by: Simon Scatton <sscatton@nvidia.com>
---------
Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
Signed-off-by: Simon Scatton <sscatton@nvidia.com>
Co-authored-by: Mrunal Patel <mrunalp@gmail.com>
* test(tmachine): add K3s conformance scenario
Signed-off-by: Simon Scatton <sscatton@nvidia.com>
* refactor(tmachine): use Helm values file for K3s installer
Signed-off-by: Simon Scatton <sscatton@nvidia.com>
* ci(tmachine): run K3s conformance in integration jobs
Signed-off-by: Simon Scatton <sscatton@nvidia.com>
* ci(tmachine): verify installer scripts and document version baseline
Signed-off-by: Simon Scatton <sscatton@nvidia.com>
---------
Signed-off-by: Simon Scatton <sscatton@nvidia.com>
* fix(policy): propose rules for unknown DNS hosts
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
* docs(policy): clarify synthetic DNS use across protocols
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
* fix(policy): harden unknown-host DNS observations
- Emit the policy_dns_ineligible denial for every unknown name and
report observation staging failures as DNS failure events.
- Refuse unknown names during fail-closed quarantine and after the
observation budget, now a quarter of each address family's pool.
- Pin transparent TCP to the mapping of the deciding policy generation
so a reload between DNS and authorization fails closed.
- Stop Docker workloads from inheriting host DNS search domains, which
let the first expanded short name claim an observation address.
- Share mechanistic draft polling in conformance, register
new-hostname-proposal in the installed suite, and update docs.
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
* test(policy): build policy DNS proxy tests on every target
The proxy tests name PolicyEndpointId, which proxy.rs imported only on
Linux, so the macOS test build failed. Import it for test builds too.
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
* fix(policy): name DNS queries and mapped hosts in OCSF denials
DNS denial and failure events attached port 53 to the queried name,
which read as a connection to that host. They now carry only the name.
Transparent TCP denials for a policy DNS address show the mapped
hostname and keep the synthetic address in dst_endpoint.ip.
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
---------
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
* fix(install): avoid installing incompatible docker snap
The work to land RFC-0012 added new restrictions when interacting with
Docker by setting `NoNewPrivs`. This prevents the `docker` snap from
transitioning its AppArmor profile from `snap.docker.dockerd` to
`docker-default` when it tries to launch a container. Thus, the `docker`
snap is currently incompatible with OpenShell.
This commit prevents `install.sh` from installing the `docker` snap
before installing the `openshell` snap, and instead requires the user to
install a non-snap Docker daemon before proceeding with installing the
snap. Systems without the `snap` command are unaffected, since they
install native packages without checking for the presence of Docker or
other compute providers.
Signed-off-by: Oliver Calder <oliver.calder@canonical.com>
* fix(snap): only snapd 2.76 for openshell snap since store installs work
Signed-off-by: Oliver Calder <oliver.calder@canonical.com>
* fixup! fix(install): avoid installing incompatible docker snap
Signed-off-by: Oliver Calder <oliver.calder@canonical.com>
* fixup! fix(snap): only snapd 2.76 for openshell snap since store installs work
Signed-off-by: Oliver Calder <oliver.calder@canonical.com>
---------
Signed-off-by: Oliver Calder <oliver.calder@canonical.com>
* fix(snap): update stale snap docs and tests
Previously, the `openshell` snap required the `docker` snap. Now, it
works with any Docker daemon running on the system. Furthermore, the
`snap-declaration` assertion on the `openshell` snap when installed from
the Snap Store causes the `openshell` snap to always connect to the
system `:docker` slot, rather than a slot provided by the `docker` snap.
This commit updates the documentation, including the `description` field in
`snapcraft.yaml`, to ensure that all information is correct and
up-to-date.
Additionally, some tests connected the `openshell:docker` plug to the
`docker` snap's `docker:docker-daemon` slot, which is inconsistent with
how the `openshell` snap operates when installed from the store. Update
those tests to connect to the system `:docker` slot as well.
Signed-off-by: Oliver Calder <oliver.calder@canonical.com>
* fix(snap): require snapd 2.76 for openshell snap
Signed-off-by: Oliver Calder <oliver.calder@canonical.com>
* fix(nix): align snap gateway reproducer timeout with release-canary
Signed-off-by: Oliver Calder <oliver.calder@canonical.com>
* fix(snap): require snapd 2.77 for openshell snap
Signed-off-by: Oliver Calder <oliver.calder@canonical.com>
* fixup! fix(snap): require snapd 2.77 for openshell snap
Signed-off-by: Oliver Calder <oliver.calder@canonical.com>
* feat(snap): install openshell snap via install.sh when snap available
Change `install.sh` to install the `openshell` snap by default when
snapd is installed on the host. This installs the snap from the
`latest/stable` channel, which should match the most up-to-date release
tag on github.
If `OPENSHELL_VERSION=dev` is set for `install.sh`, then it will install
the `openshell` snap from the `latest/edge` channel, which matches the
latest dev release available on github.
The `openshell` snap currently requires Docker in order to function. If
a Docker daemon is already installed on the system, it will be used by
the `openshell` snap. Otherwise, `install.sh` will install the `docker`
snap first, wait for the Docker daemon to be ready, and then install the
`openshell` snap.
Also, update the `release-canary.yml` to split the `ubuntu-snap` job
into `ubuntu-snap-system-docker` and `ubuntu-snap-provisions-docker`,
which test the two aforementioned scenarios. Previously, `ubuntu-snap`
manually installed a given snap artifact as built from CI, but with
these new jobs, it instead uses `install.sh` to install the published
`openshell` snap from the `latest/edge` track, thus matching the
behavior of the other release canary jobs.
Make corresponding changes to the `nix` guest reproducer, documentation,
and tests.
Signed-off-by: Oliver Calder <oliver.calder@canonical.com>
* fix(snap): configure local gateway authentication
The `openshell` snap runs the gateway as a systemd system service, which
runs as root. Thus, the mTLS certs are generated by root and stored in a
root-owned directory to which non-root users do not have access. For
this reason, the snap's `openshell-gateway-wrapper` script sets
`OPENSHELL_DISABLE_TLS=true`.
This commit ensures that the `openshell` snap's gateway allows
unauthenticated local access by writing a default `gateway.toml`
configuration file during the install hook, which runs after the snap is
first installed but before services are started. The config file
contains sets `allow_unauthenticated_users = true`.
Signed-off-by: Oliver Calder <oliver.calder@canonical.com>
* fix(install): configure snap gateway authentication
Recently, a new install hook was added which writes a default config
file for the `openshell` snap to allow unauthenticated local access to
the gateway. This is because the gateway service runs as root and the
mTLS certificates are not accessible to non-root users.
However, the `install.sh` script installs the `openshell` snap from the
snap store, and the published version may not yet have that new install
hook. Or, the user may already have the snap installed, in which case
the install hook does not run. In either case, we need `install.sh` to
ensure that the config file is written to set the gateway auth to
`allow_unauthenticated_users = true`, and then restart the openshell
gateway service.
Signed-off-by: Oliver Calder <oliver.calder@canonical.com>
* fixup! feat(snap): install openshell snap via install.sh when snap available
Signed-off-by: Oliver Calder <oliver.calder@canonical.com>
---------
Signed-off-by: Oliver Calder <oliver.calder@canonical.com>
* feat(credentials): add provider credential storage drivers
Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
* fix(credentials): harden credential update handling
Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
* fix(credentials): harden credential driver security, correctness, and performance
Address review findings from the credential storage drivers PR:
- Route additional_credentials through the driver on refresh to prevent
silent data loss for multi-credential providers (e.g. AWS STS)
- Clean up stored credential handles on CAS failure during refresh to
prevent orphaned secrets in external backends
- Enforce namespace validation in the Kubernetes Secrets driver to
prevent cross-namespace credential access when allow_reference_namespace
is not enabled
- Cache Vault Kubernetes auth tokens with 80% TTL to avoid re-authenticating
on every credential operation
- Parallelize resolve_credentials in all three drivers using try_join_all
for faster sandbox startup
- Add existingSecret support for the KEK Secret to fix helm template/GitOps
workflows where lookup returns empty and regenerates the key
- Document RBAC blast radius for the Kubernetes Secrets credential driver
and recommend a dedicated namespace
Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com>
Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>
* fix(credentials): add optimistic concurrency, fix thundering herd, parallelize operations
Use resourceVersion optimistic concurrency with retry loop for K8s
Secret ownership checks to prevent TOCTOU races. Switch Vault token
cache from RwLock to Mutex with double-check pattern to prevent
thundering herd on cache miss. Parallelize credential store and delete
operations across independent keys using try_join_all.
Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com>
Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>
* fix(credentials): handle partial failures, add delete retry, consolidate cleanup
Replace try_join_all with join_all in credential store/delete operations
to handle partial failures — successfully-stored handles are cleaned up
when another key fails. Add retry loop with conflict detection to
db-credstore delete_credential, matching the K8s driver pattern.
Consolidate 4 manual cleanup_pre_stored_provider_credentials call sites
into a single error handler using an async block. Remove inconsistent
.trim() from db-credstore validate_handle_owner.
Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com>
Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>
* fix(credentials): fix retry loop guard and remove unprotected validation
Remove attempt-count guard from 409/Aborted match arms in retry loops
so the post-loop Status::aborted error is reachable after exhausting
retries. Previously, last-attempt conflicts fell through to the
catch-all error arm, producing misleading Status::unavailable errors.
Remove duplicate validation calls that ran after
prepare_provider_credential_update but outside the cleanup-protected
async block, which would leak pre-stored handles on failure.
Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com>
Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>
* fix(credentials): add workspace/provider UUID to credential backend paths
Include workspace and provider ID in credential backend object paths to ensure
cross-workspace uniqueness and prevent credential collision (GATOR-1806c9be-01).
- Updated credential driver proto to include workspace and provider_id fields
- Modified Vault driver to include workspace/provider_id in managed_secret_path
- Modified Kubernetes Secrets driver to include workspace/provider_id in
credential_owner_id and managed_secret_name
- Updated all credential runtime calls to pass workspace/provider_id
- Updated tests to use the new signatures
This prevents two workspaces sharing the same external credential store from
colliding on provider names, which was a critical security issue (CWE-639).
* fix(credentials): preserve provider-level expiration for handle-backed credentials
Compute effective expiration from both provider and driver values using the
earliest non-zero timestamp and skip expired values before insertion
(GATOR-1806c9be-02).
- Modified resolve_provider_handles to check provider credential_expires_at_ms
- Skip expired credentials during resolution instead of returning them
- Use effective expiration (min of provider and driver) in resolution results
- Fix inference.rs to preserve earliest expiration when merging
This ensures handle-backed credentials respect the same expiration semantics
as inline credentials.
* fix(credentials): stage refresh changes under new handles before validation
Stage credential replacements under new immutable handles instead of reusing
existing handles to prevent overwriting committed values before validation/CAS
(GATOR-1806c9be-03).
- Stage credentials with empty existing_handles map to force new handle creation
- Validate and CAS before the new values are committed to backend storage
- Delete old handles only after successful CAS
- On CAS failure, delete only the newly staged handles
- This prevents CWE-362/CWE-367 race conditions where failed refreshes could
still modify or delete the active credential
The fix ensures that a rejected refresh cannot modify the backend object still
referenced by the committed provider record.
* fix(credentials): add timeouts to credential driver RPCs
Apply configured timeouts to both startup capability negotiation and runtime
RPCs to prevent indefinite hangs (GATOR-1806c9be-05).
- Add DEFAULT_CREDENTIAL_DRIVER_RPC_TIMEOUT_SECS constant (30s)
- Apply timeout to GetCapabilities during startup connection
- Apply timeout to all runtime RPCs (store, delete, resolve)
- Use tokio::time::timeout to bound the entire GetCapabilities operation
during startup, not just the socket connection
- Return contextual deadline errors on timeout
This prevents a faulty or overloaded driver from hanging gateway operations
indefinitely.
* fix(credentials): fix test to use consistent workspace/provider identity
The Kubernetes auth Vault resolve test was constructing a managed path with
test-workspace/test-provider-id but sending default/prov-123 in the request,
causing validation to reject the request (GATOR-18e32351-01).
- Update test to use test-workspace and test-provider-id in the request to
match the logical_path construction
- This ensures the test exercises the intended code path and validates
Kubernetes auth resolution properly
The test now passes and correctly validates identity enforcement.
* fix(credentials): use unique staging ID for refresh to avoid overwrites
Stage refresh replacements under genuinely distinct immutable handles using
a unique staging ID to prevent overwriting committed values (GATOR-1806c9be-03).
- Generate a unique staging ID using UUID for each refresh operation
- Use this staging ID when storing credentials instead of the real provider ID
- Pass the same staging ID during cleanup on failure to delete only staged objects
- This ensures deterministic paths (Vault) and object names (K8s) don't collide
with the committed provider's credentials
The fix prevents failed refreshes from silently replacing active credentials
or breaking providers by deleting still-referenced backend objects.
* fix(credentials): wrap credential driver RPCs in local timeouts
Add local tokio::time::timeout wrappers around credential driver RPCs to
bound non-compliant or stalled UDS peers (GATOR-1806c9be-05).
- Wrap StoreCredential, DeleteCredential, and ResolveCredentials in local timeouts
- Return contextual deadline_exceeded errors when timeouts occur
- Keep existing gRPC timeout metadata for compliant implementations
- GetCapabilities during startup was already wrapped in previous commit
This ensures a faulty local driver cannot hang gateway operations indefinitely,
even if it accepts the connection but never responds to the RPC.
* fix(credentials): preserve ownership for staged refreshes
Signed-off-by: Seth Jennings <sjenning@redhat.com>
* fix(credentials): bound startup capability probe
Signed-off-by: Seth Jennings <sjenning@redhat.com>
* test(provider): authenticate credential handler requests
Signed-off-by: Seth Jennings <sjenning@redhat.com>
* fix(ci): grant actions read to credential driver e2e
Signed-off-by: Seth Jennings <sjenning@redhat.com>
---------
Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com>
Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>
Signed-off-by: Seth Jennings <sjenning@redhat.com>
Co-authored-by: Taylor Mutch <taylormutch@gmail.com>
Co-authored-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>
* test(e2e): run VM suite in CI
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* fix(ci): configure KVM permissions directly
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* fix(e2e): flush VM overlay before restart
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* docs: simplify VM test documentation
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* test(e2e): include gateway resume in VM run
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
---------
Signed-off-by: Drew Newberry <anewberry@nvidia.com>