mirror of
https://github.com/NVIDIA/OpenShell.git
synced 2026-10-03 07:58:25 +08:00
codex/mxc-https-test-socket-owner
94
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
dbe36eaf85 |
fix(security): harden Vault credential transport (#3329)
Reject non-loopback plaintext Vault endpoints, disable redirects, and support private CA bundles without weakening hostname verification. Update Helm configuration, documentation, operator skills, and regression coverage for OSSR-002. Signed-off-by: Seth Jennings <sjenning@redhat.com> |
||
|
|
b799fccb8b |
fix(auth): harden OIDC trust root retrieval (#3332)
* fix(auth): harden OIDC trust root retrieval Signed-off-by: Mrunal Patel <mrunalp@gmail.com> * fix(e2e): pass OIDC HTTP acknowledgement value Signed-off-by: Mrunal Patel <mrunalp@gmail.com> --------- Signed-off-by: Mrunal Patel <mrunalp@gmail.com> |
||
|
|
cc780d4e17 |
feat(helm): add BackendTLSPolicy support (#2728)
* feat(helm): add optional BackendTLSPolicy for e2e TLS Add grpcRoute.backendTLSPolicy values to optionally create a BackendTLSPolicy resource that enables end-to-end TLS between the Gateway proxy and the OpenShell gateway pod. The Gateway proxy terminates client-facing TLS and re-encrypts when connecting to the backend, validating the pod's certificate against a user-supplied CA ConfigMap. This removes the requirement to set server.disableTls=true when using HTTPS at the Gateway listener. Supported on OpenShift 4.22+ and other platforms with BackendTLSPolicy support in the Gateway API implementation. Update OpenShift and ingress documentation with e2e TLS instructions. Signed-off-by: Brandon Squizzato <bsquizza@redhat.com> * feat(helm,server): auto-create backend CA ConfigMap in certgen hook Extend the generate-certs command with --backend-ca-configmap-name and --backend-ca-source-secret flags. When BackendTLSPolicy is enabled, the certgen pre-install hook creates the CA ConfigMap automatically: - pkiInitJob mode (default): uses the CA from the generated PKI bundle. Fully automatic on first install. - cert-manager mode: reads ca.crt from the server TLS Secret. On first install the Secret does not exist yet (cert-manager reconciles after templates are applied), so the ConfigMap is created on the first helm upgrade. Logs a warning on the initial skip. The caCertificateConfigMapName value now defaults to <fullname>-backend-ca when empty, so users only need to set backendTLSPolicy.enabled=true. Update certgen RBAC to include configmaps get/create when the feature is enabled. Add CLI arg parsing tests for the new flags. Signed-off-by: Brandon Squizzato <bsquizza@redhat.com> * refactor(helm): add server.tls.enableMtls flag for mTLS control Replace automatic mTLS disabling based on BackendTLSPolicy with an explicit server.tls.enableMtls flag that defaults to true. The user is now responsible for setting this to false when using BackendTLSPolicy, as ingress proxies cannot present client certificates to backends. Updated: - values.yaml: Added server.tls.enableMtls (default true) - gateway-config.yaml: Check enableMtls instead of backendTLSPolicy - _gateway-workload.tpl: Check enableMtls for client CA mount - Tests: Updated to use enableMtls flag - Docs: Added enableMtls=false to BackendTLSPolicy examples - README: Document new flag and BackendTLSPolicy requirement Signed-off-by: Brandon Squizzato <bsquizza@redhat.com> * docs(helm): clarify cert-manager backend CA ConfigMap workflow Update documentation to explain the two-step install process required when using cert-manager with BackendTLSPolicy: 1. helm install - cert-manager issues the server certificate, but the certgen hook can't create the backend CA ConfigMap yet (cert-manager reconciles after templates are applied) 2. helm upgrade - certgen hook reads the CA from the cert-manager-issued certificate and creates the ConfigMap Previously, the docs said "created on first upgrade" without explaining why or that the feature won't work until then. The updated docs now: - Explain the timing issue (cert-manager reconciles after chart install) - Provide clear steps for the cert-manager workflow - Note that pkiInitJob (default) creates it immediately on install - Clarify that users must wait for the Certificate to be Ready before running the second upgrade Signed-off-by: Brandon Squizzato <bsquizza@redhat.com> * fix(docs): remove incorrect external hostname requirement for BackendTLSPolicy BackendTLSPolicy validates the backend certificate against the service FQDN (e.g., openshell.openshell.svc.cluster.local), not the external hostname. The external hostname only needs to be on the Gateway listener certificate for client-facing TLS. The default certManager.serverDnsNames already includes all required service FQDN variants, so no configuration is needed for BackendTLSPolicy to work. Fixed incorrect documentation that claimed: - "The server certificate SAN list must include the external hostname" - Users need to "configure certManager.serverDnsNames with the external hostname" Removed the unnecessary pkiInitJob.serverDnsNames override from the example and clarified that: - Gateway listener certificate needs the external hostname (for clients) - Backend certificate needs the service FQDN (for Gateway proxy) - The service FQDN is already in the defaults Signed-off-by: Brandon Squizzato <bsquizza@redhat.com> * docs: clarify ACME with LetsEncrypt reference Change all references from "ACME issuer" to "LetsEncrypt/ACME issuer" to help users understand that LetsEncrypt is the most common ACME provider and what ACME means in practice. Updated: - docs/kubernetes/managing-certificates.mdx - docs/kubernetes/openshift.mdx - deploy/helm/openshell/values.yaml - deploy/helm/openshell/README.md - deploy/helm/openshell/ci/values-openshift-route-cert-manager.yaml Signed-off-by: Brandon Squizzato <bsquizza@redhat.com> * docs(openshift): restructure end-to-end TLS options and clarify Gateway hostname Reorganize the OpenShift production deployment documentation: 1. Changed main section from "Production Deployments" to "Options for end-to-end TLS" for better clarity 2. Renamed subsections for consistency and clarity: - "End-to-end TLS using Gateway API and BackendTLSPolicy (OpenShift 4.22+)" - "End-to-end TLS using pass-through Route (all OpenShift versions)" 3. Clarified that the Gateway hostname is typically a wildcard: "typically a wildcard like *.openshell-ingress-gw.example.com" 4. Removed the recommendation to copy the cluster's wildcard certificate from openshift-ingress namespace, as this is not a recommended security best practice These changes make it clearer that users have two end-to-end TLS options and help them understand the typical naming pattern for Gateway hostnames. Signed-off-by: Brandon Squizzato <bsquizza@redhat.com> * feat(helm): eliminate two-stage install for BackendTLSPolicy with cert-manager When using BackendTLSPolicy with cert-manager, the certgen hook now polls for up to 90 seconds waiting for cert-manager to issue the TLS certificate before creating the backend CA ConfigMap. This eliminates the need for a second `helm upgrade` in most cases. The hook polls every 2 seconds with progress logging every 10 seconds. If cert-manager takes longer than 90 seconds, the hook times out gracefully and logs a warning, preserving the fallback to manual ConfigMap creation or a second upgrade. The Job's activeDeadlineSeconds is 120s, so the 90s timeout leaves 30s margin for ConfigMap creation and hook completion. Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com> * feat(helm): add configurable timeout for certgen hook Add `pkiInitJob.timeoutSeconds` Helm value (default 120) to control how long the certgen hook Job can run. When using cert-manager with BackendTLSPolicy, the hook polls for (timeoutSeconds - 30) seconds to leave margin for ConfigMap creation and cleanup. This allows users to increase the timeout for environments where cert-manager takes longer than 90 seconds to issue certificates, without requiring code changes. Example usage: ```yaml pkiInitJob: timeoutSeconds: 180 # Hook polls for 150 seconds ``` Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com> * docs(helm): document configurable certgen timeout Update documentation to mention the pkiInitJob.timeoutSeconds value and how it affects the cert-manager polling behavior when using BackendTLSPolicy. Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com> * feat(helm): add configurable failure behavior for certgen timeout Add `pkiInitJob.failOnTimeout` Helm value (default false) to control whether the certgen hook fails or succeeds when cert-manager does not issue a certificate within the polling timeout. When false (default), the hook succeeds with a warning and users can run `helm upgrade` after cert-manager issues the certificate to create the backend CA ConfigMap. This provides backwards-compatible behavior. When true, the hook fails immediately if the timeout is reached, providing clear feedback that BackendTLSPolicy is non-functional. This is useful for strict validation requirements where incomplete installs should fail fast. Example usage: ```yaml pkiInitJob: timeoutSeconds: 180 failOnTimeout: true # Fail install if cert-manager takes >150s ``` Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com> * feat(helm): change failOnTimeout default to true and add troubleshooting docs Change `pkiInitJob.failOnTimeout` default from false to true to provide immediate feedback when cert-manager does not issue certificates within the polling timeout. This prevents silent failures where BackendTLSPolicy is non-functional but the install appears to succeed. Add comprehensive troubleshooting section to docs/kubernetes/ingress.mdx documenting the specific error "TLS error: Secret is not supplied by SDS" that occurs when the backend CA ConfigMap is missing, with step-by-step resolution instructions. Updated comments in values.yaml to clearly document the default behavior and explain when administrators might see connectivity errors if they override the default to failOnTimeout=false. BREAKING CHANGE: pkiInitJob.failOnTimeout now defaults to true. Helm installs will fail if cert-manager takes longer than (timeoutSeconds - 30) seconds to issue certificates. To restore the old behavior of allowing installs to succeed with a warning, set `pkiInitJob.failOnTimeout=false`. Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com> * fix(helm): make cert-manager resources pre-install hooks to fix ordering Make Certificate and Issuer resources run as pre-install/pre-upgrade hooks with weight -30, before the certgen hook (weight -20). This fixes the chicken-and-egg problem where the certgen hook was waiting for Secrets created by Certificates that hadn't been created yet. **Hook ordering:** 1. Certificate and Issuer resources created (weight -30) 2. cert-manager issues certificates and creates Secrets 3. certgen hook runs (weight -20), finds Secrets, creates ConfigMap 4. Main resources (StatefulSet, Service, etc.) created Previously, the certgen pre-install hook would run before any resources were created, poll for a non-existent Secret, timeout, and fail. The Certificate resources would never get created because Helm waits for all pre-install hooks to succeed before creating main resources. This fix allows single-stage installs to work reliably as long as cert-manager can issue certificates within the polling timeout. Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com> * feat(helm): add validation to prevent enableMtls with BackendTLSPolicy Add Helm chart validation that fails the install if both server.tls.enableMtls=true and grpcRoute.backendTLSPolicy.enabled=true are set, since this is an invalid configuration. BackendTLSPolicy requires mTLS to be disabled because the Gateway proxy cannot present client certificates to the backend. This validation provides immediate, clear feedback at install time rather than allowing the misconfiguration to be discovered through runtime errors. Example error message: ``` Error: grpcRoute.backendTLSPolicy requires mTLS to be disabled because the Gateway proxy cannot present client certificates to the backend; set server.tls.enableMtls=false ``` Also updated documentation to mention this validation check. Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com> * docs(helm): clarify pkiInitJob.timeoutSeconds polling behavior Improve documentation to clearly explain that pkiInitJob.timeoutSeconds controls the Job deadline, but the actual polling timeout is (timeoutSeconds - 30) to reserve 30 seconds for ConfigMap creation and cleanup. Added concrete example: "timeoutSeconds=180 allows 150 seconds of polling" to make the relationship explicit and avoid confusion where users might expect the hook to poll for the full timeout value. Updated both values.yaml inline comments and ingress.mdx documentation for consistency. Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com> * docs(openshift): remove outdated two-stage install instructions Update OpenShift documentation to reflect that single-stage installs now work with cert-manager and BackendTLSPolicy. The Certificate resources run as pre-install hooks (weight -30) before certgen (weight -20), allowing the hook to poll for and find the issued certificates. Removed the outdated two-step process: 1. helm install (cert-manager issues cert, hook logs warning) 2. helm upgrade (hook creates ConfigMap) Replaced with current single-stage behavior: - Certificate resources created as pre-install hooks - certgen hook polls for up to 90 seconds (configurable) - Single helm install succeeds in most cases - Fails fast by default if timeout reached This brings openshift.mdx in line with the already-updated ingress.mdx documentation. Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com> * fix(helm): make pkiInitJob.timeoutSeconds the actual polling duration The timeout value now represents the actual polling time that users experience when waiting for cert-manager to issue certificates. The Job activeDeadlineSeconds is set to (timeoutSeconds + 30) to allow buffer time for ConfigMap creation and cleanup. Previously, the hook polled for (timeoutSeconds - 30) seconds, which was confusing when users set timeoutSeconds=180 and only got 150 seconds of actual polling. Updated documentation in values.yaml, ingress.mdx, and openshift.mdx to reflect the clearer behavior. Signed-off-by: Brandon Squizzato <bsquizza@nvidia.com> * docs(helm): update values.yaml and README with correct polling duration Updated the caCertificateConfigMapName description to reflect that the hook polls for exactly pkiInitJob.timeoutSeconds seconds, not (timeoutSeconds - 30) seconds. Regenerated README.md with helm-docs. Signed-off-by: Brandon Squizzato <bsquizza@nvidia.com> * docs(kubernetes): add OIDC configuration to helm install and CLI examples Updated all helm install and openshell gateway add examples in ingress.mdx and openshift.mdx to include OIDC issuer and audience configuration. Examples now use concrete placeholder values: - OIDC issuer: https://keycloak.example.com/realms/openshell - OIDC audience: openshell-cli - Hostname: gateway.example.com - ClusterIssuer: letsencrypt-prod This makes it clearer how to configure OIDC authentication, which is required when using BackendTLSPolicy or HTTPS termination since the Gateway proxy cannot present client certificates to the backend. Signed-off-by: Brandon Squizzato <bsquizza@nvidia.com> * docs(kubernetes): explicitly list OIDC client ID in gateway add examples Added --oidc-client-id openshell-cli to all openshell gateway add commands in ingress.mdx and openshift.mdx, making the default client ID explicit in the examples even though it's the CLI default. This improves clarity and helps users understand the complete OIDC configuration needed for gateway registration. Signed-off-by: Brandon Squizzato <bsquizza@nvidia.com> * fix(helm): address PR review feedback for BackendTLSPolicy - Read backend CA from the authoritative server Secret instead of the in-memory PKI bundle so enabling BackendTLSPolicy on an existing release uses the CA that actually signed the server certificate. - Reconcile the backend CA ConfigMap on every hook run (compare and update) instead of skipping when it already exists, so CA rotations propagate automatically. - Remove hook annotations from cert-manager Issuer/Certificate resources so they remain regular release objects managed by Helm lifecycle. Split the cert-manager backend CA ConfigMap creation into a separate post-install/post-upgrade hook Job that polls after cert-manager Certificate resources are applied. - Update architecture/gateway.md, docs/reference/gateway-config.mdx, debug-openshell-cluster skill, and helm-dev-environment skill with BackendTLSPolicy, backend CA ConfigMap, enableMtls, and timeout documentation. Signed-off-by: Brandon Squizzato <bsquizza@nvidia.com> Signed-off-by: Brandon Squizzato <bsquizza@redhat.com> * fix(docs): resolve markdown lint errors in helm README and kubernetes docs Escape inline HTML angle brackets in README.md template placeholders, remove trailing spaces, and add blank lines around fenced code blocks in numbered lists. Signed-off-by: Brandon Squizzato <bsquizza@redhat.com> * Update docs * fix(helm): escape inline HTML in values.yaml descriptions and sync mise lockfile Wrap `<fullname>` and `<namespace>` template placeholders in backticks so markdownlint does not flag them as inline HTML (MD033). Regenerate mise.lock to match current mise.toml after rebase onto main. Signed-off-by: Brandon Squizzato <bsquizza@redhat.com> * Run 'mise lock' * fix: align mise.lock with CI mise version output The lockfile was regenerated locally with mise 2026.8.10 which resolves uv Linux artifacts to gnu variants and adds provenance_verified fields, but CI uses v2026.4.25 which produces musl variants without those fields. Signed-off-by: Brandon Squizzato <bsquizza@redhat.com> * fix(docs): correct cert-manager hook ordering and clientCaSecretName comment Update ingress.mdx and openshift.mdx to describe Certificate resources as regular release objects with a post-install/post-upgrade Job, matching the current implementation and architecture/gateway.md. Fix values.yaml clientCaSecretName comment to state that "" disables client certificate verification, matching the helper and access-control docs. Signed-off-by: Brandon Squizzato <bsquizza@redhat.com> --------- Signed-off-by: Brandon Squizzato <bsquizza@redhat.com> Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com> Signed-off-by: Brandon Squizzato <bsquizza@nvidia.com> |
||
|
|
99e83a535e |
feat(helm): scope ClusterRole/ClusterRoleBinding names by release namespace (#2939)
* feat(helm): scope ClusterRole/ClusterRoleBinding names by release namespace The chart creates cluster-scoped ClusterRole and ClusterRoleBinding resources with a fixed name derived from the release name. When multiple Helm releases coexist on the same cluster (multi-tenant), only one release can own these resources due to Helm ownership annotations -- the second install fails with a conflict. Append .Release.Namespace to the ClusterRole and ClusterRoleBinding names so each release gets its own cluster-scoped resources. The duplication is harmless (the rules are identical and small) and eliminates multi-tenant conflicts entirely without requiring external RBAC management. Signed-off-by: Brandon Squizzato <bsquizza@redhat.com> * test(helm): add regression tests for namespace-scoped ClusterRole names Assert the generated ClusterRole name, ClusterRoleBinding name, and roleRef all include the release namespace suffix so multi-namespace installations cannot silently regress to conflicting fixed names. Signed-off-by: Brandon Squizzato <bsquizza@redhat.com> --------- Signed-off-by: Brandon Squizzato <bsquizza@redhat.com> |
||
|
|
02b664bb0d |
refactor(config): normalize and enforce gateway schema v2 (#2814)
* refactor(config): normalize compute driver field names Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * refactor(config): introduce canonical gateway fields Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * refactor(config): enforce gateway schema version 2 Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): preserve compute driver runtime guarantees Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): address schema v2 review regressions Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): complete schema v2 migration safeguards Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(config): expand schema v2 regression coverage Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(config): add schema v2 parity manifest Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): correct parity manifest inventory Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * docs(config): record schema v2 intentional changes Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * docs(config): disposition schema v2 parity gaps Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): add dual schema parity harness Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): establish compute lifecycle parity baseline Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): preserve gateway option compatibility Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): record gateway option parity Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * docs(config): close gateway-wide parity gaps Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(podman): apply configured pids limit Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): validate Podman option parity Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): add Kubernetes option parity harness Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): record Kubernetes option parity Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): disposition VM parity lanes Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): add external driver parity lane Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(e2e): preserve external driver pull policy Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): attest parity artifacts and launches Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): require clean parity build sources Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): bind parity runtime artifacts Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(e2e): use isolated supervisor tags Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(e2e): qualify parity image tags Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(e2e): serve parity supervisor locally Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): isolate parity podman services Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): harden parity evidence provenance Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): pin parity sandbox artifacts Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): attest parity runtime inputs Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): bind parity runtime evidence Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): record compute boundary parity Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): disposition cross-cutting parity lanes Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(packaging): preflight gateway config upgrades Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): preserve rebase integration guarantees Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(ci): isolate temporary git signing config Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): update remaining schema v2 consumers Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(ci): provide e2fs tools to VM tests Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): align preflight with gateway startup Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(vm): preserve rootfs tar configuration Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * chore(config): adopt duration unit constructors Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(packaging): preflight RPM gateway config Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): address driver review findings Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(e2e): require fresh semantic parity evidence Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(docker): update tests for renamed sandbox label Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(gateway): preserve selective driver coverage after rebase Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> Signed-off-by: Drew Newberry <anewberry@nvidia.com> Co-authored-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
519e5eb35f |
feat(e2e): make e2e:kubernetes work transparently on OpenShift (#3183)
* feat(e2e): make e2e:kubernetes work transparently on OpenShift
Running `mise run e2e:kubernetes` on OpenShift required manual namespace
creation, SCC grants, Helm value overrides, and cleanup. A separate
`e2e:openshift` task existed but only checked pod readiness without
running the Rust e2e test suite, and even with the suite wired up the
SSH-relay `sandbox connect` path stalled to the ready timeout because
`kubectl port-forward` cannot carry round-trip-heavy SSH over the
internet.
The harness now auto-detects OpenShift via the `route.openshift.io` API
group and, on OpenShift, both configures the cluster and switches the
gateway transport automatically:
- Drives the gateway through a passthrough OpenShift Route secured with
mandatory mTLS instead of port-forward, so the connect suites
(live_policy_update, port_forward, sync, connect-based
sandbox_lifecycle, settings_management) actually pass. Computes the
Route host from the cluster ingress domain, extracts client mTLS
material from the openshell-client-tls secret, waits for the Route to
serve mTLS, asserts a certless caller is rejected at the TLS
handshake, and registers an mTLS CLI gateway pointing at the Route.
- Applies an SCC-compatible Helm values overlay that removes hardcoded
runAsUser/fsGroup, letting OpenShift assign UIDs from the namespace
range.
- Grants the privileged SCC to openshell-sandbox before Helm install
and removes it during cleanup.
- Grants the anyuid SCC to the PostgreSQL fixture service account in
DB scenarios and removes it during cleanup.
- All oc commands use --context to target the correct cluster.
The OpenShift e2e overlay (ci/values-openshift-e2e.yaml) turns TLS back
on, enables the Route, promotes the cert-verified caller to a dev
principal, and forces `image.pullPolicy`/`supervisor.image.pullPolicy`
to Always so runs against the `latest` upstream image use it instead of
a stale copy cached on the cluster nodes. Every OpenShift branch is
gated on OPENSHIFT_DETECTED, so the vanilla-Kubernetes port-forward path
is unchanged.
The Helm template for podSecurityContext is wrapped with {{- with }} so
null values omit the block instead of rendering invalid YAML.
The separate e2e:openshift task and e2e-openshift.sh script are removed
since e2e:kubernetes now covers OpenShift.
TESTING.md is updated with Kubernetes e2e documentation including
OpenShift auto-detection, dropping the e2e-host-gateway feature on
remote clusters, pinning IMAGE_TAG when the CLI and image versions
differ, task variants, and environment variables.
The debug-openshell-cluster skill gains an OpenShift platform row and
two SCC failure patterns (gateway rejected over hardcoded runAsUser,
sandbox missing the privileged SCC) covering the SCC handling and
podSecurityContext behavior this change introduces.
Signed-off-by: Jorge Garcia Oncins <jgarciao@redhat.com>
* fix(e2e): harden OpenShift SCC cleanup, mTLS gate, and Route timeout
Track the anyuid SCC grant for the PostgreSQL fixture with a dedicated
OPENSHIFT_POSTGRES_SCC_GRANTED flag set before the fixture apply, so a
failed apply no longer leaks the binding; cleanup now revokes it whenever
the grant succeeded, independent of deploy state.
Validate the Route server cert in the certless security gate (curl
--cacert instead of -k) and classify curl's exit code so only a TLS
client-auth rejection (35/56) counts as the expected certless rejection;
an unrelated DNS/timeout/TLS failure now fails loudly instead of masking
a potential mTLS hole.
Raise the OpenShift Route timeout in the e2e overlay. The default HAProxy
Route timeout is 30s, which severed long-lived transfers (large sandbox
upload/download, SSH-relay `sandbox connect`) mid-stream and failed the
sync e2e tests. Set both haproxy.router.openshift.io/timeout and
timeout-tunnel to 300s: a passthrough Route proxies in TCP mode, so
timeout-tunnel governs the established tunnel while timeout covers the
pre-tunnel phase.
Document the OpenShift transport exception, oc prerequisites and SCC
grants, and make the skopeo tag-check example copy-safe in TESTING.md.
Signed-off-by: Jorge Garcia Oncins <jgarciao@redhat.com>
---------
Signed-off-by: Jorge Garcia Oncins <jgarciao@redhat.com>
|
||
|
|
457f5dfae7 |
fix(helm): omit podSecurityContext block when value is null (#3034)
The gateway pod template rendered `securityContext:` unconditionally, so
setting `podSecurityContext: null` (e.g. to let OpenShift's SCC assign the
UID/GID range) produced `securityContext: null` instead of omitting the
block. Wrap the block in `{{- with .Values.podSecurityContext }}` so a null
value omits it and an explicit value renders unchanged.
Add a helm-unittest suite covering the default, explicit, and null cases.
Fixes #3033
Signed-off-by: Jorge Garcia Oncins <jgarciao@redhat.com>
|
||
|
|
7cc9551677 |
feat(server): support EC and EdDSA keys in OIDC JWKS validation (#2593)
Signed-off-by: Yuedong Wu <dwcn22@outlook.com> |
||
|
|
cc4ded2088 |
feat(helm): split gateway and workspace charts (#2643)
* feat(helm): split gateway and workspace charts Signed-off-by: Dhiraj Bokde <dbokde@nvidia.com> * fix(helm): preserve split chart upgrade compatibility Keep workspace manifests valid after value validation and default legacy reused values to the combined resource topology. * fix(ci): preserve VM runtime for E2E The Rust cache restores target/ after VM runtime artifacts are staged, overwriting target/vm-runtime-compressed before openshell-driver-vm is built. Stage the compressed runtime outside target and pass that location through OPENSHELL_VM_RUNTIME_COMPRESSED_DIR so build.rs can embed the supervisor. Also locate the Helm split-ownership test repository root from the script path rather than git rev-parse. The test runs in a container where the GitHub checkout can be owned by a different UID and rejected as dubious ownership. Signed-off-by: Dhiraj Bokde <dbokde@nvidia.com> * fix(ci): install yq for Helm ownership test The split-chart ownership regression uses yq to inspect rendered YAML, but the Helm CI container installs only tools declared in mise. Declare and lock yq so mise install --locked provides the test dependency. Signed-off-by: Dhiraj Bokde <dbokde@nvidia.com> --------- Signed-off-by: Dhiraj Bokde <dbokde@nvidia.com> |
||
|
|
e508c169ea |
fix(helm): honor empty clientCaSecretName for HTTPS-only mode (#2235)
* fix(helm): honor empty clientCaSecretName for HTTPS-only mode Signed-off-by: Yuedong Wu <dwcn22@outlook.com> * docs(skills): document HTTPS-only clientCaSecretName in debug-openshell-cluster Signed-off-by: Yuedong Wu <dwcn22@outlook.com> --------- Signed-off-by: Yuedong Wu <dwcn22@outlook.com> |
||
|
|
9b6d904e88 |
feat(compute): delegate sandbox authentication to drivers (#2968)
* feat(compute): delegate sandbox authentication to drivers Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> * fix(kubernetes): align sandbox identity annotation Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> * test(auth): restore sandbox bootstrap coverage Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> * fix(kubernetes): satisfy ownership test lint Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> --------- Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> |
||
|
|
f68867b869 |
feat(gateway): identify gateways in exported traces (#2647)
* feat(gateway): add installation name configuration Add a first-class operator-assigned gateway name with TOML, CLI, environment, and Helm configuration surfaces. Local gateways default to openshell, while Helm defaults to the chart fullname; operators sharing a collector across namespaces or clusters can set a globally distinct name. Signed-off-by: Kris Hicks <khicks@nvidia.com> * feat(gateway): identify gateways in exported traces Attach the configured gateway installation name and compute driver to the gateway OpenTelemetry resource so operators can filter traces from multiple installations that share a collector. Forward the gateway name and OTLP endpoint to managed external drivers so their distinct service resources carry the same installation identity. Keep service.name stable per process type, omit blank resource values, and leave per-span operation names and request attributes unchanged. Refs #2507 Signed-off-by: Kris Hicks <khicks@nvidia.com> --------- Signed-off-by: Kris Hicks <khicks@nvidia.com> |
||
|
|
d0dfb22baf |
feat(kubernetes): export driver traces over OTLP (#2958)
Mirror the VM, Podman, and Docker driver tracing setup for Kubernetes. Export standalone driver spans through OTLP/gRPC as the distinct openshell-driver-kubernetes service, preserve gateway trace context, record lifecycle operations and gRPC failures, and flush spans on shutdown. Kubernetes currently runs in-process when selected as a built-in gateway driver. Use the temporary server-boundary shim shared with Podman and Docker so traces retain the shape they will have when Kubernetes moves to a separate process. Move the common ComputeDriver RPC tracing layer into openshell-otel to keep all drivers aligned. Propagate the active W3C context through the controller-reserved Sandbox annotation and enable Agent Sandbox OTLP export in the local k3s workflow. This connects asynchronous controller reconciliation spans to the originating OpenShell create trace. Expose gateway OTLP configuration through Helm and add an Aspire collector to the local k3s workflow. Extend helm:k3s:forward with OTLP ingest and trace UI forwarding for Kubernetes and local container gateway development. Signed-off-by: Kris Hicks <khicks@nvidia.com> |
||
|
|
40d1b48666 |
feat(provider): support for SPIFFE backed token exchange (#1970)
* feat(provider): add ability to request token exchange instead of client credentials as OAuth grant_type Signed-off-by: Gordon Sim <gsim@redhat.com> * test(proxy): add further tests for token exchange Signed-off-by: Gordon Sim <gsim@redhat.com> * test(provider): add runnable example for token exchange Signed-off-by: Gordon Sim <gsim@redhat.com> * test(e2e): cover Podman token exchange grants Signed-off-by: Gordon Sim <gsim@redhat.com> * refactor(oauth): extract duplicated functionality from server and supervisor Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(provider): evict nearest-to-expiry entry from intermediate token cache Signed-off-by: Gordon Sim <gsim@redhat.com> * doc(supervisor): add podman example for token exchange Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(provider): withhold token-exchange subject credentials Signed-off-by: Gordon Sim <gsim@redhat.com> --------- Signed-off-by: Gordon Sim <gsim@redhat.com> |
||
|
|
3be2cd8a29 |
fix(helm): preflight Agent Sandbox APIs (#2867)
* fix(helm): preflight Agent Sandbox APIs Signed-off-by: Evan Lezar <elezar@nvidia.com> * test(kubernetes): share Agent Sandbox setup Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(e2e): wait for Agent Sandbox CRD status Signed-off-by: Evan Lezar <elezar@nvidia.com> * ci(canary): sparse-checkout sandbox helper Signed-off-by: Evan Lezar <elezar@nvidia.com> --------- Signed-off-by: Evan Lezar <elezar@nvidia.com> |
||
|
|
59479f492a |
feat(k8s): add namespace-per-workspace support (RFC 0011 Phase 3) (#2656)
* feat(k8s): add namespace-per-workspace support (RFC 0011 Phase 3) Implement three workspace namespace modes for the Kubernetes compute driver: shared (default, preserves current single-namespace behavior), managed (auto-creates/deletes namespaces per workspace), and operator (pre-provisioned namespaces with dynamic discovery via label selector or drop-in allowlist file). Key changes: - WorkspaceMode enum and namespace resolution in driver config - Managed namespace lifecycle with ServiceAccount and OpenShift SCC annotation propagation - Cluster-wide sandbox CR watchers for managed/operator modes - NamespaceValidator (Exact/Prefix/Allowlist) for SA token auth - Workspace-aware credential secret storage - Helm ClusterRole for multi-namespace RBAC - Gateway config, architecture, and reference docs Signed-off-by: Derek Carr <decarr@redhat.com> * test(k8s): add e2e tests for workspace namespace modes Add end-to-end tests for managed and operator workspace modes introduced in RFC 0011 Phase 3. The managed mode tests verify namespace creation with correct labels, ServiceAccount provisioning, sandbox CR placement, and namespace survival with remaining sandboxes. The operator mode tests verify rejection of unlabeled and nonexistent namespaces. The positive operator path (sandbox in labeled namespace) is known to fail due to an RBAC gap and will be addressed separately. Also fixes Helm 4 compatibility: move SPDX license headers inside conditional guards in 8 chart templates to prevent empty comment-only documents, and fix a trailing whitespace trimmer in clusterrole.yaml that concatenated the license header with apiVersion. Adds cleanup sweep in with-kube-gateway.sh to remove managed and operator namespaces before Helm uninstall, and mise tasks for running each mode independently. Signed-off-by: Derek Carr <decarr@redhat.com> * feat(k8s): add operator namespace label watcher Spawn a background kube::runtime::watcher in the K8s driver that watches namespaces matching the configured label selector and populates the OperatorNamespaceAllowlist at runtime. The driver owns the allowlist and exposes its Arc so the server can share the same set with the SA token authenticator. create_sandbox now gates pod creation on the allowlist in operator mode — workspaces whose namespace is not yet labeled are rejected at resource render time rather than silently proceeding. Workspace lifecycle itself is unaffected; only sandbox (resource) creation is gated. Signed-off-by: Derek Carr <decarr@redhat.com> * fix(k8s): harden operator mode and address review findings Close the fail-open gap in operator mode when only operator_namespace_file is configured: the allowlist is now created unconditionally in operator mode (fail-closed from startup). Implement the namespace file watcher using the notify crate, following the TLS hot-reload pattern (parent-directory watch, 1s debounce, ConfigMap symlink-swap safe). The file format is a JSON array of namespace name strings. Additional fixes from the 10-reviewer audit: - Change allowlist rejection from InvalidArgument to FailedPrecondition so callers know the request may succeed later once the namespace is provisioned. - NamespaceValidator::Allowlist now holds the OperatorNamespaceAllowlist newtype instead of a raw Arc<RwLock<BTreeSet>>, eliminating silent denial on RwLock poison. - Verify LABEL_MANAGED_BY and LABEL_GATEWAY_ID ownership before deleting a managed namespace. - Replace fixed 5s sleep in operator e2e test with a 30s poll loop. - Add Helm validation for workspaceMode values. - Fix Helm README type column and description for operator fields. - Add insert/remove methods to OperatorNamespaceAllowlist; label watcher now uses them instead of reaching through shared(). - Reject configs with both operator_namespace_label and operator_namespace_file set. Signed-off-by: Derek Carr <decarr@redhat.com> * feat(k8s): add workspace-level compute driver RPCs and harden RBAC Decouple namespace lifecycle from sandbox lifecycle by adding EnsureWorkspace/DeleteWorkspace RPCs to the ComputeDriver service. Namespace creation now happens before credential storage and namespace deletion happens on workspace delete, fixing credential storage in managed workspace mode. - Add EnsureWorkspace and DeleteWorkspace proto RPCs with implementations across all compute drivers (K8s managed delegates to ensure_namespace/delete_namespace_if_empty; others no-op) - Wire ensure_workspace into provider create/update/refresh paths so the namespace exists before the credential driver writes secrets - Wire delete_workspace into workspace deletion for cleanup - Remove delete_namespace_if_empty from sandbox deletion path - Scope ClusterRole secrets access to non-shared workspace modes - Add TODO for TLS cert hot-reload in sandbox gRPC client - Harden e2e tests with control-plane sandbox resolution assertions - Fix docker image save --platform flag for OCI index manifests Signed-off-by: Derek Carr <decarr@redhat.com> * fix(k8s): address re-review findings and add test coverage - Use server-side apply for TLS secret sync (fixes second sandbox creation failure when TLS is enabled) - Scope gateway-ID label selector unconditionally across all workspace modes (fixes operator reads/watches/deletes seeing foreign sandboxes) - Validate operator allowlist in EnsureWorkspace and DeleteWorkspace RPCs (prevents credential writes to namespaces outside the allowlist) - Extend ClusterRole secrets patch+delete to all non-shared modes with credential driver enabled (fixes operator credential storage RBAC) - Validate namespace ownership on 409 conflict in ensure_namespace (prevents adopting unowned namespaces in managed mode) - Replace delete_namespace_if_empty with unconditional delete_namespace letting Kubernetes cascade cleanup (fixes stuck terminating CRs) - Strengthen NetworkPolicy TODO to cover both managed and operator modes - Extract selector and ownership logic into testable free functions - Add unit tests for gateway-ID selectors and namespace ownership - Add Helm ClusterRole RBAC tests for operator credential driver Signed-off-by: Derek Carr <decarr@redhat.com> * ci(k8s): add workspace managed and operator mode e2e to CI Wire the existing e2e:kubernetes:workspace-managed and e2e:kubernetes:workspace-operator mise tasks into the branch-e2e workflow so they run alongside the other core Kubernetes e2e suites. Both are gated by run_core_e2e and included in the Core E2E result gate. Signed-off-by: Derek Carr <decarr@redhat.com> * test(k8s): add e2e tests for workspace namespace modes Add 7 new e2e tests covering workspace namespace lifecycle, TLS secret copying, ownership conflict detection, DNS-1123 validation, operator namespace preservation, and dynamic label watcher behavior. Fix async sandbox deletion race condition in existing tests by polling sandbox list instead of asserting immediately after delete. Signed-off-by: Derek Carr <decarr@redhat.com> * fix(k8s): grant secrets/patch unconditionally and backfill gateway-id labels Address two review findings: 1. RBAC: server-side apply (PATCH) is used for TLS secret sync in multi-namespace modes, but the ClusterRole only granted patch when the kubernetes-secrets credential driver was enabled. Grant patch unconditionally for non-shared modes since TLS sync always needs it; keep delete gated on the credential driver. 2. Upgrade safety: the new gateway-id label selector would orphan legacy Sandbox CRs that predate its introduction. Add a startup backfill in shared mode that patches any managed Sandbox CR missing the gateway-id label before the driver begins serving requests. Signed-off-by: Derek Carr <decarr@redhat.com> * fix(k8s): address workspace namespace review findings Signed-off-by: Derek Carr <decarr@redhat.com> * fix(k8s): address follow-up review findings Signed-off-by: Derek Carr <decarr@redhat.com> * fix(k8s): preserve workspace lookup after rebase Signed-off-by: Derek Carr <decarr@redhat.com> * fix(helm): allow managed secret creation Signed-off-by: Derek Carr <decarr@redhat.com> * fix(k8s): stop pods in workspace namespace Signed-off-by: Derek Carr <decarr@redhat.com> * test(k8s): scope pod deletion check to v1alpha1 Signed-off-by: Derek Carr <decarr@redhat.com> * fix(k8s): address workspace namespace review findings Signed-off-by: Derek Carr <decarr@redhat.com> --------- Signed-off-by: Derek Carr <decarr@redhat.com> |
||
|
|
d0c6dc3fd8 |
feat(kubernetes): support corporate upstream proxy (#2633)
* feat(kubernetes): support corporate upstream proxy Signed-off-by: loveRhythm1990 <qiuweimin@126.com> * fix(kubernetes): reject proxy_auth_secret_key values Kubernetes cannot create Gateway validation accepted proxy_auth_secret_key values that Kubernetes rejects when creating the Secret (keys longer than 253 bytes, or the reserved "."/".." names), turning an invalid deployment setting into repeated sandbox Pod-provisioning failures instead of a startup error. Reject them in validate_upstream_proxy_config so they fail closed at gateway startup. Signed-off-by: loveRhythm1990 <qiuweimin@126.com> * docs(skill): add corporate upstream proxy checks to debug-openshell-cluster Add a Kubernetes corporate upstream proxy troubleshooting section covering rendered [openshell.drivers.kubernetes] configuration, credential Secret volume events, supervisor arguments and mounts confined to the network- supervising container, and proxy reachability. Signed-off-by: loveRhythm1990 <qiuweimin@126.com> --------- Signed-off-by: loveRhythm1990 <qiuweimin@126.com> |
||
|
|
c4b500a7de |
feat(helm): cert-manager external issuer + OpenShift passthrough Route (#2468)
* fix(core): trust public root CAs alongside the sandbox mTLS CA The supervisor gRPC client only trusted the CA configured via OPENSHELL_TLS_CA, since tonic ClientTlsConfig starts with an empty root store unless with_native_roots()/with_webpki_roots() is also enabled. Deployments where the gateway server certificate is issued by a public CA (e.g. cert-manager against an ACME issuer) caused every supervisor connection to fail the TLS handshake with "UnknownCA", since the sandbox mTLS CA and the server cert issuer were no longer the same. Enable both native and webpki roots in addition to the configured CA. tonic root store is a union of all configured sources, so this does not weaken verification for existing self-signed deployments. webpki-roots (compiled in) is enabled alongside native-roots since the supervisor binary may run in minimal sandbox images without a populated system CA bundle. Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * feat(helm): support external cert-manager issuers and OpenShift Route passthrough Add certManager.serverIssuerRef/clientIssuerRef so the gateway and mTLS client certificates can be issued by a real Issuer/ClusterIssuer (e.g. ACME) instead of only the chart built-in self-signed CA. Add openshiftRoute template for exposing the gateway via a TLS passthrough Route so the gateway keeps terminating its own TLS/mTLS. The server Certificate excludes internal-only SANs (cluster-local, localhost, loopback) when an external issuer is configured, since ACME issuers reject those per CA/Browser Forum baseline requirements. A template-time fail guard catches the misconfiguration at helm install time rather than asynchronously at cert-manager issuance time. Includes Helm unittest coverage for both issuerRef overrides and Route rendering, plus a CI values overlay for lint coverage. Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * docs: document cert-manager external issuer and OpenShift Route Update managing-certificates.mdx with the serverIssuerRef workflow and install-time validation behavior. Add a production section to the OpenShift guide covering passthrough Route with a real certificate. Regenerate Helm README for new certManager and openshiftRoute values. Sync debug-openshell-cluster skill with new troubleshooting steps for ACME issuance failures and supervisor UnknownCA from mismatched CAs. Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(helm,core): address PR review feedback on cert-manager external issuer Addresses all five blocking review items from #2468: 1. Remove .with_native_roots() from supervisor gRPC client -- the supervisor runs inside the user-selected sandbox image, so the image CA bundle is not operator-controlled. Keep .with_webpki_roots() (compiled-in, not user-controlled) alongside the configured CA. 2. Fail at render time when serverIssuerRef.name is set but clientCaFromServerTlsSecret is still true. Add negative Helm test. 3. Remove clientIssuerRef -- changing only clientIssuerRef breaks both directions because trust bundles are not modeled separately. Change serverIssuerRef.kind default from ClusterIssuer to Issuer. 4. Add server.oidc.issuer and server.oidc.audience to the documented OpenShift production Helm command. Add Access Control prerequisite. 5. Fail at render time when openshiftRoute.enabled and disableTls are both true. Add negative Helm test. Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(drivers): strip GATEWAY_TLS_SERVER_NAME from Docker and Podman env Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(helm,drivers): guard default clientCaSecretName and add env-strip tests Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * feat(tls): SNI-based dual certificate for internal and external server TLS Split the gateway server certificate into two: an internal cert issued by the chart's own CA (for supervisor connections via cluster-local SANs) and an external cert issued by an operator-configured Issuer such as ACME/Let's Encrypt (for CLI and Route access via public SANs). The gateway uses SNI-based certificate selection: connections whose SNI hostname matches external_server_names receive the external cert; all others (including those with no SNI) receive the internal cert. Security improvement: remove .with_webpki_roots() from the supervisor gRPC client so supervisors trust only the chart CA, closing a MITM vector via publicly-trusted certificates in user-supplied container images. Key changes: - Add DualCertResolver with SNI-based cert selection and full test coverage - Add external_cert_path, external_key_path, external_server_names to TlsConfig - Validate partial external cert config (error on cert-without-key or vice versa) - Validate empty external_server_names when external cert is configured - Split cert-manager templates into internal + external Certificate resources - Add Helm guards for misconfigured external issuer (empty serverDnsNames, internal-only SANs with external issuer, conflicting clientCaFromServerTlsSecret) - Update gateway-config.mdx, managing-certificates.mdx, openshift.mdx docs - Update debug-openshell-cluster skill for dual-cert troubleshooting Signed-off-by: Pi Agent <agent@openshell.local> * fix(drivers): strip GATEWAY_TLS_SERVER_NAME in VM driver and correct comments Add the same GATEWAY_TLS_SERVER_NAME environment stripping to the VM compute driver that Docker, Podman, and Kubernetes drivers already perform. Without this, a sandbox user on the VM driver could override the TLS server name the supervisor verifies. Fix stale comments in Docker and Podman drivers that referenced 'with WebPKI roots trusted' — WebPKI roots are explicitly not trusted after the tls-webpki-roots removal. Use tls-ring instead of bare channel for tonic in openshell-core so the TLS API (ClientTlsConfig, Endpoint::tls_config) is available without pulling in any root certificate store. Signed-off-by: Pi Agent <agent@openshell.local> * fix(tls,helm): wildcard SNI matching and Route host validation Add RFC 6125 single-level wildcard matching to DualCertResolver so external_server_names entries like *.example.com correctly match SNI hostnames like gw.example.com. Previously only exact matches worked, silently falling back to the internal cert for wildcard configurations. Add a Helm fail guard in route.yaml that rejects openshiftRoute.host values not listed in certManager.serverDnsNames when an external issuer is configured — catches cert/route hostname mismatches at install time instead of at TLS connect time. Quote the host field in route.yaml for robustness. Signed-off-by: Pi Agent <agent@openshell.local> * fix(helm): address blocking review items — client-CA guard, wildcard Route, serverIssuerRef gate 1. Remove the obsolete guard rejecting serverIssuerRef + clientCaFromServerTlsSecret=true. The internal server certificate is always signed by the chart CA (the same CA that signs the client cert), so clientCaFromServerTlsSecret=true is correct — its filtered ca.crt is exactly the right trust anchor. The old workaround (mounting openshell-ca-tls directly) unnecessarily exposed the CA private key to the gateway container. Remove the client-CA overrides from docs, CI overlay, and production examples. 2. Route host validation now supports wildcard certificates per RFC 6125: single-level wildcards like *.example.com match gateway.example.com but not deep.sub.example.com. Require an explicit openshiftRoute.host when an external issuer is configured — without one, OpenShift generates a hostname absent from serverDnsNames. 3. Reject serverIssuerRef.name when certManager.enabled is false — the external certificate, its Secret mount, and the gateway TLS config all require cert-manager to be enabled. Validated on ROSA (dev.dyee.p3) with branch-built images: - Fresh install with letsencrypt-prod ClusterIssuer - SNI dual-cert: external hostname served Let's Encrypt cert - Supervisor mTLS via internal cert path: ConnectSupervisor accepted - Client CA volume: filtered ca.crt from internal server secret (no key) - CLI connected via Route + OIDC Helm tests: 81 pass across 7 suites. Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> --------- Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> Signed-off-by: Pi Agent <agent@openshell.local> |
||
|
|
170961997f |
chore(ci): disable telemetry in internal test runs (#2648)
* chore(ci): disable telemetry in internal test runs Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> * test(ci): remove brittle telemetry wiring test Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> * docs: trim CI telemetry guidance Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> * test(e2e): share telemetry default with OpenShift Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> --------- Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> |
||
|
|
5548405fcb |
feat(credentials): add provider credential storage drivers (#2437)
* feat(credentials): add provider credential storage drivers Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(credentials): harden credential update handling Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(credentials): harden credential driver security, correctness, and performance Address review findings from the credential storage drivers PR: - Route additional_credentials through the driver on refresh to prevent silent data loss for multi-credential providers (e.g. AWS STS) - Clean up stored credential handles on CAS failure during refresh to prevent orphaned secrets in external backends - Enforce namespace validation in the Kubernetes Secrets driver to prevent cross-namespace credential access when allow_reference_namespace is not enabled - Cache Vault Kubernetes auth tokens with 80% TTL to avoid re-authenticating on every credential operation - Parallelize resolve_credentials in all three drivers using try_join_all for faster sandbox startup - Add existingSecret support for the KEK Secret to fix helm template/GitOps workflows where lookup returns empty and regenerates the key - Document RBAC blast radius for the Kubernetes Secrets credential driver and recommend a dedicated namespace Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com> Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> * fix(credentials): add optimistic concurrency, fix thundering herd, parallelize operations Use resourceVersion optimistic concurrency with retry loop for K8s Secret ownership checks to prevent TOCTOU races. Switch Vault token cache from RwLock to Mutex with double-check pattern to prevent thundering herd on cache miss. Parallelize credential store and delete operations across independent keys using try_join_all. Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com> Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> * fix(credentials): handle partial failures, add delete retry, consolidate cleanup Replace try_join_all with join_all in credential store/delete operations to handle partial failures — successfully-stored handles are cleaned up when another key fails. Add retry loop with conflict detection to db-credstore delete_credential, matching the K8s driver pattern. Consolidate 4 manual cleanup_pre_stored_provider_credentials call sites into a single error handler using an async block. Remove inconsistent .trim() from db-credstore validate_handle_owner. Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com> Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> * fix(credentials): fix retry loop guard and remove unprotected validation Remove attempt-count guard from 409/Aborted match arms in retry loops so the post-loop Status::aborted error is reachable after exhausting retries. Previously, last-attempt conflicts fell through to the catch-all error arm, producing misleading Status::unavailable errors. Remove duplicate validation calls that ran after prepare_provider_credential_update but outside the cleanup-protected async block, which would leak pre-stored handles on failure. Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com> Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> * fix(credentials): add workspace/provider UUID to credential backend paths Include workspace and provider ID in credential backend object paths to ensure cross-workspace uniqueness and prevent credential collision (GATOR-1806c9be-01). - Updated credential driver proto to include workspace and provider_id fields - Modified Vault driver to include workspace/provider_id in managed_secret_path - Modified Kubernetes Secrets driver to include workspace/provider_id in credential_owner_id and managed_secret_name - Updated all credential runtime calls to pass workspace/provider_id - Updated tests to use the new signatures This prevents two workspaces sharing the same external credential store from colliding on provider names, which was a critical security issue (CWE-639). * fix(credentials): preserve provider-level expiration for handle-backed credentials Compute effective expiration from both provider and driver values using the earliest non-zero timestamp and skip expired values before insertion (GATOR-1806c9be-02). - Modified resolve_provider_handles to check provider credential_expires_at_ms - Skip expired credentials during resolution instead of returning them - Use effective expiration (min of provider and driver) in resolution results - Fix inference.rs to preserve earliest expiration when merging This ensures handle-backed credentials respect the same expiration semantics as inline credentials. * fix(credentials): stage refresh changes under new handles before validation Stage credential replacements under new immutable handles instead of reusing existing handles to prevent overwriting committed values before validation/CAS (GATOR-1806c9be-03). - Stage credentials with empty existing_handles map to force new handle creation - Validate and CAS before the new values are committed to backend storage - Delete old handles only after successful CAS - On CAS failure, delete only the newly staged handles - This prevents CWE-362/CWE-367 race conditions where failed refreshes could still modify or delete the active credential The fix ensures that a rejected refresh cannot modify the backend object still referenced by the committed provider record. * fix(credentials): add timeouts to credential driver RPCs Apply configured timeouts to both startup capability negotiation and runtime RPCs to prevent indefinite hangs (GATOR-1806c9be-05). - Add DEFAULT_CREDENTIAL_DRIVER_RPC_TIMEOUT_SECS constant (30s) - Apply timeout to GetCapabilities during startup connection - Apply timeout to all runtime RPCs (store, delete, resolve) - Use tokio::time::timeout to bound the entire GetCapabilities operation during startup, not just the socket connection - Return contextual deadline errors on timeout This prevents a faulty or overloaded driver from hanging gateway operations indefinitely. * fix(credentials): fix test to use consistent workspace/provider identity The Kubernetes auth Vault resolve test was constructing a managed path with test-workspace/test-provider-id but sending default/prov-123 in the request, causing validation to reject the request (GATOR-18e32351-01). - Update test to use test-workspace and test-provider-id in the request to match the logical_path construction - This ensures the test exercises the intended code path and validates Kubernetes auth resolution properly The test now passes and correctly validates identity enforcement. * fix(credentials): use unique staging ID for refresh to avoid overwrites Stage refresh replacements under genuinely distinct immutable handles using a unique staging ID to prevent overwriting committed values (GATOR-1806c9be-03). - Generate a unique staging ID using UUID for each refresh operation - Use this staging ID when storing credentials instead of the real provider ID - Pass the same staging ID during cleanup on failure to delete only staged objects - This ensures deterministic paths (Vault) and object names (K8s) don't collide with the committed provider's credentials The fix prevents failed refreshes from silently replacing active credentials or breaking providers by deleting still-referenced backend objects. * fix(credentials): wrap credential driver RPCs in local timeouts Add local tokio::time::timeout wrappers around credential driver RPCs to bound non-compliant or stalled UDS peers (GATOR-1806c9be-05). - Wrap StoreCredential, DeleteCredential, and ResolveCredentials in local timeouts - Return contextual deadline_exceeded errors when timeouts occur - Keep existing gRPC timeout metadata for compliant implementations - GetCapabilities during startup was already wrapped in previous commit This ensures a faulty local driver cannot hang gateway operations indefinitely, even if it accepts the connection but never responds to the RPC. * fix(credentials): preserve ownership for staged refreshes Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(credentials): bound startup capability probe Signed-off-by: Seth Jennings <sjenning@redhat.com> * test(provider): authenticate credential handler requests Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(ci): grant actions read to credential driver e2e Signed-off-by: Seth Jennings <sjenning@redhat.com> --------- Signed-off-by: Taylor Mutch <taylormutch@gmail.com> Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com> Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> Signed-off-by: Seth Jennings <sjenning@redhat.com> Co-authored-by: Taylor Mutch <taylormutch@gmail.com> Co-authored-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> |
||
|
|
905b554c7c |
refactor(network): consolidate proxy egress pipeline (#2373)
* refactor(network): introduce shared egress pipeline Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(network): cover shared proxy egress paths Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * refactor(network): make destination authorization explicit Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * refactor(network): pin proxy relay policy context Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(network): lock relay generation contracts Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(network): establish phase zero compatibility baseline Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * feat(policy): detect ambiguous network endpoints Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * feat(sandbox): fail closed on invalid policy updates Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * refactor(network): invalidate relays on policy changes Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * docs(policy): document validation failure posture Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(network): cover validation and middleware egress Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(config): move policy failure mode to gateway toml Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(network): name proxy contracts by behavior Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(network): align overlap validation with endpoint selection Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(network): expect hard loopback denial Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(network): match declared endpoint denial Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(policy): preserve path-specific endpoint overrides Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(network): respect hard-blocked host gateways Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(network): reconcile proxy refactor with main Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(network): preserve CONNECT policy generation Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(policy): cover runtime endpoint glob semantics Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(server): reject ambiguous policies before persistence Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * docs(policy): explain ambiguity preflight behavior Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(policy): compare body limits within protocol Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(proxy): avoid global tracing capture race Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * chore(server): format rebased provider tests Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(server): authenticate rebased policy requests Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(sandbox): retain runtime on middleware outage Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(server): preflight provider composition activation Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(sandbox): distinguish runtime failure transitions Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> --------- Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> Signed-off-by: John Myers <johntmyers@users.noreply.github.com> Co-authored-by: John Myers <johntmyers@users.noreply.github.com> |
||
|
|
7955c8309b |
feat(k8s): support configuring workspace PVC storageClassName (#2463)
The Kubernetes driver's default workspace PVC never set storageClassName, so on clusters with no default StorageClass the PVC stayed Pending and sandbox creation failed. Add a workspace_storage_class option to KubernetesComputeConfig, wired through SandboxPodParams into the generated volumeClaimTemplates. When non-empty it sets storageClassName; empty preserves the current behavior of relying on the cluster default StorageClass. Expose it via the OPENSHELL_K8S_WORKSPACE_STORAGE_CLASS env var on both the standalone driver and the embedded gateway runtime defaults, and via the server.workspaceStorageClass Helm value. Closes #2442 Signed-off-by: lr90 <qiuweimin@matrixorigin.cn> |
||
|
|
8eacb4779f |
feat(kubernetes): add sidecar supervisor topology (#2076)
* feat(kubernetes): add sidecar supervisor topology Add the Kubernetes sidecar supervisor topology, its Helm/Skaffold configuration, topology documentation, and sidecar e2e matrix coverage. Skip root-only sandbox identity rewriting when process enforcement is network-only so the low-permission sidecar process container can start successfully. Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(supervisor): avoid similar process id names Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(supervisor): avoid similar process id names Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(sandbox): avoid similar proxy id names Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * docs(kubernetes): clarify sidecar topology limits Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(kubernetes): keep sidecar process leaf capless Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(kubernetes): refresh sidecar provider env snapshots Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * test(supervisor): align hot-swap identity regression Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(kubernetes): stage sidecar mtls files before proxy chown Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(kubernetes): simplify sidecar supervisor topology Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * chore(helm): reuse sidecar skaffold values Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(supervisor): avoid similar iptables helper names Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(e2e): harden kube gateway wrapper setup Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(supervisor): avoid nft batch rollback on OCP Run nftables setup as individual commands so optional conntrack and log expressions can fail without rolling back required table, chain, and reject rules. Signed-off-by: Seth Jennings <sjenning@redhat.com> Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(kubernetes): preserve process identity in sidecar topology Render sidecar pods with a shared process namespace, keep binary-aware network policy enabled, and move Kubernetes sidecar settings under the nested sidecar config table. Also apply unprivileged Landlock/seccomp setup in NetworkOnly supervisor mode so sidecar topology keeps sandbox child hardening without privileged process setup. Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * refactor(kubernetes): replace sidecar snapshots with control socket Coordinate sidecar policy and provider bootstrap over a local Unix socket so the process leaf no longer reads policy/provider snapshot files. Report entrypoint startup through the control channel and keep gateway credentials confined to the network sidecar. Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * feat(kubernetes): support relaxed sidecar network identity Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(sandbox): satisfy sidecar clippy lint Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * refactor(kubernetes): standardize topology naming Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(sandbox): satisfy linux clippy timeout import Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(kubernetes): support kata sidecar on ipv4 pods Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(kubernetes): satisfy linux clippy for sidecar fallback Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * chore(kubernetes): remove stale supervisor topology references Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(kubernetes): enable sidecar binary policy inspection Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(kubernetes): harden sidecar control boundary Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(kubernetes): couple sidecar supervisor lifecycles Signed-off-by: Taylor Mutch <taylormutch@gmail.com> --------- Signed-off-by: Taylor Mutch <taylormutch@gmail.com> Signed-off-by: Seth Jennings <sjenning@redhat.com> Co-authored-by: Seth Jennings <sjenning@redhat.com> |
||
|
|
bebf440b25 |
fix(helm): propagate supervisor image overrides (#2216)
Signed-off-by: Taylor Mutch <taylormutch@gmail.com> |
||
|
|
8c0ecac8cf |
docs(openshift): simplify install steps and add Helm README entries for OpenShift overrides (#2125)
Signed-off-by: ChristianZaccaria <christian.zaccaria.cz@gmail.com> |
||
|
|
6461677c32 |
feat(policy): accept numeric UIDs for sandbox process identity (#1973)
* feat(policy): accept numeric UIDs in sandbox process identity validation Allow run_as_user and run_as_group to be either the literal 'sandbox' or a numeric UID/GID within [1000, 2_000_000_000]. This removes the hard dependency on a baked-in 'sandbox' user in container images, enabling compute drivers to inject resolved UIDs at sandbox creation. Phase 1 of #1959. Signed-off-by: Seth Jennings <sjenning@redhat.com> * feat(supervisor): accept numeric UIDs for process identity dropping Allow run_as_user and run_as_group to be numeric UIDs/GIDs, removing the hard dependency on a baked-in 'sandbox' user in container images. Changes: - validate_sandbox_user(): accepts numeric UIDs without passwd lookup (logs OCSF event); keeps passwd check for "sandbox" name; rejects non-numeric non-sandbox strings that fail passwd lookup - prepare_filesystem(): passes numeric UIDs/GIDs directly to chown() instead of requiring a passwd entry - drop_privileges(): resolves numeric UIDs/GIDs directly via UID::from_raw / Gid::from_raw; skips initgroups when target uid matches current euid; uses guard conditions before setgid/setuid calls - session_user_and_home(): falls back to ("{uid}", "/sandbox") for numeric UIDs, avoiding a passwd lookup that will fail Re-exports MIN_SANDBOX_UID and MAX_SANDBOX_UID from openshell-policy so callers have consistent range constants. Phase 2 of #1959. Signed-off-by: Seth Jennings <sjenning@redhat.com> * feat(driver-kubernetes): resolve sandbox UID/GID from config or OpenShift SCC annotations Phase 3 of the numeric-UID plan: allow operators to specify explicit sandbox_uid/sandbox_gid in Kubernetes driver config, auto-detect from OpenShift SCC namespace annotations, and propagate resolved values to supervisor container env vars and PVC init container securityContext. Changes: - Add sandbox_uid/sandbox_gid fields to KubernetesComputeConfig - Add SANDBOX_UID/SANDBOX_GID env var constants to openshell-core - Implement resolve_sandbox_identity() to fetch namespace annotations and auto-detect OpenShift SCC UID ranges (sa.scc.uid-range) - Pass resolved UID/GID through SandboxPodParams to pod spec builder - Inject SANDBOX_UID/SANDBOX_GID env vars into supervisor container - Update PVC init container securityContext with resolved UID/GID instead of hard-coded root - Add comprehensive unit tests for resolution logic and annotation parsing (resolve_sandbox_uid, resolve_sandbox_gid, OpenShift SCC annotation parsing) Signed-off-by: Seth Jennings <sjenning@redhat.com> * feat(driver-vm): add configurable sandbox UID/GID and update docs/examples Phase 4 of the numeric-UID plan: replace hardcoded SANDBOX_UID (10001) in VM rootfs preparation with configurable sandbox_uid/sandbox_gid fields. Changes: - Add sandbox_uid/sandbox_gid to VmDriverConfig with serde derives - Pass resolved UID/GID through prepare_sandbox_rootfs_from_image_root to ensure_sandbox_guest_user which writes /etc/passwd/group/gshadow - Update BYOC Dockerfile: remove groupadd/useradd, document runtime UID injection and the ability to skip baked-in sandbox user - Update gateway-config.mdx: document sandbox_uid/sandbox_gid for both Kubernetes (with OpenShift SCC autodetection) and VM drivers - Update sandbox-compute-drivers.mdx: add Sandbox User Identity section explaining numeric UID support across all compute drivers - Update rootfs tests to use non-default UIDs, verify config passthrough Signed-off-by: Seth Jennings <sjenning@redhat.com> * code review changes * fix(supervisor): harden tests for restricted CI container environments Guard tests against CI-specific constraints: root without CAP_SETPCAP, UIDs with no /etc/passwd entry, and restricted /proc access. Signed-off-by: Seth Jennings <sjennings@nvidia.com> Signed-off-by: Seth Jennings <sjenning@redhat.com> --------- Signed-off-by: Seth Jennings <sjenning@redhat.com> Signed-off-by: Seth Jennings <sjennings@nvidia.com> |
||
|
|
abcd15d1f2 |
feat(helm): add TLS termination for Envoy Gateway ingress (#2015)
The chart's optional Gateway API ingress only rendered a plaintext HTTP listener, so the gateway could not be exposed over TLS. Add an HTTPS listener option that terminates TLS at the Envoy Gateway and forwards plaintext gRPC to the gateway pod. - gateway.yaml renders an HTTPS listener with `tls.mode: Terminate` and `certificateRefs` when `grpcRoute.gateway.listener.protocol=HTTPS`, keeping the default HTTP listener unchanged. Guards fail the render when `certificateRefs` is empty or `server.disableTls` is not true (the chart does not render a BackendTLSPolicy for re-encryption). - values.yaml adds `grpcRoute.gateway.listener.tls.certificateRefs`. - ci/values-gateway-tls.yaml exercises the HTTPS branch in lint/render. - docs/kubernetes/ingress.mdx documents HTTPS setup and clarifies that Envoy Gateway only terminates TLS (no OIDC SecurityPolicy); client identity uses OIDC bearer tokens, with the client-credentials grant for headless agents. - debug-openshell-cluster skill gains HTTPS-ingress troubleshooting rows. - Regenerated the chart README values table. Signed-off-by: Huabing (Robin) Zhao <zhaohuabing@gmail.com> |
||
|
|
914da339b4 |
feat(kubernetes): add combined topology config surface (#2074)
* feat(kubernetes): add combined topology config surface Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * docs(kubernetes): clarify topology defaults Signed-off-by: Taylor Mutch <taylormutch@gmail.com> --------- Signed-off-by: Taylor Mutch <taylormutch@gmail.com> |
||
|
|
ed0026aae1 |
fix(helm): generate namespace-aware SANs in certgen and cert-manager templates (#2062)
The certgen hook and cert-manager Certificate template hardcoded openshell.openshell.svc.cluster.local in server certificate SANs, breaking deployments in any namespace other than openshell. Use .Release.Namespace in the templates so the SANs match the actual service FQDN regardless of the target namespace. Closes #2060 Signed-off-by: Akram <akram.benaissi@gmail.com> |
||
|
|
8cb16de9ea | chore(deploy): use OCI registry for cert-manager Helm chart (#2041) | ||
|
|
b6428cb9cd |
fix(build): align container engine selection (#1944)
Signed-off-by: Evan Lezar <elezar@nvidia.com> |
||
|
|
b6c87a76ab |
feat(server): add grpc rate limiting gateway-wide (#1566)
Signed-off-by: Adrien Langou <alangou@nvidia.com> |
||
|
|
7dab612feb |
feat(helm): support Deployment kind in HA gateway workloads (#1867)
* feat(helm): support Deployment kind in HA gateway workloads Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(helm): handle null workload values Signed-off-by: Taylor Mutch <taylormutch@gmail.com> --------- Signed-off-by: Taylor Mutch <taylormutch@gmail.com> |
||
|
|
4b44d629cf |
fix(helm): use stable gateway container name (#1864)
Signed-off-by: Taylor Mutch <taylormutch@gmail.com> |
||
|
|
702cbc4f63 |
feat(providers): support SPIFFE-backed token grants (#1784)
* feat(providers): support SPIFFE-backed token grants Add provider profile token_grant metadata and expand endpoint-specific dynamic credentials so sandbox supervisors can request SPIFFE JWT-SVIDs, exchange them with an OAuth-style token endpoint, cache returned access tokens, and inject bearer tokens into matching HTTP requests. Wire Kubernetes and Helm deployments to mount the provider SPIFFE Workload API socket into sandbox pods for token grant exchange. Signed-off-by: Taylor Mutch <taylormutch@gmail.com> Signed-off-by: Gordon Sim <gsim@redhat.com> * test(examples): add SPIFFE token grant demo Add a reusable alpha/beta demo that deploys a SPIFFE-verifying token issuer and protected services, imports a token-grant provider profile, creates a sandbox, and verifies endpoint-specific bearer tokens. The script leaves Kubernetes workloads in place, deletes sandboxes through openshell unless KEEP_SANDBOX=1, and prints protected service logs as proof of life. Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(providers): harden SPIFFE token grants * fix(providers): harden dynamic token grants Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(providers): harden token grant handling --------- Signed-off-by: Taylor Mutch <taylormutch@gmail.com> Signed-off-by: Gordon Sim <gsim@redhat.com> |
||
|
|
c4ca283c1a | refactor(helm): require external postgres for ha (#1844) | ||
|
|
e26a1b1ffc |
fix(kubernetes): configure sandbox apparmor profile (#1767)
Signed-off-by: Taylor Mutch <taylormutch@gmail.com> |
||
|
|
5e32403dbc |
feat(k8s-driver): add default_runtime_class_name config for sandbox pods (#1729)
Allow operators to configure a default Kubernetes runtimeClassName that is applied to sandbox pods when the CreateSandbox request does not specify one. This avoids requiring every API caller to explicitly set the runtime class for clusters that always need a specific RuntimeClass (e.g. kata-containers, nvidia). The fallback is applied in the Kubernetes driver only — per-request values still take priority, and an empty default (the built-in) preserves existing behavior (field omitted, cluster default applies). |
||
|
|
5f58cb018f |
fix(helm): create sandbox JWT secret when cert-manager is enabled (#1700)
* fix(helm): create sandbox JWT secret under cert-manager The cert-manager install path (certManager.enabled=true, pkiInitJob.enabled=false) left the gateway StatefulSet unable to start because nothing created the openshell-jwt-keys Secret: cert-manager owns TLS Secrets but does not mint the sandbox JWT signing key, and the certgen hook only rendered when pkiInitJob.enabled was true. Separate JWT signing-key provisioning from TLS PKI provisioning: - certgen: add a --jwt-only mode that creates only the Opaque JWT signing Secret, for use when another controller owns TLS Secrets. - certgen.yaml: render the hook when pkiInitJob.enabled OR certManager.enabled is true. cert-manager takes precedence and runs the hook with --jwt-only even if pkiInitJob.enabled remains true. Remove the mutual-exclusion failure between the two values. - _helpers.tpl: add openshell.sandboxJwtSecretName, shared by the hook and the StatefulSet mount. - Update values, README, docs, architecture, and the debug-openshell-cluster skill to reflect the new precedence; the documented cert-manager install no longer needs pkiInitJob.enabled=false. Closes #1691 * fix(helm): honor cert-manager precedence for client CA volume The client CA volume logic treated pkiInitJob.enabled as proof that built-in PKI owns the client CA. With cert-manager precedence now allowing certManager.enabled=true alongside the default pkiInitJob.enabled=true, that assumption mounts the server TLS cert secret as the client CA and ignores certManager.clientCaFromServerTlsSecret=false, which can break mTLS or trust the wrong CA. Gate the pkiInitJob.enabled term with (not certManager.enabled) in all three client CA conditions (volume mount, volume definition, and secret selection) so cert-manager owns TLS when enabled. Add a Helm test suite covering built-in PKI, cert-manager shared CA, the regression config (cert-manager + clientCaFromServerTlsSecret=false + default pkiInitJob), and the no-client-CA case. |
||
|
|
d9908222f2 | feat(kubernetes): support sandbox image pull secrets (#1671) | ||
|
|
99ca85afbb |
ci(kubernetes): stabilize HA e2e setup (#1659)
* ci(kubernetes): pin mise in e2e workflow Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * ci(kubernetes): mirror postgres image for ha e2e Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * ci(kubernetes): reuse e2e workflow for ha Signed-off-by: Taylor Mutch <taylormutch@gmail.com> --------- Signed-off-by: Taylor Mutch <taylormutch@gmail.com> |
||
|
|
269dbc6d8a |
ci(kubernetes): add HA e2e workflow (#1598)
Signed-off-by: Taylor Mutch <taylormutch@gmail.com> |
||
|
|
5007042e79 |
feat(helm): add optional PostgreSQL backing store (#1579)
* feat(helm): add optional PostgreSQL backing store with Secret-based credentials - Add postgres.enabled and postgres.deploy values to control database backend (SQLite vs PostgreSQL) and subchart deployment independently. - Introduce db-secret.yaml template for Opaque Secret with assembled postgresql:// connection string injected via OPENSHELL_DB_URL env var. - Add Bitnami PostgreSQL as optional subchart dependency keyed on postgres.deploy to prevent subchart deployment in external mode. - Externalize JWT signing key file mode via sandboxJwt.secretDefaultMode with 0400 default matching upstream. - Add validation guard for postgres.deploy=true without postgres.enabled. - Add helm unit tests covering internal, external, URL-override, special character encoding, and misconfiguration error paths. - Update README with Kubernetes and OpenShift install examples for bundled and external PostgreSQL configurations. - Add helm dependency build to lint and unittest tasks. * fix(helm): add database backend docs to README.md.gotmpl and regenerate The helm-docs CI check failed because the Database backend section was added directly to README.md instead of README.md.gotmpl. Move the content to the template and regenerate so the check passes. * fix(helm): use Secret-based DB credentials and support existingSecret Replace the inline db-url stringData pattern with a proper Secret containing individual fields plus a uri key. When postgres.deploy=true the Bitnami service-binding secret is referenced directly; when deploy=false users can supply postgres.external.existingSecret to bring their own Secret, or let the chart generate one from the external field values. Also restructures the README database section for clarity, adds helm-unittest coverage for the new secret resolution paths, and fixes a markdown lint issue in the root README. * refactor(helm): move OpenShift e2e script to e2e/rust/ and add mise task Move test-openshift-scenarios.sh from deploy/helm/openshell/ci/ to e2e/rust/e2e-openshift.sh, matching the existing e2e script naming convention. Register it as `e2e:openshift` in tasks/test.toml — not wired into the `test` or `e2e` aggregates so it only runs on explicit invocation against a live OpenShift cluster. * feat(e2e): add database backend scenarios to Kubernetes e2e Extend with-kube-gateway.sh with an optional multi-scenario loop gated by OPENSHELL_E2E_KUBE_DB_SCENARIOS=1. When enabled, the script installs the Helm chart three times — SQLite (default), bundled PostgreSQL, and external PostgreSQL with existingSecret — running the full test suite against each backend. When unset, existing single-install behavior is unchanged. Also adds helm dependency build before helm install, fixing CI failures caused by the missing PostgreSQL subchart dependency. * refactor(helm): simplify PostgreSQL config to two orthogonal controls Replace postgres.deploy and postgres.external.* with two simple controls: - postgres.enabled: deploy the bundled Bitnami PostgreSQL subchart - server.externalDbSecret: name of a pre-existing Secret with a uri key Delete db-secret.yaml — the chart no longer generates Secrets from individual credential fields. Users either get the Bitnami service-binding secret (bundled) or bring their own via server.externalDbSecret. Add validation that postgres.serviceBindings.enabled must stay true when using bundled PostgreSQL, preventing a confusing runtime failure. |
||
|
|
863d2a2ea9 |
chore(helm): add missing SPDX header to gateway-config template (#1545)
* chore(helm): add missing SPDX header to gateway-config template * chore(scripts): remove helm templates from license header exclusions The bypass had no known rationale. Removing it ensures the header script covers deploy/helm/openshell/templates uniformly going forward. Signed-off-by: mesutoezdil <mesudozdil@gmail.com> --------- Signed-off-by: mesutoezdil <mesudozdil@gmail.com> |
||
|
|
a3b16c18ab | feat(auth): per-sandbox authentication to gateway (#1404) | ||
|
|
c527341d8a |
feat(k8s): make default workspace PVC storage size configurable (#1436)
The Kubernetes driver hardcoded the workspace PVC size to 2Gi. Add a workspace_default_storage_size field to KubernetesComputeConfig so operators can tune it via TOML config, Helm values, or the OPENSHELL_K8S_WORKSPACE_DEFAULT_STORAGE_SIZE environment variable. |
||
|
|
a7cd1608f3 |
docs(helm): add chart readme generation (#1437)
Signed-off-by: Taylor Mutch <taylormutch@gmail.com> |
||
|
|
b61a98dbad |
feat(gateway): add TOML configuration file (RFC 0003) (#1317)
* feat(gateway): add TOML configuration file (RFC 0003)
Introduces an opt-in --config / OPENSHELL_GATEWAY_CONFIG flag that loads a
TOML file with gateway-wide settings and per-driver tables. Source
precedence is CLI > env > file > built-in default, implemented via clap's
ValueSource so existing flags and env vars keep their priority.
Driver crates (kubernetes, docker, podman, vm) now derive Deserialize on
their config structs. SupervisorSideloadMethod gains Deserialize with
kebab-case rename. A per-driver inheritance allowlist on the loader side
overlays [openshell.gateway] shared defaults (default_image,
supervisor_image, image_pull_policy, guest_tls_*, ssh_handshake_skew_secs,
client_tls_secret_name, host_gateway_ip, enable_user_namespaces) onto
each [openshell.drivers.<name>] table before deserialization.
The Helm chart renders a new gateway-config ConfigMap and mounts it at
/etc/openshell/gateway.toml. The migrated OPENSHELL_* env entries are
dropped from the StatefulSet — only the Secret-backed
OPENSHELL_SSH_HANDSHAKE_SECRET remains. database_url stays on --db-url.
Adds examples/gateway/gateway.example.toml and updates architecture/gateway.md
with the source precedence and inheritance rules.
* docs(gateway): drop ssh_handshake_skew_secs and ssh_handshake_secret from examples
Both fields are scheduled for removal. Remove the example values and the
env-only note so the gateway.toml example and the architecture doc stop
recommending settings that will not exist much longer.
* docs(rfc): correct OPENSHELL_CONFIG to OPENSHELL_GATEWAY_CONFIG in RFC 0003
* docs(gateway): add per-driver TOML example configurations
Adds focused single-driver examples next to the comprehensive
gateway.example.toml: kubernetes, docker, podman, and microvm. Each one
demonstrates the realistic settings for that driver plus how shared
[openshell.gateway] defaults inherit into the driver table.
A new unit test (`checked_in_examples_parse`) loads every example through
the config_file loader so schema drift fails CI rather than silently
shipping a broken example.
* refactor(gateway): drop image_pull_policy from shared inheritance
Kubernetes and Podman use mutually-incompatible vocabularies for the same
TOML key:
- Kubernetes: `Always | IfNotPresent | Never` (free-form string passed
verbatim to the K8s API).
- Podman: `always | missing | never | newer` (strict lowercase enum
deserialised into `ImagePullPolicy`).
No value means the same thing in both drivers. Sharing the key at
`[openshell.gateway]` scope and inheriting it into every active driver's
table meant any value safe for one driver was either wrong or silently
dropped for the other (`IfNotPresent` → `ImagePullPolicy::Missing` after
`.unwrap_or_default()`). Operators run one driver per gateway, so the
"shared default" never pays for itself.
Make `image_pull_policy` driver-local:
- Remove the field from `GatewayFileSection` and from
`inheritable_keys()` for both Kubernetes and Podman.
- Drop the file→`RunArgs` merge for the gateway-scope key.
- Stop unconditionally clobbering the driver value with
`config.sandbox_image_pull_policy` in the runtime wiring — only apply
the CLI/env override when it was set (and, for Podman, only when it
parses into the lowercase enum).
- Move the key under `[openshell.drivers.kubernetes]` and
`[openshell.drivers.podman]` in every example, the RFC, the
architecture doc, and the Helm-rendered gateway ConfigMap.
The supervisor pull policy follows the same shape: it is K8s-only and
moves into `[openshell.drivers.kubernetes]` alongside `image_pull_policy`
in the Helm template.
* fix(gateway): address review feedback on TOML configuration
Resolves the P1 and P2 issues raised in PR #1317:
- Helm gateway ConfigMap moves `grpc_endpoint` under
`[openshell.drivers.kubernetes]` so the default install no longer fails
the gateway's `deny_unknown_fields` schema check.
- `kubernetes_config_from_file` and `podman_config_from_file` only let
the gateway-wide CLI/env `grpc_endpoint` overwrite the driver-table
value when it was actually supplied, preserving file-only configs.
- Kubernetes driver default `image_pull_policy` is now empty (was Podman
vocabulary "missing"), so default deployments let the Kubernetes API
apply its own policy instead of being rejected.
- New `disable_tls` gateway field plumbs `.Values.server.disableTls`
through the TOML ConfigMap instead of relying on env vars dropped from
the StatefulSet.
- StatefulSet pod template now carries a `checksum/gateway-config`
annotation so `helm upgrade` rolls pods when the ConfigMap changes.
- Auxiliary listener resolution preserves the full `SocketAddr` from
`health_bind_address` / `metrics_bind_address`, so a loopback-pinned
health port is not silently relocated onto the public bind address.
- `ssh_session_ttl_secs` from the file is now applied to `Config` (it
was previously accepted by the loader but never read).
New regression coverage: cli-level merge tests for the new fields plus
helm-unittest assertions for the ConfigMap shape, checksum annotation,
and `disable_tls` rendering.
* docs(gateway): consolidate gateway TOML examples into docs reference
Replaces the per-driver example files under examples/gateway/ with a
single published reference page at docs/reference/gateway-config.mdx
covering source precedence, layout, the full example, and the four
per-driver examples (Kubernetes, Docker, Podman, microVM). Drew flagged
during PR #1317 review that the examples belong with the user-facing
docs rather than in a sibling examples/ directory.
The cross-references in architecture/gateway.md and RFC 0003 are updated
to point at the new docs page; the round-trip test in config_file.rs is
removed (schema coverage stays on the inline parses_full_example test
and per-field merge tests — doc-snippet drift belongs in a separate
docs-lint, not in a cross-tree Rust unit test).
* refactor(core): move DEFAULT_K8S_NAMESPACE into K8s driver
The constant is Kubernetes-specific (used only by KubernetesComputeConfig's
Default impl) and does not belong in openshell-core. Relocate it to the
driver crate that owns the K8s vocabulary; openshell-core retains only
truly cross-cutting defaults.
* refactor(core): move Podman bridge default into Podman driver
DEFAULT_NETWORK_NAME is Podman vocabulary, consumed only by the Podman
driver. Also drops the unused DEFAULT_IMAGE_PULL_POLICY constant.
* docs(auth): scrub remaining SSH handshake secret references
Sweeps the trailing mentions left after the rebase: the gateway
config-file module doc, the Helm gateway-config ConfigMap header,
the gateway-config.mdx env-only note, and the RPM systemd unit
comment for init-gateway-env.sh.
* docs(gateway): clarify OPENSHELL_GRPC_ENDPOINT applies to all drivers
The previous comment implied the callback endpoint was Kubernetes-only,
but the value is propagated to every compute driver (Kubernetes, Docker,
Podman, VM) and must be reachable from wherever the sandbox runs.
* refactor(gateway): move driver options into config (#1394)
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* fix(e2e): regenerate gateway config via TOML for docker + podman harnesses
The gateway CLI flags moved into TOML config tables in 560550d2 (#1394),
which made every existing e2e/with-{docker,podman}-gateway.sh invocation
fail with "unexpected argument '--sandbox-namespace'" (and a long tail
of similar driver-specific options) before the gateway could even bind.
Replace the obsolete CLI flags with a synthesized
`[openshell.drivers.<driver>]` table written to `${STATE_DIR}/gateway.toml`
and passed via the new `--config` flag. Only the gateway-wide flags
that survived 560550d2 (bind-address, port, drivers, db-url, tls-*,
disable-tls, log-level, health-port) stay on the command line.
Both scripts get a small `toml_string` helper to properly TOML-quote
the values (the previous `%q` printf format produced bash-escape, not
TOML-escape). The Docker harness also corrects two field names that
diverged from the driver schema: `docker_network_name` →
`network_name`, and the supervisor binary/image plumbing now reads
through to `supervisor_bin` / `supervisor_image` in the same table.
The Podman harness drops `--ssh-gateway-port` (deleted in 560550d2 —
gRPC + SSH are multiplexed on the same port now) and substitutes
`network_name` + `gateway_port` for the obsolete `--sandbox-namespace`
(which the Podman driver never had as a typed field).
* fix(core): swap bind-only 0.0.0.0 SSH gateway host for cluster URL host
CLI's resolve_ssh_gateway treated 0.0.0.0 as a loopback "keep as-is"
when the cluster URL was also loopback, so the SSH proxy connected to
0.0.0.0:port. The unspecified address is never a valid connect target
and is not present in any TLS cert SAN, which produced BadCertificate
TLS handshake failures during `openshell sandbox create -- ...` in
docker/podman e2e (e.g. bypass_detection).
Resolution: when the server returns 0.0.0.0 or :: as the gateway host
and both endpoints are loopback, fall back to the cluster URL's host
(which the CLI is already using to reach the gateway, so it must
resolve and match the cert).
* fix(e2e): repair podman harness on macOS
Podman 5.x with the applehv/libkrun provider no longer creates the legacy
~/.local/share/containers/podman/machine/podman.sock symlink, and
`podman system service` is a Linux-only subcommand — the macOS client
delegates the API service to the VM. Both assumptions in the harness
were stale, so the script tried to start a temporary service that podman
rejected with "unknown flag: --time".
- Discover the macOS socket via `podman machine inspect` instead of the
hardcoded path.
- On Darwin, fail fast with a "start podman machine" message rather than
attempting the Linux-only `podman system service` fallback.
- Write socket_path into [openshell.drivers.podman] so the in-process
driver picks up the discovered socket; the driver reads TOML only
after the config refactor (560550d2), so OPENSHELL_PODMAN_SOCKET alone
was no longer enough.
* fix(server): use clone_from for TLS client CA assignment
clippy 1.95.0 rejects assigning the result of `Clone::clone()` to an
existing variable under `-D warnings` (`assigning_clones`). Switch to
`clone_from(&...)` to satisfy the lint and avoid the redundant
allocation.
---------
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
|
||
|
|
c94cddbfb8 |
feat(server): separate HTTPS from mTLS authentication (#1351)
Make --tls-client-ca optional and make client certificates always optional when a CA is configured. This decouples HTTPS encryption from mTLS authentication, allowing mTLS and OIDC bearer tokens to coexist as parallel authentication mechanisms. When --tls-client-ca is provided, client certificates are validated against the CA when presented but never required. Clients may connect with or without a certificate — authentication is handled at the application layer (e.g. OIDC). Two TLS modes are now supported: - HTTPS with optional mTLS (--tls-client-ca provided) - HTTPS-only (--tls-client-ca omitted) The --disable-gateway-auth flag is preserved for backward compatibility but is now a no-op. The allow_unauthenticated field has been removed from TlsConfig. The Helm chart conditionally includes the client-ca volume and env var based on whether clientCaSecretName is configured. |
||
|
|
7a0c444445 |
refactor!(auth): drop SSH handshake secret (#1274)
* refactor!(auth): drop SSH handshake secret in favor of mTLS The OPENSHELL_SSH_HANDSHAKE_SECRET / x-sandbox-secret mechanism was misnamed: it does not authenticate SSH (which flows over the RelayStream gRPC RPC and is gated by mTLS plus supervisor Unix-socket permissions). It only gated a small set of sandbox-to-gateway control-plane RPCs, and production deployments already enforce mTLS on that channel — so the shared secret was redundant. Replace the secret check with an mTLS-presence marker. Sandbox-class methods (ReportPolicyStatus, PushSandboxLogs, GetSandboxProviderEnvironment, SubmitPolicyAnalysis, GetSandboxConfig, GetInferenceBundle) accept callers without a Bearer token; the gRPC mTLS handshake is the trust boundary. Dual-auth methods treat Bearer-present as full-scope CLI access and Bearer-absent as sandbox-restricted scope via validate_sandbox_caller_update. Drops the secret from all drivers (K8s, Podman, VM), the sandbox gRPC interceptor, the Helm chart (values + pre-install hook + StatefulSet env), the RPM bootstrap script, the man pages, and the debug-openshell-cluster skill. Also removes the never-read ssh_handshake_skew_secs flag and config field. BREAKING CHANGE: --ssh-handshake-secret / OPENSHELL_SSH_HANDSHAKE_SECRET and --ssh-handshake-skew-secs / OPENSHELL_SSH_HANDSHAKE_SKEW_SECS are removed from the gateway, sandbox, and all driver binaries. The openshell-ssh-handshake K8s Secret is no longer managed by the chart; operators may delete the orphan. Deployments using --disable-gateway-auth must enforce caller authentication at the fronting proxy, since the gateway no longer validates a per-request secret on sandbox-class methods. Refs OS-174. * docs(auth): scrub residual SSH handshake secret references Sweep across docs, e2e scripts, the Podman driver README/NETWORKING notes, the RPM/Helm/setup guides, the gateway man page, and RFC 0003 to remove instructions and examples that still referenced OPENSHELL_SSH_HANDSHAKE_SECRET / --ssh-handshake-secret / ssh_handshake_skew_secs. The mechanism is gone; nothing should still suggest setting it. Negative-assertion regression tests are kept so the env var cannot silently be re-introduced. * test(drivers): drop SSH handshake secret negative-assertion tests The supporting code, env vars, CLI flags, and config plumbing are gone — these tests asserted absence of strings that no longer have any path to being set. Remove the guards from the Docker, Podman, Kubernetes, and VM driver test modules. |